Topological perception and information weighted power distribution network multi-objective optimization scheduling method and system

CN122844155APending Publication Date: 2026-09-29STATE GRID SHANDONG ELECTRIC POWER CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611059076.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-16
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]针对分布式资源大规模接入背景下配电网潮流非线性增强、局部电压越限以及传统强化学习算法在复杂物理拓扑下感知薄弱与信用分配不均的问题,提出拓扑感知与信息加权的配电网多目标优化调度方法和系统

Benefits of technology

[0044]本发明利用图卷积网络与多智能体深度强化学习方法,针对分布式资源大规模接入导致的配电网潮流非线性增强、局部电压越限以及传统模型难以有效捕捉物理拓扑特征与解决信用分配难题的问题,提出的拓扑感知与信息加权的配电网多目标优化调度方法,能够提取台区节点间的电气耦合关系与空间相关性,并基于环境基线与动态权重实现全局奖励的分解,提升智能体的全局态势感知能力以及模型在高比例非平稳工况下的收敛效率与鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122844155A_ABST
    Figure CN122844155A_ABST
Patent Text Reader

Abstract

This invention discloses a multi-objective optimization scheduling method and system for distribution networks based on topology sensing and information weighting. The method includes: acquiring the distribution network topology and related historical operational characteristic data, and preprocessing the characteristic data; constructing a multi-agent collaborative optimization decision model for the distribution network based on a cooperative Markov decision process, transforming the economic optimization scheduling task of the distribution network into multi-agent collaborative decision optimization to minimize operating costs and maintain grid security constraints; training the constructed multi-agent collaborative optimization decision model to obtain a trained decision model; and using the trained decision model for optimal distribution network scheduling. This invention can extract the electrical coupling relationship and spatial correlation between transformer substation nodes, and decompose the global reward based on environmental baselines and dynamic weights, improving the global situational awareness capability of the agents and the convergence efficiency and robustness of the model under high-proportion non-stationary operating conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of distribution network technology, and more specifically, to a method and system for multi-objective optimization scheduling of distribution networks based on topology sensing and information weighting. Background Technology

[0002] With the large-scale integration of distributed resources such as distributed photovoltaics, distributed energy storage (ESS), and electric vehicles into distribution networks, the traditional unidirectional power flow of distribution networks is gradually evolving into complex bidirectional nonlinear power flow. The grid connection of such a high proportion of stochastic power sources poses problems affecting the safe operation of the power grid, such as local voltage exceeding limits in transformer areas, source-load power imbalance, and power flow backflow. While mathematical models based on optimization algorithms are rigorous and can ensure system safety, they are computationally time-consuming when dealing with massive resources, making real-time scheduling difficult. In contrast, intelligent scheduling based on reinforcement learning possesses stronger dynamic adaptability, complex problem-solving capabilities, and self-learning mechanisms, enabling real-time strategy adjustments and exploration of unknown spaces. However, traditional multi-agent reinforcement learning algorithms, when dealing with distribution networks with complex topologies, often treat each control resource as an isolated node, ignoring the strong coupling and spatial constraints inherent in the underlying physical topology of the power grid. Furthermore, in the process of multi-device collaboration, a simple global reward mechanism cannot accurately measure the differences in the contribution of each heterogeneous resource to system stability, leading to a credit allocation problem and limiting its convergence performance under high-proportion non-stationary operating conditions. Therefore, how to achieve efficient and coordinated management and control of multiple types of distributed control resources within the distribution network has become a key issue that urgently needs to be addressed. Summary of the Invention

[0003] To address the challenges of enhanced power flow nonlinearity, local voltage exceedances, and weak perception and uneven credit allocation in complex physical topologies of distribution networks under large-scale distributed resource integration, this paper proposes a multi-objective optimization scheduling method and system for distribution networks based on topology perception and information weighting. This method utilizes graph convolutional networks to extract electrical coupling relationships and spatial correlation features between transformer substation nodes. It then combines this with a weighted evaluation mechanism consisting of environmental baselines, dynamic topology weights, and individual values ​​to achieve accurate decomposition of global rewards. This effectively enhances the agent's global situational awareness and improves the model's convergence efficiency and scheduling robustness under non-stationary conditions.

[0004] Specifically, this invention provides a multi-objective optimization scheduling method for distribution networks based on topology sensing and information weighting, comprising the following steps:

[0005] S1. Obtain the distribution network topology and related historical operation characteristic data of the distribution network, and preprocess the characteristic data;

[0006] S2. Construct a multi-agent collaborative optimization decision model for distribution networks based on collaborative Markov decision process, transforming the economic optimization scheduling task of distribution networks into multi-agent collaborative decision optimization of distribution networks, so as to minimize operating costs and maintain grid security constraints.

[0007] S3. Train the constructed multi-agent collaborative optimization decision-making model to obtain the trained decision-making model;

[0008] S4. The trained decision model is used for distribution network optimization scheduling, and the actions generated by each agent are specifically mapped to output active and reactive power output control commands for the current step to the distributed photovoltaic agent. The energy storage intelligent agent outputs charging and discharging power commands. The actual charging load control target after the linear deviation of the benchmark charging demand output by the intelligent agent of the electric vehicle charging station. This enables multi-objective optimized scheduling of the power distribution network.

[0009] Furthermore, in the multi-agent collaborative optimization decision-making model, each agent... The Actor network is based on the local observations of each agent. Graph topology embedding And add exploration noise Generate actions:

[0010] ;

[0011] Among them, graph topology embedding The method for obtaining it is as follows:

[0012] Node features With the topologically normalized Laplace matrix of the transformer area The input is a GCN encoder, which outputs a high-dimensional graph topological embedding by spatially aggregating the features of neighboring nodes. ;

[0013] Among them, the Laplace matrix , This is the topological adjacency matrix of the transformer area. This is the node degree matrix.

[0014] Furthermore, the current joint evaluation value in the multi-agent collaborative optimization decision-making model The method for obtaining it is as follows:

[0015] Calculate the current environmental baseline Dynamic topology weights and individual value To obtain the current joint evaluation value ;

[0016] ;

[0017] ;

[0018] ;

[0019] ;

[0020] in, For intelligent agents The estimated value of local action. , , Feature embedding functions representing state, action, and global information, respectively, employ an MLP layer. This is the global state; The normalization exponential function is a global transformation function that maps a set of arbitrary real vectors to a probability distribution. is the activation function used in the hidden layers of a neural network; It is determined by parameters The defined deep neural network; , The distribution represents the learnable weight matrix and bias vector of the dynamic weight network; N represents the number of agents.

[0021] Furthermore, a multi-agent collaborative optimization decision-making model for distribution networks based on collaborative Markov decision processes is constructed. This transforms the economic optimization scheduling task of the distribution network into multi-agent collaborative decision optimization, aiming to minimize operating costs and maintain grid security constraints.

[0022] To achieve the goal of economically optimizing the dispatch of the distribution network, a reinforcement learning approach is adopted. First, the task of economically optimizing the dispatch of the distribution network is transformed into a multi-agent collaborative optimization of the distribution network based on a collaborative Markov decision process. By setting energy storage, distributed photovoltaic, and electric vehicle charging station resources as independent agents, a decision model containing a joint state space and action space is constructed. Based on sharing global state information, collaborative decision-making is carried out to minimize operating costs and maintain grid security constraints.

[0023] Furthermore, the state space of the intelligent agent Based on time characteristics Global load Energy storage state of charge Distributed photovoltaic power output EV state of charge and critical node voltage and real-time electricity price Composition, represented as:

[0024] ;

[0025] Distributed photovoltaic intelligent agents in time step Action space It can be represented as follows:

[0026] ;

[0027] In the formula, and These represent the DPV agent at time step [missing information]. Those who contribute, whether they achieve something or not.

[0028] The action spaces of the energy storage agent and the electric vehicle charging station agent can be represented as follows: , , Represented as the charging and discharging load of the energy storage system, This represents the actual charging load after a linear shift from the base charging demand.

[0029] Furthermore, in the decision-making model, the agent adopts a globally shared reward signal, and the reward function... It consists of voltage safety, EV service quality incentives, electricity purchase cost, and photovoltaic integration rewards, and the reward function is shown below:

[0030] ;

[0031] ;

[0032] ;

[0033] ;

[0034]

[0035] In the formula, These are the normalization coefficients; for The cost of purchasing electricity at any given time for Purchase power from the main grid at all times. for Time-of-use electricity pricing at any given moment For time step, for Incentives for photovoltaic utilization at any time This is the incentive factor for photovoltaic utilization. For the number of photovoltaic power plants, For the first A photovoltaic power station in Actual output at any moment for Incentives for electric vehicle charging at all times Incentive factor for electric vehicle charging For the number of charging stations, For the first One charging station State of charge at time t, Penalty for exceeding voltage limits, This is the voltage over-limit penalty coefficient. For the number of nodes, For nodes exist The voltage deviation value at any given time.

[0036] A topology-sensing and information-weighted multi-objective optimization scheduling system for distribution networks, used to implement the method described in any one of claims 1-6, comprising:

[0037] The data acquisition and processing module is used to acquire the distribution network topology and related historical operation characteristics data of the distribution network, and to preprocess the characteristic data.

[0038] The multi-agent collaborative optimization decision model construction module is used to construct a multi-agent collaborative optimization decision model for distribution networks based on the collaborative Markov decision process. This transforms the economic optimization scheduling task of the distribution network into multi-agent collaborative decision optimization of the distribution network, so as to minimize operating costs and maintain grid security constraints.

[0039] The multi-agent collaborative optimization decision model training module is used to train the constructed multi-agent collaborative optimization decision model and obtain the trained decision model.

[0040] The distribution network optimization and scheduling module is used to perform distribution network optimization and scheduling using the trained decision model, and to specifically map the actions generated by each agent, outputting active and reactive power output control commands for the current step to the distributed photovoltaic agent. The energy storage intelligent agent outputs charging and discharging power commands. The actual charging load control target after the linear deviation of the benchmark charging demand output by the intelligent agent of the electric vehicle charging station. This enables multi-objective optimized scheduling of the power distribution network.

[0041] A non-transitory computer-readable storage medium storing a computer program that, when executed by a processor, implements the multi-objective optimal scheduling method for distribution networks based on topology sensing and information weighting as described above.

[0042] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the multi-objective optimization scheduling method for distribution networks based on topology sensing and information weighting as described above.

[0043] The beneficial technical effects of this invention are as follows:

[0044] This invention utilizes graph convolutional networks and multi-agent deep reinforcement learning methods to address the problems of enhanced power flow nonlinearity, local voltage exceedance, and the difficulty of traditional models in effectively capturing physical topology features and solving the credit allocation problem caused by large-scale access of distributed resources in distribution networks. The proposed topology-aware and information-weighted multi-objective optimization scheduling method for distribution networks can extract the electrical coupling relationship and spatial correlation between transformer nodes, and decompose the global reward based on the environmental baseline and dynamic weights, thereby improving the global situational awareness capability of the agents and the convergence efficiency and robustness of the model under high proportion of non-stationary operating conditions. Attached Figure Description

[0045] Figure 1 This is a flowchart illustrating the multi-objective optimization scheduling method for distribution networks based on topology sensing and information weighting provided in an embodiment of the present invention.

[0046] Figure 2 This is the power distribution network system topology provided in Embodiment 2 of the present invention;

[0047] Figure 3 This is a comparison of the reward training process of different algorithms provided in Embodiment 2 of the present invention;

[0048] Figure 4 This is a comparison of the average cost training process of different algorithms provided in Embodiment 2 of the present invention. Detailed Implementation

[0049] The proposed multi-objective optimization scheduling method for distribution networks based on topology awareness and information weighting, firstly, utilizes graph convolutional networks (GCNs) to deeply encode the physical connectivity and operating status of distribution substations. By extracting the electrical coupling relationships and spatial correlation features between nodes in the network topology, a global situational awareness capability is constructed, enhancing the model's generalization ability. Secondly, a topology-driven dynamic information weighting allocation strategy is employed. Utilizing global topology embedding features, the physical importance and adjustment sensitivity of each node under the current operating conditions are dynamically evaluated, and the global reward is precisely decomposed to each local control device, significantly improving the model's convergence efficiency and robustness when facing a high proportion of non-stationary operating conditions. Finally, through the effects of the global situational awareness and dynamic information weighting allocation strategy, the model's convergence efficiency and system robustness are effectively improved when facing a high proportion of non-stationary operating conditions, ensuring the stability of the control strategy in complex dynamic environments. The following detailed explanation of the invention is provided in conjunction with the accompanying drawings.

[0050] Example 1

[0051] This embodiment discloses a multi-objective optimization scheduling method for distribution networks based on topology sensing and information weighting, including the following steps:

[0052] S1. Obtain the distribution network topology and related historical operation characteristic data of the distribution network, and preprocess the characteristic data;

[0053] Specifically, the network topology of the target distribution area is obtained through the distribution network geographic information system or energy management system, the node connection relationship is established, and the rated technical parameters of each distributed energy hardware in the distribution network are obtained, including the rated capacity of distributed photovoltaic inverters, the rated charging and discharging power and capacity limit of energy storage systems, and the maximum dispatchable load and number of charging piles of electric vehicle charging stations; relevant operating data are obtained based on the historical operating data of the distribution network and electric vehicle charging operation platforms, including the historical basic load curve of each node, the historical output curve of distributed photovoltaics, the historical time-of-use / real-time electricity price data of the power grid, and the behavioral characteristic data of historical EV users, including vehicle entry and exit time, initial state of charge (SOC), and expected target SOC, etc.

[0054] The obtained distribution network operation characteristic data are processed using max-min normalization and used as input for model training. The normalization mapping formula is as follows:

[0055]

[0056] In the formula, This represents the normalized per-unit value of the feature, with a numerical range of [0,1]. and These are the historical maximum and historical minimum values ​​for this feature, respectively.

[0057] The preprocessed data is divided proportionally into training and testing sets.

[0058] S2. Construct a multi-agent collaborative optimization decision model for distribution networks based on collaborative Markov decision process, transforming the economic optimization scheduling task of distribution networks into multi-agent collaborative decision optimization of distribution networks, so as to minimize operating costs and maintain grid security constraints.

[0059] To achieve the goal of economically optimizing the dispatching of the distribution network, a reinforcement learning approach is adopted. First, the task of economically optimizing the dispatching of the distribution network is transformed into multi-agent collaborative optimization based on a cooperative Markov decision process. By setting energy storage, distributed photovoltaic power, and electric vehicle charging station resources as independent agents, a decision model containing a joint state space and action space is constructed. Collaborative decision-making is achieved based on shared global state information to minimize operating costs and maintain grid security constraints. The state space, action space, and reward function of the decision model are shown below:

[0060] The state space Based on time characteristics Global load Energy storage state of charge Distributed photovoltaic power output EV state of charge and critical node voltage and real-time electricity price Composition, represented as:

[0061] ;

[0062] ;

[0063] In the formula, Expressed in hours, time characteristics Used to represent Euclidean distance in feature space;

[0064] The action space of each agent includes the charging and discharging power and power control of various resources. All actions of the agents are continuous. Taking a distributed photovoltaic (DPV) agent as an example, the DPV agent's actions are continuous. Action space It can be represented as follows:

[0065] ;

[0066] In the formula, and These represent the DPV agent at time step [missing information]. Those who contribute, whether they achieve something or not.

[0067] The action spaces of the energy storage agent and the electric vehicle charging station agent can be represented as follows: , , Represented as the charging and discharging load of the energy storage system, This represents the actual charging load after a linear shift from the base charging demand.

[0068] In the decision-making model, the agent uses a globally shared reward signal and reward function. It consists of voltage safety, EV service quality incentives, electricity purchase cost, and photovoltaic integration rewards, which encourage the agent to pay more attention to voltage safety and stability constraints during policy updates. The reward function is as follows:

[0069] ;

[0070] ;

[0071] ;

[0072] ;

[0073]

[0074] In the formula, These are the normalization coefficients; for The cost of purchasing electricity at any given time for Purchase power from the main grid at all times. for Time-of-use electricity pricing at any given moment For time step, for Incentives for photovoltaic utilization at any time This is the incentive factor for photovoltaic utilization. For the number of photovoltaic power plants, For the first A photovoltaic power station in Actual output at any moment for Incentives for electric vehicle charging at all times Incentive factor for electric vehicle charging For the number of charging stations, For the first One charging station State of charge at time t, Penalty for exceeding voltage limits, This is the voltage over-limit penalty coefficient. For the number of nodes, For nodes exist The voltage deviation value at any given time;

[0075] S3. Train the constructed multi-agent collaborative optimization decision-making model to obtain the trained decision-making model; specifically, this includes the following steps:

[0076] S301. Initialize the parameters of the decision model, including the agent network. Critic Network Priority Experience Replay Pool (PER) And set the training hyperparameters, and use the Adam optimizer to update the parameters;

[0077] S302, Define the topological adjacency matrix of the transformer area as follows: The node degree matrix is The Laplace matrix is ​​processed using symmetric normalization, and the normalized adjacency matrix is... for:

[0078] ;

[0079] Input training set: distribution network topology and transformer substation topology normalized Laplace matrix Data on distributed photovoltaic power generation and residential electricity consumption, and the number of intelligent agents. Training rounds Current training round Maximum number of moves in a round Current number of steps Batch size Learning rate Discount Factor Network soft update coefficient ;

[0080] S303. Initialize the power distribution network environment and obtain the global state. Local observations of each intelligent agent Node features Among them, global state Local observations of each intelligent agent Both are state spaces at the current moment. Node features This includes data such as voltage amplitude, active power, reactive power, and node type for each node in the current distribution network.

[0081] S304. Obtain the current step number. To extract the spatial correlation features of the distribution network, the node features... With the topologically normalized Laplace matrix of the transformer area The input is a GCN encoder, which outputs a high-dimensional graph topological embedding by spatially aggregating the features of neighboring nodes. ;

[0082] Each intelligent agent The Actor network is based on the local observations of each agent. Graph topology embedding And add an action to explore noise generation:

[0083] ;

[0084] Joint execution action set Perform power flow calculations in the distribution network environment; obtain rewards. The set of states at the next moment Next time step node features and the termination mark ; transfer tuple Stored in the PER pool with the highest initial priority. ;

[0085] S305, If the experience pool Calculate the current environmental baseline Dynamic topology weights and individual value To obtain the current joint evaluation value ;

[0086] ;

[0087] ;

[0088] ;

[0089] ;

[0090] in, For intelligent agents The estimated value of local action. , , Feature embedding functions representing state, action, and global information, respectively, employ an MLP layer. This is the global state. The normalization exponential function is a global transformation function that maps a set of arbitrary real vectors to a probability distribution. This is the activation function used in the hidden layers of a neural network. It is determined by parameters The defined deep neural network. , The distribution represents the learnable weight matrix and bias vector of a dynamic weight network.

[0091] Based on priority probability from Medium sampling For each sample, calculate the sampling weight. ,Will and Input GCN encoder to obtain ,Will Input target Actor network Output the target action for the next moment. ;Will , , Input the target Critic network and calculate the joint objective value using the value decomposition formula:

[0092] ;

[0093] in, Indicates the first The joint target value of each sample. This is expressed as the total action value function of the objective. Indicates the first The immediate reward an agent receives from the power distribution network environment after performing an action. This indicates whether the current state is the end state of the task sequence.

[0094] Calculate the TD error, combined with the importance sampling weights. Minimize the TD error using Huber loss, and update the parameters using backpropagation with the Adam optimizer. :

[0095] ;

[0096] In the formula, This is expressed as the loss function of the joint Critic network. Represented as the Huber loss function, it manifests as mean squared error when the error is small, and as mean absolute error when the error is large. This represents the current Critic network.

[0097] Using the temporal difference error of the joint action value function Reverse update the priority of samples in the replay pool ;

[0098] In dynamic weights Under the scaling effect, the chain rule is used to maximize the total expected return of the system and update the parameters. :

[0099] ;

[0100] In the formula, Represented as the system's total expected revenue For intelligent agents Actor network parameters policy gradient, Indicates adjusting network parameters Post-action change, Indicates action Value of current local action after change change.

[0101] Soft update target network parameters:

[0102] ;

[0103] If not satisfied, then training rounds Increase by 1;

[0104] S306. If the current step number satisfies If the condition is not met, then repeat step S303. If not, then the current training round... Increase by 1;

[0105] S307, If the current training round If the condition is not met, repeat steps S301, S302, and S303. If not, complete the model training and output the trained model parameters.

[0106] S308. Test the trained model using the test set. The specific steps are as follows:

[0107] Only the trained Actor model needs to be used. Given its time step Local observation The formula for generating the action is:

[0108] ;

[0109] intelligent agent The score of a single test sequence is expressed as: ,run The average value of the tests was taken to measure the stability of the strategy and used as a performance evaluation metric.

[0110] After completing multiple rounds of iterative training, the performance of the Actor policy network generated in each training cycle is verified using a pre-set test set. The Actor network parameters with the best performance evaluation index are selected as the decision model for distribution network scheduling.

[0111] S4. Use the trained agents (Actors) to perform power distribution network optimization scheduling. The specific steps are as follows:

[0112] In the current actual operating time step of the distribution network The system collects various data at the current moment, including time characteristics, global base load, actual voltage of each key node, and real-time electricity price. It then processes the data using a normalization mapping formula to construct the local observations of each agent at the current moment. Node feature matrix of the entire network and graph topology embedding .

[0113] The local Actor network with completed training parameters is loaded, and each agent is in the current state. The Actor network is based on the local observations of each agent. Graph topology embedding Without introducing any exploration noise, the current optimal continuous control action is directly calculated and output through the network's forward propagation:

[0114]

[0115] The actions generated by each intelligent agent are specifically mapped, and the active and reactive power output control commands for the current step are output to the distributed photovoltaic intelligent agent respectively. The energy storage intelligent agent outputs charging and discharging power commands. The actual charging load control target after the linear deviation of the benchmark charging demand output by the intelligent agent of the electric vehicle charging station. This enables multi-objective optimized scheduling of the power distribution network.

[0116] Example 2

[0117] As an example, in this embodiment, the method of Embodiment 1 is used for multi-objective optimization scheduling of the distribution network. This embodiment adopts the IEEE European Low Voltage Test Feeder Standard System, which includes 905 distribution lines. For ease of visualization, this embodiment simplifies the distribution network by using a key node selection method to simplify it into a 123-node network topology, while preserving the system's electrical characteristics and topological connections. The system reference voltage is 0.416kV, and the safe range of the per-unit voltage values ​​for the nodes is set to [0.95, 1.05] pu. The distribution network system topology is as follows: Figure 2 As shown, to simulate the power grid dispatching process, the optimization period is set to 24 hours, the step size is 5 minutes, and there are a total of 288 decision periods.

[0118] Distributed photovoltaic (DPV), electric vehicle charging station (EVCS), and energy storage (ES) are set up within the low-voltage distribution transformer area to achieve active and reactive power interactive scheduling and realize the economical and stable operation of the system. The system includes 2 sets of DPV, 1 set of ES, and 2 sets of EVCS, which are coordinated and optimized through four intelligent agents. The specific grid connection and capacity configuration of various resources are shown in Table 1.

[0119] Table 1 Resource Distribution Location and Capacity Configuration Parameters

[0120]

[0121] To verify the effectiveness of the method proposed in this invention, relevant comparison algorithms and ablation experiments were set based on the task set in this embodiment. The comparison algorithms were DDPG algorithm, DDPG-GCN algorithm, PPO algorithm, SAC algorithm, and TD3 algorithm, and the ablation experiments were MADDPG algorithm and MADDPG-GCN algorithm.

[0122] The changes in the overall reward and average cost of the above RL algorithm during the training process are as follows: Figure 3 , Figure 4 As shown.

[0123] Table 2 presents the average reward, average operating cost, and average power purchase of each algorithm during the test period, providing a clear picture of the overall effectiveness of different control strategies.

[0124] Table 2 Comparison of comprehensive performance indicators of each algorithm

[0125]

[0126] As shown in Table 3, the proposed MADDPG-GCN-Topo algorithm exhibits the best overall performance, with an average reward of 81.09 on the test set, significantly higher than MADDPG-GCN's 78.46, TD3's 75.12, and DDPG-GCN's 75.99. In ablation experiments, the MADDPG-GCN-Topo algorithm with the added topology-aware module showed an average reward increase of 2.63 compared to the MADDPG-GCN algorithm, while the MADDPG-GCN algorithm with the added GCN module showed an average reward increase of 3.46 compared to the MADDPG algorithm, verifying the performance improvement of the proposed method.

[0127] As shown in Table 3, the proposed MADDPG-GCN-Topo algorithm exhibits the best overall performance. Its average reward on the test set reaches 81.09, significantly outperforming MADDPG-GCN (78.46), TD3 (75.12), and DDPG-GCN (75.99). Ablation experiments further demonstrate that, compared to MADDPG-GCN, adding topology-aware dynamic credit allocation improves the average reward by 2.63; while compared to the basic MADDPG algorithm, introducing the GCN module improves the average reward by 3.46. These data fully validate the effectiveness of the proposed improvement strategy in enhancing model decision-making performance, and also demonstrate the crucial role of graph convolutional networks and topology-aware dynamic credit allocation in handling complex multi-agent cooperative tasks.

[0128] Example 3

[0129] This embodiment provides a multi-objective optimization scheduling system for distribution networks based on topology sensing and information weighting, used to implement the method described, including:

[0130] The data acquisition and processing module is used to acquire the distribution network topology and related historical operation characteristics data of the distribution network, and to preprocess the characteristic data.

[0131] The multi-agent collaborative optimization decision model construction module is used to construct a multi-agent collaborative optimization decision model for distribution networks based on the collaborative Markov decision process. This transforms the economic optimization scheduling task of the distribution network into multi-agent collaborative decision optimization of the distribution network, so as to minimize operating costs and maintain grid security constraints.

[0132] The multi-agent collaborative optimization decision model training module is used to train the constructed multi-agent collaborative optimization decision model and obtain the trained decision model.

[0133] The distribution network optimization and scheduling module is used to perform distribution network optimization and scheduling using the trained decision model, and to specifically map the actions generated by each agent, outputting active and reactive power output control commands for the current step to the distributed photovoltaic agent. The energy storage intelligent agent outputs charging and discharging power commands. The actual charging load control target after the linear deviation of the benchmark charging demand output by the intelligent agent of the electric vehicle charging station. This enables multi-objective optimized scheduling of the power distribution network.

[0134] A non-transitory computer-readable storage medium storing a computer program that, when executed by a processor, implements the multi-objective optimal scheduling method for distribution networks based on topology sensing and information weighting as described above.

[0135] Furthermore, the present invention adopts the following technical solution:

[0136] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the multi-objective optimization scheduling method for distribution networks based on topology sensing and information weighting as described above.

[0137] From the above description of the embodiments, those skilled in the art will clearly understand that the facilities of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Embodiments of the present invention can be implemented using existing processors, or by dedicated processors used for this or other purposes for suitable systems, or by hardwired systems. Embodiments of the present invention also include non-transitory computer-readable storage media, comprising machine-readable media for carrying or having machine-executable instructions or data structures stored thereon; such machine-readable media can be any available medium accessible by a general-purpose or special-purpose computer or other machine with a processor. For example, such machine-readable media can include RAM, ROM, EPROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to carry or store the required program code in the form of machine-executable instructions or data structures and is accessible by a general-purpose or special-purpose computer or other machine with a processor. When information is transmitted or provided to a machine via a network or other communication connection (hardwired, wireless, or a combination of hardwired and wireless), that connection is also considered a machine-readable medium.

[0138] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will all fall within the scope of protection of the present invention.

Claims

1. A multi-objective optimization scheduling method for distribution networks based on topology sensing and information weighting, characterized in that, Includes the following steps: S1. Obtain the distribution network topology and related historical operation characteristic data of the distribution network, and preprocess the characteristic data; S2. Construct a multi-agent collaborative optimization decision-making model for distribution networks based on collaborative Markov decision-making processes, transforming the economic optimization scheduling task of distribution networks into multi-agent collaborative decision optimization of distribution networks, so as to minimize operating costs and maintain grid security constraints. S3. Train the constructed multi-agent collaborative optimization decision-making model to obtain the trained decision-making model; S4. The trained decision model is used for distribution network optimization scheduling, and the actions generated by each agent are specifically mapped to output active and reactive power output control commands for the current step to the distributed photovoltaic agent. The energy storage intelligent agent outputs charging and discharging power commands. The actual charging load control target after the linear deviation of the benchmark charging demand output by the intelligent agent of the electric vehicle charging station. This enables multi-objective optimized scheduling of the power distribution network.

2. The multi-objective optimization scheduling method for distribution networks based on topology sensing and information weighting according to claim 1, characterized in that: Each agent in the multi-agent collaborative optimization decision-making model The Actor network is based on the local observations of each agent. Graph topology embedding And add exploration noise Generate actions: ; Among them, graph topology embedding The method for obtaining it is as follows: Node features With the topologically normalized Laplace matrix of the transformer area The input is a GCN encoder, which outputs a high-dimensional graph topological embedding by spatially aggregating the features of neighboring nodes. ; Among them, the Laplace matrix , This is the topological adjacency matrix of the transformer area. This is the node degree matrix.

3. The multi-objective optimization scheduling method for distribution networks based on topology sensing and information weighting according to claim 1, characterized in that: The current joint evaluation value in the multi-agent cooperative optimization decision-making model The method for obtaining it is as follows: Calculate the current environmental baseline Dynamic topology weights and individual value To obtain the current joint evaluation value ; ; ; ; ; in, For intelligent agents The estimated value of local action. , , Feature embedding functions representing state, action, and global information, respectively, employ an MLP layer. This is the global state; The normalization exponential function is a global transformation function that maps a set of arbitrary real vectors to a probability distribution. is the activation function used in the hidden layers of a neural network; It is determined by parameters The defined deep neural network; , The distribution represents the learnable weight matrix and bias vector of the dynamic weight network; N represents the number of agents.

4. The multi-objective optimization scheduling method for distribution networks based on topology sensing and information weighting according to claim 1, characterized in that: A multi-agent collaborative optimization decision-making model for distribution networks based on collaborative Markov decision processes is constructed. This transforms the economic optimization scheduling task of the distribution network into a multi-agent collaborative decision optimization model to minimize operating costs and maintain grid security constraints. To achieve the goal of economically optimizing the dispatch of the distribution network, a reinforcement learning approach is adopted. First, the task of economically optimizing the dispatch of the distribution network is transformed into a multi-agent collaborative optimization of the distribution network based on a collaborative Markov decision process. By setting energy storage, distributed photovoltaic, and electric vehicle charging station resources as independent agents, a decision model containing a joint state space and action space is constructed. Based on sharing global state information, collaborative decision-making is carried out to minimize operating costs and maintain grid security constraints.

5. The multi-objective optimization scheduling method for distribution networks based on topology sensing and information weighting according to claim 4, characterized in that: The state space of the intelligent agent Based on time characteristics Global load Energy storage state of charge Distributed photovoltaic power output EV state of charge and critical node voltage and real-time electricity price Composition, represented as: ; Distributed photovoltaic intelligent agents in time step Action space It can be represented as follows: ; In the formula, and These represent the DPV agent at time step [missing information]. Those who contribute, whether they achieve something or not. The action spaces of the energy storage agent and the electric vehicle charging station agent can be represented as follows: , , Represented as the charging and discharging load of the energy storage system, This represents the actual charging load after a linear shift from the base charging demand.

6. The multi-objective optimization scheduling method for distribution networks based on topology sensing and information weighting according to claim 1, characterized in that: In the decision-making model, the agent uses a globally shared reward signal and a reward function. It consists of voltage safety, EV service quality incentives, electricity purchase cost, and photovoltaic integration rewards, and the reward function is shown below: ; ; ; ; In the formula, These are the normalization coefficients; for The cost of purchasing electricity at any given time for Purchase power from the main grid at all times. for Time-of-use electricity pricing at any given moment For time step, for Incentives for photovoltaic utilization at any time This is the incentive factor for photovoltaic utilization. For the number of photovoltaic power plants, For the first A photovoltaic power station in Actual output at any moment for Incentives for electric vehicle charging at all times Incentive factor for electric vehicle charging For the number of charging stations, For the first One charging station State of charge at time t, Penalty for exceeding voltage limits, This is the voltage over-limit penalty coefficient. For the number of nodes, For nodes exist The voltage deviation value at any given time.

7. A multi-objective optimization scheduling system for distribution networks based on topology sensing and information weighting, used to implement the method described in any one of claims 1-6, characterized in that, include: The data acquisition and processing module is used to acquire the distribution network topology and related historical operation characteristics data of the distribution network, and to preprocess the characteristic data. The multi-agent collaborative optimization decision model construction module is used to construct a multi-agent collaborative optimization decision model for distribution networks based on the collaborative Markov decision process. This transforms the economic optimization scheduling task of the distribution network into multi-agent collaborative decision optimization of the distribution network, so as to minimize operating costs and maintain grid security constraints. The multi-agent collaborative optimization decision model training module is used to train the constructed multi-agent collaborative optimization decision model and obtain the trained decision model. The distribution network optimization and scheduling module is used to perform distribution network optimization and scheduling using the trained decision model, and to specifically map the actions generated by each agent, outputting active and reactive power output control commands for the current step to the distributed photovoltaic agent. The energy storage intelligent agent outputs charging and discharging power commands. The actual charging load control target after the linear deviation of the benchmark charging demand output by the intelligent agent of the electric vehicle charging station. This enables multi-objective optimized scheduling of the power distribution network.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the multi-objective optimization scheduling method for distribution networks based on topology sensing and information weighting as described in any one of claims 1 to 6.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the multi-objective optimization scheduling method for distribution networks based on topology sensing and information weighting as described in any one of claims 1 to 6.