MARL-based flexible distributed energy network multi-agent collaborative optimization method

By employing a multi-agent collaborative optimization method for flexible distributed energy networks based on MARL, combined with intelligent soft switching and network topology reconfiguration, the global collaborative optimization problem of flexible interconnected power distribution systems with a high proportion of renewable energy access is solved, achieving efficient and secure distributed decision-making and user data privacy protection.

CN121769896APending Publication Date: 2026-03-31INST OF ELECTRICAL ENG CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Traditional centralized dispatch and trading models are difficult to adapt to flexible interconnected distribution systems with a high proportion of renewable energy. Existing MARL methods fail to deeply couple the physical network model, resulting in high computational complexity, neglect of user privacy in information transparency, and difficulty in achieving global collaborative optimization.

Method used

A multi-agent collaborative optimization method based on MARL for flexible distributed energy networks is adopted. By constructing a unified framework and combining intelligent soft switching and network topology reconfiguration technologies, a learning and decision-making process for multiple agents is established to ensure that the strategy follows the physical laws of the distribution network and achieves decentralized collaboration under local observation, thus protecting user data privacy.

Benefits of technology

It effectively reduces the total active power loss of the system, increases the proportion of local consumption of distributed energy, enhances the resilience of the community power grid, adapts to the uncertainty of source and load, and realizes efficient global strategy formation and safe and reliable distributed decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121769896A_ABST
    Figure CN121769896A_ABST
Patent Text Reader

Abstract

The invention provides a flexible distributed energy network multi-agent collaborative optimization method based on MARL, and the method comprises the steps: building a distributed energy network multi-agent interaction framework which comprises a local energy transaction model for distinguishing an intelligent soft switching path, a network cost model based on a power transmission distribution factor, and a power distribution network operator and user income model; a physical system operation constraint model is established and covers intelligent soft switch operation characteristics, network topology radial constraint, energy storage charging and discharging dynamic states, controllable distributed power supply and photovoltaic output limitation, translational load scheduling and DistFlow power flow constraint. And finally, constructing a collaborative optimization mechanism based on a multi-agent near-end strategy optimization algorithm, modeling the system into a partially observable Markov decision process, adopting a centralized training distributed execution framework, and guaranteeing transaction security and data privacy through block chain evidence storage, RSA encryption and an intelligent contract. The system economy and the operation flexibility are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of energy optimization, specifically relating to a multi-agent collaborative optimization method for flexible distributed energy networks based on MARL. Background Technology

[0002] With the high proportion of distributed renewable energy integrated into the grid, traditional centralized dispatching and trading models face challenges. Peer-to-peer (P2P) energy trading, as a decentralized model, has emerged, allowing direct transactions between producers and consumers to improve local consumption and economic efficiency. To coordinate massive and heterogeneous trading entities, multi-agent reinforcement learning (MARL) has been introduced as a data-driven distributed decision-making tool. Its agents learn strategies through autonomous interaction with the environment, eliminating the need for precise global models. Simultaneously, utilizing smart soft switching (SOP) and network topology reconfiguration techniques can improve the operational flexibility of distribution networks and optimize power flow distribution. However, existing research typically optimizes trading mechanisms, physical networks, and agent decision-making in a fragmented manner, or uses simplified models, making it difficult to achieve global coordination and efficient operation in complex, flexible interconnected network environments.

[0003] Traditional P2P transaction coordination methods based on game theory or optimization models have high computational complexity, making them difficult to adapt to large-scale, highly dynamic scenarios. They also typically assume complete transparency of information, neglecting user data privacy. Secondly, the application of existing MARL methods in energy networks is mostly focused on market strategy research, failing to deeply couple with the refined physical models of new distribution networks (such as smart soft-switching operation, topology constraints, and power flow security). Summary of the Invention

[0004] To address the challenges of multi-agent distributed energy collaborative optimization in flexible interconnected distribution systems with high renewable energy integration, this invention provides a multi-agent collaborative optimization method for flexible distributed energy networks based on MARL. This method designs a unified framework that ensures the learning and decision-making processes of multiple agents strictly adhere to the physical operating laws of the distribution network (including power regulation by intelligent soft switching, radial constraints of network topology, power flow safety boundaries, and energy storage dynamics), guaranteeing that all autonomously decided scheduling schemes are physically feasible and reliable. Furthermore, in systems containing heterogeneous stakeholders (such as grid operators seeking minimum network losses and diverse users seeking maximum returns), this invention utilizes MARL to achieve decentralized collaboration. This allows each agent to autonomously learn and form efficient global strategies with only local observation information, while simultaneously protecting their private data, such as load and generation data, from leakage.

[0005] To achieve the above objectives, the present invention adopts the following technical solution:

[0006] A multi-agent collaborative optimization method for flexible distributed energy networks based on MARL includes the following steps:

[0007] Step 1: Construct a multi-agent interaction framework for distributed energy networks, establish a local energy trading model that distinguishes smart soft-switching paths, a network access cost model based on power transmission distribution factors, and a revenue model that minimizes network losses for distribution system operators and maximizes economic benefits for user agents.

[0008] Step 2: Establish a physical system operation constraint model, including intelligent soft switching operation constraints, network topology radial constraints, electrochemical energy storage operation constraints, controllable distributed power generation and photovoltaic output constraints, transferable load constraints, node power balance constraints, and DistFlow power flow constraints.

[0009] Step 3: Construct a collaborative optimization mechanism based on a multi-agent near-end policy optimization algorithm. Model the system as a partially observable Markov decision process. The state space includes the grid topology, electricity price, photovoltaic forecast, load demand, energy storage charge state and point-to-point transaction state. The action space includes energy storage charging and discharging power, point-to-point transaction power, smart soft switch injection power, transferable load power and topology switching state. The reward function corresponds to the profit target of each agent. A centralized training and distributed execution framework is used to achieve decentralized collaborative optimization.

[0010] Furthermore, in step one, the local energy trading model divides energy trading paths into two categories: trading paths that do not pass through smart soft switches and trading paths that do pass through smart soft switches. Paths that do not pass through smart soft switches are transmitted only through power distribution lines, while paths that do pass through smart soft switches use smart soft switches as reference nodes. Each transaction clearly defines the energy flow path and prohibits users from trading with themselves.

[0011] Furthermore, in step one, the local energy trading model includes user trading participation status identifiers, self-trading prohibition constraints, energy balance constraints, and upper and lower limits of trading volume constraints. Among them, the energy balance constraint requires that the energy sold by the user is equal to the energy received, and the upper and lower limits of trading volume constraints restrict the energy exchange of each line to a preset range.

[0012] Furthermore, in step one, the network access cost calculation model uses the power transmission distribution factor method to calculate the distribution cost of the energy trading path set without smart soft switches and the energy trading path set with smart soft switches respectively. The model takes into account the transmission cost per unit power per unit distance and the loss cost of power transmitted through smart soft switches, and introduces parameters reflecting electrical distance for cost quantification.

[0013] Furthermore, in step one, the revenue model of the power distribution system operator calculates the sum of the products of the current amplitude of all branches and the resistance of the branches as the total active power loss of the entire network. The revenue model of the user agent includes the net power revenue from the main grid transaction, the power revenue from point-to-point transactions, the grid access fee cost, and the generation cost of the controllable distributed power source. The point-to-point transaction revenue is calculated based on the transaction benchmark price, and the generation cost of the controllable distributed power source is represented by a quadratic function.

[0014] Furthermore, in step two, the operating constraints of the smart soft switch include port active power balance constraints, port reactive power constraints, line power capacity constraints, converter power injection upper and lower limit constraints, and power exchange constraints under the connected state of the smart soft switch. Among them, the active power balance constraint takes into account the converter loss, and the power capacity constraint limits the total active power through the smart soft switch to be less than or equal to the total transaction power flowing through the smart soft switch.

[0015] Furthermore, in step two, the radial network topology constraints use binary variables to represent the branch connectivity status and node parent-child relationships. Through branch quantity constraints, node power balance constraints, and voltage support node constraints, the reconstructed distribution network topology is ensured to meet the radial operation requirements.

[0016] Furthermore, in step two, the constraints on the operation of electrochemical energy storage include the dynamic equation of the energy storage state of charge, the upper and lower limits of the energy storage charge, the upper and lower limits of the charging power, the upper and lower limits of the discharging power, and the mutual exclusion constraint of charging and discharging. The dynamic equation of the energy storage state of charge takes into account the self-loss rate, charging efficiency, and discharging efficiency.

[0017] Furthermore, in step three, the state space of the partially observable Markov decision process includes the state of the grid topology switches, the point-to-point transaction price information, the photovoltaic predicted output value, the demand of fixed loads and movable loads, the energy storage charge state, and the point-to-point transaction state information. The action space includes the energy storage charging and discharging power value, the point-to-point transaction power value, the active and reactive power injection power value of the smart soft switch, the movable load power value, and the discrete actions of the topology switches. The reward function is defined as the negative network loss value of the distribution system operator and the economic benefit value of the user agent.

[0018] Furthermore, in step three, the multi-agent proximal policy optimization algorithm adopts an Actor-Critic framework with a centralized evaluation network and a distributed execution network. A policy pruning mechanism is introduced to stabilize the training process. During the centralized training phase, each agent inputs local observation information and global state information into the evaluation network to evaluate the value of actions. The execution network outputs actions based on local observations. After experience collection, the execution network is updated by minimizing the pruning objective function, and the evaluation network is updated by minimizing the temporal difference error. The method also includes a blockchain-based weakly centralized transaction notarization mechanism. A practical Byzantine fault-tolerant consensus algorithm is used to ensure data consistency. The RSA asymmetric encryption algorithm is used to protect the confidentiality of the transmission and storage of transaction quotes and load data. The rules for network access fee calculation, transaction clearing and settlement, and revenue distribution are encoded into smart contracts for automatic execution.

[0019] Beneficial effects:

[0020] 1. This invention achieves optimal policy learning for agents by deeply integrating the MAPPO (Multi-Agent Proximal Policy Optimization) algorithm with a refined physical model, while satisfying all network security constraints. This method effectively reduces total active power loss, smooths node voltage fluctuations, and maximizes total user benefits.

[0021] 2. The P2P trading mechanism proposed in this invention works in conjunction with network flexible regulation (intelligent soft switching, topology reconfiguration) to guide energy to be preferentially consumed locally. This method can significantly increase the local consumption ratio of distributed energy sources such as photovoltaics, reduce dependence on the main grid and reverse power, and enhance the resilience of the community-level power grid.

[0022] 3. The MARL-based framework of this invention does not rely on a precise analytical model and can adapt to source load uncertainties through a data-driven approach. The distributed execution architecture makes the system easily scalable, allowing new users or devices to quickly join and participate in collaboration as new intelligent agents. Attached Figure Description

[0023] Figure 1 A schematic diagram illustrating the classification of local energy trading paths to take into account flexible interconnection;

[0024] Figure 2 This is a schematic diagram of the training process for the test system. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0026] The present invention provides a multi-agent collaborative optimization method for flexible distributed energy networks based on MARL, comprising the following steps:

[0027] Step 1. Construct a multi-agent interaction framework for distributed energy networks, including:

[0028] Step 1.1 Construct a local energy trading model:

[0029] Users who install photovoltaic (PV) or other distributed energy resources (DERs) can participate in local energy sharing, which is implemented in a peer-to-peer (P2P) trading format. Energy sharing typically occurs between users with surplus power generation and users with unmet demand, and each transaction requires a clearly defined energy flow path. In flexible interconnected networks, the introduction of smart soft switches increases the options for trading paths, but due to the different connection characteristics of distribution feeders and smart soft switches, it is necessary to distinguish whether a path passes through a smart soft switch in order to accurately calculate costs. Based on the above description, local energy trading paths can be divided into two categories, such as... Figure 1 As shown: 1, 2, 3, n, N all represent user IDs, and the arrows indicate that energy trading is bidirectional.

[0030] Transaction path without smart soft switches: Transactions only go through the power distribution line, and smart soft switches can be considered non-existent.

[0031] The transaction path via smart soft switches: the smart soft switch serves as the reference point, and the grid root node is considered to be removed.

[0032] A collection of transaction participation statuses of P2P users (i.e., P2P participants). Defined as:

[0033] (1)

[0034] in, Let be the variable representing whether node i participates in the transaction. If user i participates in the transaction, then... Otherwise, it is 0; N is the total number of users. User energy sharing must meet the following constraints:

[0035] Self-transaction prohibited: Users cannot transact with themselves, that is:

[0036] (2)

[0037] Energy balance: The energy sold and the energy received are equal, that is:

[0038] (3)

[0039] Trading volume limits: Energy exchange on each line must be within upper and lower limits, i.e.:

[0040] (4)

[0041] (5)

[0042] in, This indicates the amount of energy traded without using a smart soft switch (upper and lower limits are...). ), This indicates the amount of energy traded through smart soft switches (upper and lower limits are...). ( ), both refer to the energy from user i to user j at time h; For the set of users participating in the transaction, This represents the energy from user j to user i at time h. The total transaction volume of user i can be expressed as:

[0043] (6)

[0044] Step 1.2 Constructing the network access fee model:

[0045] Network access fees quantify the distribution costs of local energy transactions, using PTDF (Power Transfer Distribution Factor) to calculate the cost of each route. This network access fee model improves cost accuracy and is suitable for complex, large-scale systems. The integration of smart soft switches alters the network topology, requiring differentiation of transaction paths and consideration of smart soft switch loss costs.

[0046] The network access fee cost model that takes into account different transaction paths can be expressed as:

[0047] (7)

[0048] (8)

[0049] in, and These represent the sets of energy trading paths that do not use smart soft switches and those that do not. This represents the grid connection cost for energy transactions between user i and user j at time h; This represents the total grid connection cost for user i at time h in energy trading; It is the transmission cost per unit power per unit distance; It is the cost of power loss transmitted through intelligent soft switching; It is the set of all branches of the system; and These are the PTDFs of the l-th branch corresponding to the two paths; vectors The element definition satisfies: the i-th element is That is, the energy trading volume of users i and j at time h, where the j-th element is... All other elements are 0; and A parameter reflecting electrical distance.

[0050] Step 1.3 Construct a multi-agent benefit model:

[0051] The system comprises one distribution system operator (DSO) agent and N user agents.

[0052] DSO agent objective: Minimize total active power loss across the entire network.

[0053] (9)

[0054] in, For the set of all branches, Let h represent the current amplitude of branch ij during time period h. H is the branch resistance, and H is the total operating time.

[0055] User agent objective: Maximize the overall economic benefits of the i-th user. .

[0056] (10)

[0057] in: This represents the net interaction power between user i and the main grid (positive for electricity sales, negative for electricity purchases). / Main grid electricity sales price / electricity purchase price; The total P2P transaction power for user i; The benchmark price for P2P transactions is the average of the power grid purchase and sale prices; the last item is... The generation cost of a controllable distributed power source (CDG) is a quadratic function, where a, b, and c are all generation cost parameters.

[0058] Step 2. Perform physical system modeling and operational constraints, including:

[0059] Step 2.1 Constructing an intelligent soft-switching operation model:

[0060] Intelligent soft switching optimizes network power flow by precisely controlling power regulation and providing necessary reactive power compensation based on tracking of distributed energy resources and loads. Its operating constraints are as follows:

[0061] (11)

[0062] (12)

[0063] (13)

[0064] (14)

[0065] (15)

[0066] (16)

[0067] Where i and j represent the ports connected to the smart soft switch, This is a collection of all intelligent soft switches; The set of nodes involved in intelligent soft switches; This refers to the set of circuits containing intelligent soft switches; and These represent the active and reactive power injections of the intelligent soft switch at node i and during time period h, respectively. For the loss of the intelligent soft switch on the node i side, Loss rate; Indicates the capacity of the intelligent soft switch; for The presence and state of the intelligent soft switch on the branch line. Equation (11) represents the power balance requirement on the intelligent soft switch line and takes into account converter losses. Equation (13) indicates that the power must not exceed the maximum capacity of the line. Equations (14) and (15) specify the constraints on active and reactive power at the intelligent soft switch node. Energy exchange is only allowed when the intelligent soft switch is in the connected state, otherwise it is 0. Equation (16) ensures that the total active power exchanged through the intelligent soft switch must be equal to the total power traded through the intelligent soft switch.

[0068] Step 2.2 Construct network topology reconstruction constraints:

[0069] (17)

[0070] (18)

[0071] (19)

[0072] (20)

[0073] (twenty one)

[0074] In the above formula, Binary variables: if branch A value of 1 indicates connectivity, otherwise a value of 0. : Binary variable; if node It is a node The parent node is 1 if it is 1, otherwise it is 0. The set of all branches; The set of all nodes; The set of nodes that provide voltage support within the formed island. Parameters without names are auxiliary variables.

[0075] Step 2.3 Constructing an electrochemical energy storage model:

[0076] (twenty two)

[0077] (twenty three)

[0078] (twenty four)

[0079] (25)

[0080] (26)

[0081] In the above formula, It is the self-loss rate. and For charging and discharging efficiency; This is the maximum energy storage capacity; Let be the stored energy of the i-th energy storage unit at time h. It is the charging / discharging power. It is the maximum charging / discharging power.

[0082] Step 2.4 Constructing Controllable Distributed Power Generation (CDG) and Photovoltaic (PV) Models:

[0083] CDG output constraint:

[0084] (27)

[0085] In the formula: The output of the i-th CDG at time h; This represents the maximum output of the i-th CDG;

[0086] PV output: Using the maximum power point tracking (MPPT) model, the predicted value is... .

[0087] Step 2.5 Constructing the load model:

[0088] This represents the total load of the i-th user at time h, including fixed load. and transferable load .

[0089] (28)

[0090] (29)

[0091] (30)

[0092] In the above formula, For the j-th movable load of the i-th user, the selectable working period is... These are the upper and lower limits of the transferable load. This represents the total portable load of the i-th user.

[0093] Step 2.6 Construct node power balance and power flow constraints:

[0094] (31)

[0095] Step 2.7 Constructing power flow constraints (DistFlow model):

[0096] Using the DistFlow power flow model, the operating constraints for node injected power, node voltage, and branch current are as follows:

[0097] (32)

[0098] In the above formula, , / Let i be the active / reactive power of branch i,j. / For active / reactive power injection at node j, Let i be the reactance of branch i,j. Let be the voltage of the i-th node.

[0099] Step 3. Construct a multi-agent collaborative optimization method based on MARL, including:

[0100] Step 3.1 Perform Markov Decision Process (MDP) modeling:

[0101] The system is modeled as a partially observable Markov decision process (POMDP). Each agent (DSO or user) observes the local state at time t. Execute actions Receive rewards The environment transitions to the next state. .

[0102] The state space S includes the grid topology (switching states), electricity trading price, PV predicted output, fixed and movable load demand, energy storage SOC, P2P trading state, etc.

[0103] Action space A is a mixed space, where continuous actions encompass energy storage charging and discharging power. P2P transaction power Intelligent soft switching injected power Shiftable load power Discrete actions should include the state of the topology switch.

[0104] The reward function R is consistent with the objectives of each agent. The DSO reward is the negative of the network loss; the user reward is their economic gain. .

[0105] The state transition function is determined by both operational constraints and uncertainties (such as PV and load).

[0106] Step 3.2 Construct a multi-agent proximal policy optimization algorithm:

[0107] The policy network of each agent is trained using a multi-agent proximal policy optimization algorithm. By employing a framework of centralized evaluation (Critic) and distributed execution (Actor), and introducing a PPO pruning mechanism, training stability is ensured. Its core objective function is:

[0108] (33)

[0109] in, , For the estimation of the advantage function, As a clipping parameter, the `clip` function is primarily used to restrict input values ​​to between a specified lower and upper limit, truncating any values ​​exceeding the boundary value. This represents the parameters of the policy neural network. The Critic network updates by minimizing the error of the value function. The training and execution flow of the multi-agent proximal policy optimization algorithm is as follows:

[0110] Centralized training: Each agent inputs local observations and global state information (which can be obtained through communication or blockchain) into the Critic network to evaluate the value of actions. The Actor network outputs actions based on local observations.

[0111] Experience collection: The agent interacts with the environment and stores experience tuples.

[0112] Strategy Update: Utilizing the collected experience, update each Actor network by minimizing the above pruning objective function, and update the Critic network by minimizing the temporal difference error.

[0113] Distributed execution: After training, each agent can make decisions based on local observations, achieving decentralized operation.

[0114] Step 3.3: Constructing a security and privacy protection mechanism:

[0115] Blockchain-based evidence storage: Employing a weakly centralized architecture, the transaction network is constructed using blockchain technology. All P2P transaction contracts, network fee records, and key scheduling instructions are stored on the blockchain, and the PBFT (Practical Byzantine Fault Tolerance) consensus mechanism is used to ensure data immutability and consistency.

[0116] Data encryption: Employs the RSA asymmetric encryption algorithm. When transmitting sensitive information such as transaction quotes and private payload data, users use the counterparty's public key for encryption to ensure the confidentiality of data transmission and storage.

[0117] Smart contracts: The rules for calculating network access fees, transaction clearing and settlement, and revenue distribution are encoded into smart contracts, which are executed automatically, reducing human intervention and disputes.

[0118] The method model proposed in this invention is applied to the IEEE 118-node system. Figure 2 This paper presents the curve showing the change of the total reward (TRP) of a multi-agent proximal policy optimization algorithm in the collaborative optimization task of distributed energy networks over training time steps. This curve visually reflects the evolution of the overall performance of the multi-agent system and the convergence of the policy during training. In the early stages of training (approximately the first 2000 steps), the TRP value is low and fluctuates significantly, indicating that the agents are exploring the environment and learning initial policies. As the number of training steps increases, the TRP value begins to show a significant upward trend, entering a rapid climbing phase after approximately 4000 steps. This phase indicates that the agents have mastered effective policies, and the collaborative performance of the system is rapidly improving. Notably, after approximately 7000 training steps, the growth of the TRP value tends to level off, and the curve enters a relatively stable plateau. This signifies that policy learning is nearing convergence, the agent coordination behavior is stabilizing, and the overall system performance has reached a relatively optimal level. Subsequent small fluctuations reflect the algorithm's continuous exploration and fine-tuning near the optimal policy.

[0119] The test results verify that the multi-agent proximal policy optimization algorithm framework can effectively solve the multi-agent cooperative optimization problem. The agents successfully learn cooperative strategies that can significantly improve the total system benefit through interaction with the environment, providing key empirical evidence for the effectiveness of the method in the simulation environment.

[0120] In summary, this invention presents a distributed energy collaborative management method integrating multi-agent reinforcement learning and flexible network interconnection. By constructing a multi-agent interaction model and employing a multi-agent near-end policy optimization algorithm to achieve decentralized collaborative decision-making, and by coupling physical models such as intelligent soft switching and energy storage with blockchain security mechanisms, it effectively solves the challenges of economic efficiency, security, and flexibility in system operation under high-proportion renewable energy access.

Claims

1. A MARL-based flexible distributed energy network multi-agent collaborative optimization method, characterized in that, The method comprises the following steps: Step one: constructing a multi-agent interaction framework of a distributed energy network, establishing a local energy transaction model distinguishing intelligent soft switch paths, a cross-network fee cost model based on a power transmission distribution factor, and a benefit model of a power distribution system operator minimizing network loss and a user agent maximizing economic benefits; Step two: establishing a physical system operation constraint model, including intelligent soft switch operation constraints, network topology radial constraints, electrochemical energy storage operation constraints, controllable distributed power and photovoltaic output constraints, translatable load constraints, node power balance constraints, and DistFlow power flow constraints; Step three: constructing a collaborative optimization mechanism based on a multi-agent proximal policy optimization algorithm, modeling the system as a partially observable Markov decision process, with a state space including grid topology, transaction price, photovoltaic prediction, load demand, energy storage state of charge, and point-to-point transaction state, an action space including energy storage charging and discharging power, point-to-point transaction power, intelligent soft switch injection power, translatable load power, and topology switch state, and a reward function corresponding to the benefit objectives of each agent, and implementing decentralized collaborative optimization using a centralized training and distributed execution framework.

2. The MARL-based flexible distributed energy network multi-agent collaborative optimization method according to claim 1, characterized in that, In step one, the local energy transaction model divides the energy transaction path into two categories: a transaction path that does not pass through an intelligent soft switch and a transaction path that passes through an intelligent soft switch. The path that does not pass through the intelligent soft switch is transmitted only through the distribution line, and the path that passes through the intelligent soft switch uses the intelligent soft switch as the reference node. Each transaction explicitly specifies the energy flow path and prohibits users from transacting with themselves.

3. The MARL-based flexible distributed energy network multi-agent collaborative optimization method according to claim 2, characterized in that, In step one, the local energy transaction model includes user transaction participation state identification, self-transaction prohibition constraints, energy balance constraints, and transaction volume upper and lower limit constraints. The energy balance constraint requires that the energy sold by a user be equal to the energy received, and the transaction volume upper and lower limit constraint limits the energy exchange on each line within a predetermined range.

4. The MARL-based flexible distributed energy network multi-agent collaborative optimization method according to claim 1, characterized in that, In step one, the cross-network fee cost calculation model uses the power transmission distribution factor method to calculate the distribution cost of the energy transaction path set that does not pass through the intelligent soft switch and the energy transaction path set that passes through the intelligent soft switch. The model takes into account the transmission cost per unit power per unit distance and the loss cost of power transmission through the intelligent soft switch, and introduces a parameter reflecting the electrical distance to quantify the cost.

5. The MARL-based flexible distributed energy network multi-agent collaborative optimization method according to claim 1, characterized in that, In step one, the benefit model of the power distribution system operator calculates the product of all branch current amplitudes and branch resistances as the total active network loss. The benefit model of the user agent includes net power benefits from transactions with the main grid, point-to-point transaction power benefits, cross-network fee costs, and controllable distributed power generation costs. The point-to-point transaction benefit is calculated based on the transaction benchmark price, and the controllable distributed power generation cost is represented by a quadratic function.

6. The MARL-based flexible distributed energy network multi-agent collaborative optimization method according to claim 1, characterized in that, In step two, the intelligent soft switch operation constraints include port active power balance constraints, port reactive power constraints, line power capacity constraints, converter power injection upper and lower limit constraints, and power exchange constraints under the connected state of the intelligent soft switch. The active power balance constraint takes into account the converter loss, and the power capacity constraint limits the total active power through the intelligent soft switch to be less than or equal to the total transaction power through the intelligent soft switch.

7. The MARL-based flexible distributed energy network multi-agent collaborative optimization method according to claim 1, characterized in that, In step two, the radial constraint of network topology is represented by binary variables to indicate the branch connectivity state and the parent-child relationship of nodes, and the radial operation requirement of the reconstructed distribution network topology is ensured by branch number constraint, node power balance constraint and voltage support node constraint.

8. The MARL-based flexible distributed energy network multi-agent collaborative optimization method according to claim 1, characterized in that, In step two, the electrochemical energy storage operation constraint includes the state of charge dynamic equation, the upper and lower limits of state of charge, the upper and lower limits of charging power, the upper and lower limits of discharging power, and the charging and discharging exclusion constraint, wherein the state of charge dynamic equation takes into account the self-loss rate, charging efficiency and discharging efficiency.

9. The MARL-based flexible distributed energy network multi-agent collaborative optimization method according to claim 1, characterized in that, In step three, the state space of the partially observable Markov decision process includes the grid topology switch state, the point-to-point transaction electricity price information, the photovoltaic predicted output value, the fixed load and the movable load demand, the state of charge of the energy storage and the point-to-point transaction state information, the action space includes the energy storage charging and discharging power value, the point-to-point transaction power value, the intelligent soft switch active and reactive power injection power value, the movable load power value and the topology switch discrete action, and the reward function is defined as the negative network loss value of the distribution system operator and the economic benefit value of the user agent.

10. The MARL-based flexible distributed energy network multi-agent collaborative optimization method according to claim 1, characterized in that, In step three, the multi-agent near-optimal strategy optimization algorithm adopts the Actor-Critic framework of centralized evaluation network and distributed execution network, introduces the policy clipping mechanism to stabilize the training process, inputs the local observation information and global state information into the evaluation network to evaluate the action value in the centralized training stage, and outputs the action according to the local observation by the execution network. After experience collection, the execution network is updated by minimizing the clipping objective function, and the evaluation network is updated by minimizing the time difference error; the method also includes a weak centralized transaction storage mechanism based on blockchain, adopts a practical Byzantine fault tolerance consensus algorithm to ensure data consistency, uses RSA asymmetric encryption algorithm to protect the transmission and storage confidentiality of transaction quotes and load data, and encodes the over-the-network fee calculation, transaction clearing and settlement, and income distribution rules into smart contracts for automatic execution.

Citation Information

Cited By

  • Energy storage configuration optimization method and system for flexible interconnection power distribution network

    CN122068518A