Network-variable-charge voltage cooperative control method based on multi-agent deep reinforcement learning

Through the network-variable-charge voltage collaborative control method of deep reinforcement learning of multiple agents, combined with implicit communication and near-end strategy optimization, the problem of voltage control in traditional distribution networks is solved, voltage stability and economy are improved, and new energy consumption capacity and grid reliability are enhanced.

CN120357425APending Publication Date: 2025-07-22CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510197813.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When traditional distribution network voltage control methods face high permeability distributed power supplies and complex loads, it is difficult to effectively coordinate voltage fluctuations and current characteristics. The existing multi-agent reinforcement learning algorithms have not made full use of implicit communication to optimize global decision-making.

Method used

The network-variable-load voltage collaborative control method with deep reinforcement learning of multiple agents is adopted. By building a three-layer distribution network environment, the value decomposition and near-end strategy optimization algorithm of implicit communication are used, and the coordinated control of the distribution network, transformer and load is achieved by combining flexible interconnection strategies and distributed power regulation.

Benefits of technology

It has achieved the stability and economic improvement of distribution network voltage, reduced line losses, improved the new energy consumption capacity and the reliability of power grid operation, reduced dependence on real-time communication, and ensured global optimization and computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120357425A_ABST
    Figure CN120357425A_ABST
Patent Text Reader

Abstract

The invention discloses a network-transformer-load voltage cooperative control method based on multi-agent deep reinforcement learning, and the method comprises the steps: constructing a three-layer power distribution network environment of a power distribution network side, a transformer side and a load side, and setting agents in the three layers; establishing a power distribution network voltage optimization control model, controlling actions of a plurality of equipment at three layers of a power distribution network side, a transformer side and a load side, and taking minimization of voltage deviation, network loss and action cost as a target function; modeling a network-variable-load three-layer collaborative voltage optimization control problem into a Markov decision process, and defining a corresponding action space and a state space; optimizing a global optimal strategy of each agent by using value decomposition of implicit communication; and training and solving the agents by using a multi-agent near-end strategy optimization algorithm to realize collaborative optimization among the agents so as to realize global voltage collaborative control. According to the method, voltage optimization control of the power distribution network is realized, and the operation stability, the calculation efficiency, the convergence speed and the global optimality of the power grid are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distribution network control, and particularly relates to a network-transformer-load voltage coordinated control method based on multi-agent deep reinforcement learning. Background Art

[0002] With the large-scale access of distributed power sources (DGs) such as photovoltaic and wind turbines, complex loads such as electric vehicles (EVs), and reactive power compensation devices such as capacitors, the traditional power grid has gradually evolved into a complex network structure with high-penetration DGs. This evolution not only changes the power flow characteristics of the power grid but also significantly increases the difficulty of voltage control. On the one hand, the uncertainty and intermittency of DG access and the dynamic response ability of large-scale loads make voltage fluctuations frequent; on the other hand, there are a large number of voltage regulating reactive power devices such as on-load tap changers (OLTCs), capacitor banks (CBs), and static var compensators (SVCs) in today's distribution network, which pose higher requirements for traditional voltage control. Currently, the traditional distribution network voltage control methods mainly include centralized, distributed, and hybrid control. The centralized voltage control method usually relies on the global information of the power grid and uniformly schedules the entire network through a centralized optimization algorithm. The distributed voltage control method deploys multiple intelligent control nodes in different regions of the power grid. Each node makes independent decisions based on local information and communicates with adjacent nodes for coordination to achieve the voltage regulation goal. Both of the above two control methods have certain drawbacks. Therefore, the research on hybrid voltage control is more valuable for voltage control of today's complex distribution network.

[0003] With the introduction of technologies such as intelligent algorithms and data-driven, the deep reinforcement learning algorithm has gradually become a promising voltage control solution. Multi-agent reinforcement learning (MARL) is more prominent in the voltage optimization of the distribution network due to its decision-making and collaborative optimization ability of centralized training and decentralized execution (CTDE). Currently, in the research on the MARL framework for distribution network voltage optimization, it is all based on explicit communication. The learning strategy of each agent is its own historical observation information, without considering the relationship between agents. In the distributed decision-making problem of the power network, the value decomposition of implicit communication can optimize the global decision-making function, enabling each agent to operate independently while maintaining global coordination. Summary of the Invention

[0004] The present invention comprehensively considers the influencing factors such as distribution network reconfiguration, distributed power source access, transformer tap adjustment, reactive power compensation equipment operation, and load fluctuation, and proposes a network-transformer-load voltage coordinated control method based on multi-agent deep reinforcement learning. By using value decomposition of implicit communication of multi-agents and combining with the Proximal Policy Optimization (PPO) algorithm, the equipment in the three layers of network-transformer-load is coordinated to control, so as to realize the voltage optimization control of the distribution network.

[0005] The technical solution adopted by the present invention is as follows:

[0006] A network-transformer-load voltage coordinated control method based on multi-agent deep reinforcement learning includes the following steps:

[0007] Step 1: Construct a three-layer distribution network environment on the distribution network side, transformer side, and load side, and set agents on the three layers respectively;

[0008] Step 2: Comprehensively consider multiple levels such as the reconstruction of the power grid, the access of distributed power sources, the change of transformer taps, the access of different reactive power compensation equipment, and the fluctuation of the load, establish a model for optimizing the voltage control of the distribution network, control the actions of various equipment on the three layers of the distribution network side, transformer side, and load side, take the minimum of voltage deviation, network loss, and action cost as the objective function, and set corresponding constraint conditions;

[0009] Step 3: Model the voltage optimization control problem of the three-layer coordination of network-transformer-load into a Markov decision process, and define the corresponding action space and state space;

[0010] Step 4: Use value decomposition of implicit communication to optimize the global optimal strategy of each agent;

[0011] Step 5: Use the multi-agent proximal policy optimization algorithm to train and solve the agents, realize the collaborative optimization among the agents, so as to realize the global voltage coordinated control.

[0012] In the step 1, on the distribution network side, considering the output of DG and adopting flexible interconnection strategies such as Soft Open Point (SOP), the distribution network is reconstructed and reactive power is regulated.

[0013] The flexible interconnection strategy is to use SOP to interconnect the distribution network. SOP not only has the functions of power exchange and power flow optimization, so as to realize the reconstruction of the distribution network, but also can provide a certain amount of reactive power support for the distribution network to improve the voltage stability and operation efficiency of the power grid. The distribution network reconstruction model is as follows:

[0014] Taking the minimization of network loss and node voltage deviation as the goal, the established objective function:

[0015]

[0016] Where: ε is the line set of the distribution network; R ij is the resistance of line ij; I ij is the current of line ij; N is the set of nodes in the distribution network system; both i and j are nodes in the node set N; U i and U i,ref are the voltage amplitude and rated voltage of node i respectively;

[0017] Considering the output of DG and SOP control, the power flow equation of the distribution network can be expressed as:

[0018]

[0019] Where: P DG,i and Q DG,i are the active and reactive power outputs of DG at node i; P i,load and Q i,load are the load powers at node i respectively; P ij and Q ij are the active and reactive power flows between nodes i and j respectively.

[0020] The constraint conditions are as follows:

[0021]

[0022] U i,min ≤ Ui ≤ U i,max

[0023] I ij ≤ I ij,max

[0024] Where: P sop and Q sop are the active power output and reactive power output of SOP respectively; Q SOP,min and Q SOP,max are the minimum and maximum values of the reactive power output of SOP respectively; P SOP,min and P SOP,max are the minimum and maximum values of the active power output of SOP respectively; U i,min and U i,max are the minimum and maximum values of the voltage of node i respectively; I ij,max is the maximum value of the current of line ij.

[0025] Through the flexible adjustment of the active and reactive powers of the distributed power source on the transformer side, combined with the OLTC adjustment strategy; finally, adjust the action of SVC to quickly compensate the reactive power to suppress the voltage fluctuation;

[0026] The described OLTC adjustment strategy is rule-based control based on voltage setpoints, which is adjusted by setting the upper and lower voltage limits of the node. If the voltage is lower than the lower limit, the tap is increased; if the voltage is higher than the upper limit, the tap is decreased; if the voltage is within the reasonable range, no adjustment is made. As shown in the following formula:

[0027]

[0028] In the formula: ΔT is the tap change of the OLTC; U i is the set voltage of node i; U min and U max are the set upper and lower voltage limits respectively.

[0029] On the load side, fast response on the low-voltage side is achieved through load regulation and reactive power compensation of capacitor banks (CBs).

[0030] In step 2, the model of the distribution network voltage optimization control specifically includes:

[0031] Controlling the actions of various devices in the network-transformer-load three layers to minimize voltage deviation, network loss, and control the action cost. The constructed objective function is shown in formula (1):

[0032]

[0033] In formula (1): F is the objective function; t is the number of decision-making times; w1, w2, and w3 are the voltage deviation, network loss, and control action cost coefficients of each device respectively; U i,t and U i,ref are the voltage amplitude and rated voltage of node i at the t-th decision respectively; T is 24 hours; N is the set of nodes in the distribution network system; M is the set of each device in the distribution network system; P loss,t is the distribution network network loss, as shown in formula (2); A j,t is the action cost of the j-th device at the t-th decision, as shown in formulas (3)-(6):

[0034]

[0035]

[0036] In the formula: r ij is the resistance between node i and node j; I ij,t is the current between nodes i and j at the t-th decision; A SOP,t and A OLTC,t and A SVC,t and A CBs,t are the action costs of SOP, OLTC, SVC, and CBs at the t-th decision respectively; δ SOP and α OLTC and γSVC , λ SVC are the unit operating costs of SOP, OLTC, SVC, and CBs respectively; M SOP , M OLTC , M SVC , M CBs are the quantities of SOP, OLTC, SVC, and CBs within the distribution network respectively; H s,t is the adjustment action of the s-th SOP at the t-th decision; T g,t is the tap position of the g-th OLTC at the t-th decision; C q,t is the tap position of the q-th SVC at the t-th decision; R e,t is the tap position of the e-th CBs at the t-th decision; H s,t-1 is the adjustment action of the s-th SOP at the (t - 1)-th decision; T g,t-1 is the tap position of the g-th OLTC at the (t - 1)-th decision; C q,t-1 is the tap position of the q-th SVC at the (t - 1)-th decision; R e,t-1 is the tap position of the e-th CBs at the (t - 1)-th decision; s is the number of the SOP in a certain decision; g is the number of the OLTC in a certain decision; q is the number of the SVC in a certain decision; e is the number of the CBs in a certain decision.

[0037] The constraint conditions include power flow constraints, voltage constraints, current constraints, distributed power generation power constraints, SOP capacity constraints, distribution transformer power constraints, SVC and CBs capacity constraints, and action adjustment constraints of each device such as SOP, OLTC, SVC, and CBs;

[0038] 1) Power flow constraints:

[0039]

[0040] In the formula: P i , Q i are the active power and reactive power of node i respectively; U i , U j are the voltages at the beginning and end of line ij; G ij , B ij are the conductance and susceptance of line ij respectively; θ ij is the phase difference of line ij.

[0041] 2) Voltage and current constraints:

[0042] U i,min ≤U i,△t ≤U i,max (9);

[0043] I ij ≤Iij,max (10);

[0044] Where: U i,△t is the voltage of node i at time △t; U i,min , U i,max are the minimum and maximum values of the voltage of node i respectively; I ij,max is the maximum value of the current of line ij.

[0045] 3) Distributed generation power constraint:

[0046]

[0047] In formula (11): P DG,i , Q DG,i are the active power output and reactive power output of the DG connected to node i respectively; P DG,max , P DG,min are the upper and lower limits of the active power output of the DG respectively; Q DG,max , Q DG,min are the upper and lower limits of the reactive power output of the DG respectively.

[0048] 4) SOP capacity constraint:

[0049]

[0050] Where: P SOP,s , Q SOP,s are the active power output and reactive power output of the sth SOP respectively; P SOP,min , P SOP,max are the minimum and maximum values of the active power output of the SOP respectively; Q SOP,min , Q SOP,max are the minimum and maximum values of the reactive power output of the SOP respectively; S SOP,max is the maximum capacity of the SOP.

[0051] 5) Distribution transformer power constraint:

[0052]

[0053] In formula (14): is the load rate of the gth OLTC; is the maximum load rate of the OLTC.

[0054] 6) SVC and CBs capacity constraint:

[0055]

[0056] Where: is the reactive power output of the qth SVC; is the reactive power output of the eth CBs; They are the minimum and maximum reactive power outputs of the SVC, respectively; They are the minimum and maximum reactive power outputs of the CBs, respectively.

[0057] 7) Action adjustment constraints for each device of SOP, OLTC, SVC, and CBs:

[0058]

[0059] In the formula: Ω SOP,s,max , Ω OLTC,g,max , Ω SVC,q,max , Ω CBs,e,max are the upper limits of the action times of the sth SOP, the gth OLTC, the qth SVC, and the eth CBs within a day, respectively; H s,t+1 is the adjustment action of the sth SOP at the (t + 1)-th decision; T g,t+1 is the tap position of the gth OLTC at the (t + 1)-th decision; C q,t+1 is the tap position of the qth SVC at the (t + 1)-th decision; R e,t+1 is the tap position of the eth CBs at the (t + 1)-th decision.

[0060] In step 3, the voltage optimization control problem of the network-transformer-load three-layer coordinated distribution network is modeled as a Markov decision process, which is represented as an eight-tuple <η, S, A, P, Ω, O, R, γ>; where, η represents the number of agents; S represents the global state information; A represents the set of action spaces; P: S × A × S → [0, 1] represents the state transition probability function; Ω represents the observation function of the initial state; O represents the set of observation space information; R represents the reward function; γ represents the discount factor;

[0061] 1) Agent state and observation space:

[0062] In the multi-agent reinforcement learning framework, the state space S is the global state of the system at each moment. To meet the requirements of distributed decision-making, each agent only observes the local state and neighbor information. In the present invention, the state set includes all the state information of the network-transformer-load three layers, including the voltage amplitudes, active and reactive powers of each node, the actions of each device, and the provided power.

[0063] The observation space O of the agent = [X s , Y s , Z s , where, X s , Y s , Z s are the state information of the network side, transformer side, and load side of the distribution network, respectively. The observation space of agent m is O m , and O m ∈ O.

[0064] The relevant introduction to the observation space is as follows:

[0065]

[0066] In formula (21): U i,t , P i,t , Q i,t are respectively the voltage, active power, and reactive power of node i at the t-th decision of the agent; Q DG,t , P DG,t are respectively the active and reactive powers output by the DG; P SOP,s,t , Q SOP,s,t , Ω SOP,s,t are respectively the active power, reactive power, and cumulative action times output by the s-th SOP.

[0067] Y s = {U T,g,t , Q SVC,q,t , Ω OLTC,g,t , Ω SVC,q,t} (22);

[0068] In formula (22): U T,g,t , Q SVC,q,t , Ω OLTC,g,t , Ω SVC,q,t are respectively the output voltage of the g-th transformer, the reactive power provided by the q-th SVC, and the cumulative action times of the g-th OLTC and the q-th SVC at the t-th decision of the agent.

[0069] Z s = {U Load,t , Q CBs,e,t , P Load,t , Q Load,t , Ω CBs,e,t} (23);

[0070] In formula (23): U Load,t , Q CBs,e,t , P Load,t , Q Load,t , Ω CBs,e,t are respectively the load terminal voltage, the reactive power provided by the e-th CBs, the active power and reactive power at the load terminal, and the cumulative action times of the e-th CBs at the t-th decision of the agent.

[0071] 2) Agent action space:

[0072] The action space of the agent A = [X A , Y A , Z A , the action space of agent m is a m , and am ∈ A. Among them, X A , Y A , Z A are the action space information of the grid side, transformer side, and load side of the distribution network respectively, as follows:

[0073] X A = [△Q DG,t , △P DG,t , △P SOP,s,t , △Q SOP,s,t , H s,t (24);

[0074] In formula (24): △Q DG,t , △P DG,t are the active and reactive power outputs adjusted by the DG during the t-th decision-making of the agent respectively; △P SOP,s,t , △Q SOP,s,t , H s,t are the active power output, reactive power output, and action of the s-th SOP adjusted by the agent during the t-th decision-making respectively.

[0075] Y A = [△Q SVC,q,t , T g,t , C q,t (25);

[0076] In formula (25): △Q SVC,q,t , T g,t , C q,t are the reactive power output, the tap position of the g-th OLTC of the transformer, and the tap position of the q-th SVC adjusted by the agent during the t-th decision-making respectively.

[0077] Z A = [△Q CBs,e,t , △P Load,t , △Q Load,t , R e,t (26);

[0078] In formula (26): △Q CBs,e,t , R e,t are the reactive power output and tap position adjusted by the e-th CBs by the agent during the t-th decision-making respectively; △P Load,t , △Q Load,t are the active power output and reactive power output adjusted at the load end by the agent during the t-th decision-making respectively.

[0079] 3) Agent reward function:

[0080] The reward function R of the agent during the training process is as follows:

[0081]

[0082] In formula (27): A j is the action cost of the j-th device; Φ is the penalty coefficient, σ() is the judgment function; Ω SOP,s,t , Ω OLTC,g,t , Ω SVC,q,t , Ω CBs,e,t are the action times of the s-th SOP, the g-th OLTC, the q-th SVC, and the e-th CBs of the agent at the t-th decision-making respectively; Ω SOP,s,max , Ω OLTC,m,max , Ω SVC,q,max , Ω CBs,e,max are the maximum action times of the s-th SOP, the g-th OLTC, the q-th SVC, and the e-th CBs respectively.

[0083] In step 4, each agent generates an action according to its own local observation information by using the Actor network π θ (a|o). The agent m generates an action a m based on the local observation o θ,m and the policy network π m , as shown in formula (28):

[0084] π θ (a m |o m ) = Q m [f θm (o m )] (28);

[0085] In formula (28): π θ (a m |o m ) is the Actor network of the agent m; Q m [f θm (o m )] is the value sampled from the probability distribution of the action a m of the agent m; Q m is the response function of the agent to the current environment, used to sample the action a m from the probability distribution; f θm (o m ) is the action probability distribution obtained by the non-linear mapping of the local state through the Actor network; f θm is the non-linear mapping of the Actor network of the agent m; θ is the parameter of the Actor network;

[0086] Since each agent m can only perceive its own local state o m , but this information is not sufficient to directly guide the global optimal control, it is necessary to perform feature extraction and sharing through the implicit communication module, as shown in formula (29):

[0087] z m = E φ (o m ) (29);

[0088] In formula (29): z m is the feature embedding generated by the local state extraction network E φ ; φ is the Critic network parameter. The present invention needs to use each device for collaborative optimization, so interaction is required between each agent, that is, the agent has to discover the relationship between this agent and other agents. The present invention aggregates the local feature embeddings of all agents through the implicit communication module to form a globally shared feature to describe the global characteristics of the entire system, as shown in formula (30):

[0089]

[0090] In formula (30): z all is the globally shared feature; n is all agents.

[0091] Each agent m combines its own local state o m with the globally shared feature and distributes it to each agent through the set global reward function R, as shown in formula (31), for guiding policy optimization; each agent performs reward distribution and adjusts the policy through the shared global reward signal to generate the final action decision, ensuring the optimality of the global performance, as shown in formula (32):

[0092] R m = g(R, o m , a m ) (31);

[0093] π θ (a m | o m , z all ) = Q m [f θm (o m , z all )] (32);

[0094] In the formula: g() is the reward decomposition function; R m is the global reward function of agent m; π θ (a m | o m , z all ) is the action decision of agent m; f θm (o m , z all) The optimized local state for agent m undergoes a non-linear mapping through the Actor network to generate the corresponding action probability distribution; Q m [f θm (o m ,z all )] is the value sampled from the probability distribution of the optimized action a m of agent m. In step 5, the Multi-Agent Proximal Policy Optimization algorithm (MAPPO) adopts an Actor-Critic architecture, applying the PPO algorithm to solve the problem of optimizing the Actor network and the Critic network in the implicit communication multi-agent network-variable-load voltage optimization control problem;

[0095] The loss function L CLIP (θ) of the Actor network is the clipped policy objective function as follows:

[0096]

[0097]

[0098] R t = r t + γr t+1 + γ 2 r t+2 + … γ n r t+n (36)

[0099] Where: is the expectation of the data distribution at time step t; clip(r t (θ), 1 - ε, 1 + ε) is to limit the policy update amplitude; r t (θ) is to measure the change in the action selection probability between the current policy and the old policy, as shown in Equation (34); is to measure the advantage function of the current action, as shown in Equation (35); R t is the total reward accumulated from time step t, that is, the global reward R all , as shown in Equation (36); ε is the set clipping parameter; V φ (o t ) is the value estimate of the current state o t ; π θ (a t |o t ,z all ) is the probability of the current policy selecting action a t ; π θ′ (a t |o t ,z all ) is the probability of the new policy selecting action a t ; rt is the reward at time step t; r t+1 is the reward at time step t+1; r t+2 is the reward at time step t+2; r t+n is the reward at time step t+n; γ is the discount factor.

[0100] The goal of the Critic network is to minimize the current state value, and its value function error is:

[0101]

[0102] In Equation (37), L VF (φ) is the value function error.

[0103] To encourage the exploration of the agent, the PPO algorithm adds a policy entropy regularization term to the loss function:

[0104]

[0105] In Equation (38), H(π θ ) is the policy entropy, representing the uncertainty of action selection; is the expected value, representing the expected calculation under all possible states and actions.

[0106] The total loss function L(θ,φ) of the PPO algorithm combines the clipped policy objective, the value function error, and the entropy regularization term, specifically:

[0107] L(θ,φ) = L CLIP (θ) - c1L VF (φ) + c2H(π θ ) (39);

[0108] In Equation (39), L CLIP (θ) is the optimized objective of the clipped policy; c1 and c2 are the weight coefficients of the value function error and the entropy regularization term, used to adjust the weights of different loss terms.

[0109] The PPO algorithm updates the parameters of the Actor network and the Critic network respectively through the gradient descent method;

[0110] The update of the Actor network parameters is as follows:

[0111]

[0112] In Equation (40), α is the learning rate, used to control the update step size; is the update direction of the policy network parameters; ← is the parameter update; θ is the parameter of the Actor network.

[0113] The update of the Critic network parameters is as follows:

[0114]

[0115] In Equation (41): β is the learning rate, which is used to control the update step of the Critic network; φ are the parameters of the Critic network.

[0116] In Step 5, the MAPPO training process includes the following steps:

[0117] First, the distribution network environment updates its state in real time. The action information of devices such as SOP, DG, OLTC, SVC, CBs, and the load end is collected using the existing management system, including Ω SOP,s,t 、Ω OLTC,g,t 、Ω SVC,q,t 、Ω CBs,e,t 、△Q DG,t 、△P DG,t 、△P SOP,s,t 、△Q SOP,s,t 、H s,t 、△Q SVC,q,t 、T g,t 、C q,t 、△Q CBs,e,t 、R e,t 、△P Load,t 、△Q Load,t ; and the state parameters such as the voltage amplitude, active load, and reactive load of each key node in the distribution network are dynamically updated, including U i,t 、P i,t 、Q i,t 、Q DG,t 、P DG,t 、P SOP,s,t 、Q SOP,s,t 、U T,g,t 、Q SVC,q,t 、U Load,t 、Q CBs,e,t 、P Load,t 、Q Load,t 。

[0118] Second, the agent samples the collected operation state data and stores it in the experience replay buffer, and conducts centralized training by randomly extracting data from the buffer;

[0119] Then, training and solving are performed through value decomposition MAPPO of implicit communication. Each agent makes a decision based on its own local observation information. Through the implicit communication mechanism, the agents indirectly interact in the shared environment and obtain decision compensation feedback from the environment, such as in Equations (28)-(30);

[0120] Subsequently, the system comprehensively evaluates the behaviors of the agents, generates a global reward according to the overall goal as shown in Equation (31), and reasonably distributes it to each agent to prompt the individual to optimize its own strategy as shown in Equation (32). In this stage, the agents continuously learn and optimize through the Actor-Critic architecture of the PPO algorithm, and finally form a trained action strategy as shown in Equations (33)-(41).

[0121] Finally, each agent cancels the implicit communication. Each agent uses a distributed method to perform relevant action selections on the trained strategy as shown in Equations (24)-(26), and transmits it to the security layer for correction to ensure the safety and reliability of its actions. The corrected action set is then sent to each device in the three layers for coordinated control to achieve the optimal control of the distribution network voltage.

[0122] The technical effects of a network-transformer-load voltage coordinated control method based on multi-agent deep reinforcement learning of the present invention are as follows: 1) The present invention comprehensively considers the output characteristics of DGs on the distribution network side and utilizes the adjustable power flow capacity of the SOP to realize the dynamic reconstruction and reactive power regulation of the distribution network. By optimizing the grid topology structure, the absorption capacity of DGs is maximally improved, the line loss is effectively reduced, the voltage stability is maintained, and the economy and reliability of the grid operation are improved.

[0123] 2) The present invention optimizes the voltage from multiple levels of the distribution network, comprehensively considers the distribution network reconstruction, the access of distributed power sources, the adjustment of transformer tap changers, the operation of reactive power compensation devices, and the influence of load fluctuations, establishes a three-layer voltage optimization model of network-transformer-load with the minimum of voltage deviation, network loss, and action cost as the objective function, realizes voltage stability, loss reduction, cost optimization, and efficient utilization of new energy, improves the economy, reliability, and intelligent level of the grid, and has wide engineering application value.

[0124] 3) The present invention transforms the network-transformer-load voltage optimization control problem into an implicit communication value decomposition multi-agent Markov decision process. For the topology structure reconstruction of the distribution network and the grid connection strategy of distributed power sources, the corresponding action space and state space are defined, and they are incorporated into the multi-agent decision-making and state joint optimization process to realize the coordinated control of the distribution network, transformer, and load.

[0125] 4) In the distributed execution stage, each agent makes independent decisions only based on its own observations and the trained policy network, realizes the coordinated control of voltage regulation, tap switching, and reactive power compensation devices, and reduces the dependence on real-time communication.

[0126] 5) The multi-agent policy optimization method of centralized training - distributed execution in the present invention realizes efficient, intelligent, and dynamic voltage optimization control, ensuring the operation stability, calculation efficiency, convergence speed, and global optimality of the power grid, and providing an advanced solution for the voltage optimization of the distribution network. Brief Description of the Drawings

[0127] The present invention will be further described below in conjunction with the drawings and embodiments:

[0128] Figure 1 It is a flowchart of the collaborative control method of the present invention.

[0129] Figure 2 It is a schematic diagram of the distribution network environment.

[0130] Figure 3 It is a schematic diagram of the value decomposition MARL framework for implicit communication.

[0131] Figure 4 It is a schematic diagram of the training process of the network-transformer-load voltage control strategy based on multi-agent.

[0132] Figure 5 It is a schematic diagram of the improved distribution network topology.

[0133] Figure 6 It is a schematic diagram of the agent reward convergence situation.

[0134] Figure 7 It is a schematic diagram of the voltage of each node and time of Scheme 2.

[0135] Figure 8 It is a schematic diagram of the voltage of each node and time of Scheme 1.

[0136] Figure 9 It is a voltage curve diagram of 18 nodes of each scheme. Detailed Embodiment

[0137] The present invention comprehensively considers the influence of multiple factors such as distribution network reconstruction, distributed power source access, transformer tap adjustment, reactive power compensation equipment operation, and load fluctuation, and proposes a network-transformer-load voltage collaborative control strategy based on multi-agent deep reinforcement learning; uses the value decomposition of multi-agent implicit communication and combines the Proximal Policy Optimization (PPO) algorithm to perform collaborative control on each device in the three layers of network-transformer-load to achieve voltage optimization control of the distribution network. As Figure 1 shown. It includes the following steps:

[0138] Step 1: Construct a distribution network environment at three levels of the distribution network side, transformer side, and load side, and set multiple agents at each of the three layers to achieve voltage control of the entire network through hierarchical cooperation, as Figure 2as shown;

[0139] Step 2: Considering multiple levels such as power grid reconstruction, access of distributed power sources, change of transformer tap positions, access of different reactive power compensation devices, and load fluctuations, establish a voltage optimization control model for the distribution network, control the actions of various devices in the three layers of "network-transformer-load", take the minimum of voltage deviation, network loss, and action cost as the objective function, and set corresponding constraint conditions;

[0140] Step 3: Model the voltage optimization control problem with three-layer coordination of network-transformer-load into a Markov decision process, and define the corresponding action space and state space;

[0141] Step 4: Optimize the global optimal strategy of each agent by using value decomposition of implicit communication, as Figure 3 shown;

[0142] Step 5: Use the multi-agent proximal policy optimization algorithm to train and solve the agents, realize the cooperative optimization among the agents, so as to achieve the global voltage cooperative control; as Figure 4 shown.

[0143] 1. Distribution network environment:

[0144] The control strategy proposed in the present invention is divided into three levels: the distribution network side, the transformer side, and the load side. Multiple agents are set at each of the three levels, and the voltage control of the entire network is realized through hierarchical cooperation. The distribution network environment with a three-layer structure is as Figure 1 shown. Considering the output of DG and the distribution network reconstruction and reactive power regulation of flexible interconnection strategies represented by intelligent soft switches (SOP) on the distribution network side; on the transformer side, flexibly adjust the active and reactive power of distributed power sources and combine with the OLTC adjustment strategy, and finally adjust the action of SVC to quickly compensate reactive power to suppress voltage fluctuations; on the load side, achieve fast response on the low-voltage side through load regulation and CBs reactive power compensation. As Figure 1 shown.

[0145] 2. Voltage optimization control model based on network-transformer-load:

[0146] 2.1 Objective function:

[0147] The strategy proposed in the present invention is to control the actions of various devices in the three layers of "network-transformer-load", minimize the voltage deviation, minimize the network loss, and control the action cost. The constructed objective function is as shown in Equation (1):

[0148]

[0149] In Equation (1): F is the objective function; t is the number of decision-making times; w1, w2, and w3 are the cost coefficients of voltage deviation, network loss, and control action of each device; U i,t , U i,re f are the voltage amplitude and rated voltage of node i at the t-th decision-making time respectively; T is 24 hours; N is the set of nodes in the distribution network system; M is the set of each device in the distribution network system; P loss,t is the network loss of the distribution network, as shown in Equation (2); A j,t is the action cost of the j-th device at the t-th decision-making time, as shown in Equations (3)-(6):

[0150]

[0151]

[0152] In the formula: r ij is the resistance between node i and node j; I ij,t is the current between nodes i and j at the t-th decision-making time; A SOP,t , A OLTC,t , A SVC,t , A CBs,t are the action costs of SOP, OLTC, SVC, and CBs at the t-th decision-making time respectively; δ SOP , α OLTC , γ SVC , λ SVC are the unit action costs of SOP, OLTC, SVC, and CBs respectively; M SOP , M OLTC , M SVC , M CBs are the numbers of SOP, OLTC, SVC, and CBs in the distribution network respectively; H s,t is the adjustment action of the s-th SOP at the t-th decision-making time; T g,t is the gear position of the g-th OLTC at the t-th decision-making time; C q,t is the gear position of the q-th SVC at the t-th decision-making time; R e,t is the gear position of the e-th CBs at the t-th decision-making time. H s,t-1 is the adjustment action of the s-th SOP at the (t - 1)-th decision-making time; T g,t-1 is the gear position of the g-th OLTC at the (t - 1)-th decision-making time; C q,t-1 is the gear position of the q-th SVC at the (t - 1)-th decision-making time; R e,t-1 is the gear position of the e-th CBs at the (t - 1)-th decision-making time; s is the number of SOP in a certain decision-making; g is the number of OLTC in a certain decision-making; q is the number of SVC in a certain decision-making; e is the number of CBs in a certain decision-making.

[0153] 2.2. Constraints:

[0154] The constraints of the present invention mainly include power flow constraints, voltage constraints, current constraints, distributed power generation power constraints, SOP capacity constraints, distribution transformer power constraints, SVC and CBs capacity constraints, and operation adjustment constraints of each device such as SOP, OLTC, SVC, and CBs;

[0155] 1) Power flow constraints:

[0156]

[0157] In the formula: P i and Q i are the active power and reactive power of node i respectively; U i and U j are the head and tail voltages of line ij; G ij and B ij are the conductance and susceptance of line ij respectively; θ ij is the phase difference of line ij.

[0158] 2) Voltage and current constraints:

[0159] U i,min ≤U i,△t ≤U i,max (9);

[0160] I ij ≤I ij,max (10);

[0161] In the formula: U i,△t is the voltage of node i at time △t; U i,min and U i,max are the minimum and maximum values of the voltage of node i respectively; I ij,max is the maximum value of the current of line ij.

[0162] 3) Distributed power generation power constraints:

[0163]

[0164] In formula (11): P DG,i and Q DG,i are the active power output and reactive power output of the DG connected to node i respectively; P DG,max and P DG,min are the upper and lower limits of the active power output of the DG respectively; Q DG,max and Q DG,min are the upper and lower limits of the reactive power output of the DG respectively.

[0165] 4) SOP capacity constraints:

[0166]

[0167] Where: P SOP,s , Q SOP,s are the active power output and reactive power output of the s-th SOP respectively; P SOP,min , P SOP,max are the minimum and maximum values of the active power output of the SOP respectively; Q SOP,min , Q SOP,max are the minimum and maximum values of the reactive power output of the SOP respectively; S SOP,max is the maximum capacity of the SOP.

[0168] 5) Power constraint of the distribution transformer in the area:

[0169]

[0170] In formula (14): is the load rate of the g-th OLTC; is the maximum load rate of the OLTC.

[0171] 6) Capacity constraints of SVC and CBs:

[0172]

[0173] Where: is the reactive power output of the q-th SVC; is the reactive power output of the e-th CBs; are the minimum and maximum values of the reactive power output of the SVC respectively; are the minimum and maximum values of the reactive power output of the CBs respectively.

[0174] 7) Action adjustment constraints of each device of SOP, OLTC, SVC, and CBs:

[0175]

[0176] Where: Ω SOP,s,max , Ω OLTC,g,max , Ω SVC,q,max , Ω CBs,e,max are the upper limits of the number of action times of the s-th SOP, the g-th OLTC, the q-th SVC, and the e-th CBs within a day respectively; H s,t+1 is the adjustment action of the s-th SOP at the (t + 1)-th decision; T g,t+1 is the tap position of the g-th OLTC at the (t + 1)-th decision; C q,t+1 is the tap position of the q-th SVC at the (t + 1)-th decision; R e,t+1 is the tap position of the e-th CBs at the (t + 1)-th decision.

[0177] 3. Optimal Voltage Control of Network - Transformer - Load Based on MARL:

[0178] 3.1 Markov Decision Process:

[0179] The present invention models the optimal voltage control problem of the three - layer collaborative distribution network of network - transformer - load into a Markov decision process, which is represented as an eight - tuple <η, S, A, P, Ω, O, R, γ>;

[0180] Among them, η represents the number of agents; S represents the global state information; A represents the set of action spaces; P: S×A×S→[0,1] represents the state transition probability function; Ω represents the observation function of the initial state; O represents the set of observation space information; R represents the reward function; γ represents the discount factor;

[0181] 1) Agent State and Observation Space:

[0182] In the multi - agent reinforcement learning framework, the state space S is the global state of the system at each moment. To meet the needs of distributed decision - making, each agent only observes the local state and neighbor information. In the present invention, the state set includes all the state information of the three - layer network - transformer - load, including the voltage amplitude, active and reactive power of each node, the actions of each device, and the provided power.

[0183] The observation space O of the agent = [X s , Y s , Z s , where X s , Y s , Z s are the state information of the grid side, transformer side, and load side of the distribution network respectively. The observation space of agent m is O m , and O m ∈O.

[0184] The relevant introduction of the observation space is as follows:

[0185]

[0186] In formula (21): U i,t , P i,t , Q i,t are the voltage, active power, and reactive power of node i at the t - th decision of the agent respectively; Q DG,t , P DG,t are the active and reactive powers output by the DG respectively; P SOP,s,t , Q SOP,s,t , Ω SOP,s,t are the active power, reactive power, and cumulative action times output by the s - th SOP respectively.

[0187] Y s ={UT,g,t , Q SVC,q,t , Ω OLTC,g,t , Ω SVC,q,t}, (22);

[0188] In Equation (22): U T,g,t , Q SVC,q,t , Ω OLTC,g,t , Ω SVC,q,t are respectively the output voltage of the g-th transformer, the reactive power provided by the q-th SVC, and the cumulative action times of the g-th OLTC and the q-th SVC at the t-th decision of the agent.

[0189] Z s = {U Load,t , Q CBs,e,t , P Load,t , Q Load,t , Ω CBs,e,t}, (23);

[0190] In Equation (23): U Load,t , Q CBs,e,t , P Load,t , Q Load,t , Ω CBs,e,t are respectively the load terminal voltage, the reactive power provided by the e-th CBs, the active power and reactive power at the load terminal, and the cumulative action times of the e-th CBs at the t-th decision of the agent.

[0191] 2) Agent action space:

[0192] The action space A of the agent = [X A , Y A , Z A , and the action space a of agent m m , and a m ∈ A. Among them, X A , Y A , Z A are respectively the action space information of the grid side, transformer side, and load side of the distribution network, specifically as follows:

[0193] X A = [△Q DG,t , △P DG,t , △P SOP,s,t , △Q SOP,s,t , H s,t , (24);

[0194] In Equation (24): △Q DG,t , △P DG,t are respectively the active and reactive power outputs adjusted by the DG at the t-th decision of the agent; △P SOP,s,t , △Q SOP,s,t , Hs,t They are the active power output, reactive power output, and action of the SOP adjusted by the s-th SOP when the agent makes the t-th decision, respectively.

[0195] Y A = [ΔQ SVC,q,t , T g,t , C q,t (25);

[0196] In Equation (25): ΔQ SVC,q,t , T g,t , C q,t are the reactive power output adjusted by the q-th SVC, the tap position of the g-th OLTC, and the tap position of the q-th SVC when the agent makes the t-th decision, respectively.

[0197] Z A = [ΔQ CBs,e,t , ΔP Load,t , ΔQ Load,t , R e,t (26);

[0198] In Equation (26): ΔQ CBs,e,t , R e,t are the reactive power output and tap position adjusted by the e-th CBs when the agent makes the t-th decision, respectively; ΔP Load,t , ΔQ Load,t are the active power output and reactive power output adjusted at the load end when the agent makes the t-th decision, respectively.

[0199] 3) Agent reward function:

[0200] The reward function R of the agent during training is as follows:

[0201]

[0202] In Equation (27): A j is the action cost of the j-th device; Φ is the penalty coefficient, σ() is the judgment function; Ω SOP,s,t , Ω OLTC,g,t , Ω SVC,q,t , Ω CBs,e,t are the number of actions of the s-th SOP, the g-th OLTC, the q-th SVC, and the e-th CBs when the agent makes the t-th decision, respectively; Ω SOP,s,max , Ω OLTC,m,max , Ω SVC,q,max , Ω CBs,e,max are the maximum number of actions of the s-th SOP, the g-th OLTC, the q-th SVC, and the e-th CBs, respectively.

[0203] 3.2, Value Decomposition Multi-Agent Deep Reinforcement Learning for Implicit Communication:

[0204] The value decomposition MARL framework for implicit communication accomplishes the cooperative control of multi-agent systems in network-variable-load voltage control, and its framework is as shown in Figure 3 Figure 1. For example, each agent such as a transformer control unit, a distributed power node, etc. acts as an independent decision-making entity and can learn the globally optimal voltage control strategy through local states and implicitly shared information. The core function of the implicit communication module is to share embedded feature information. This mechanism uses local state propagation, environmental feedback, and cooperation mechanisms in the policy optimization process, enabling each agent to infer the behaviors and strategies of other agents based on its own local observations, and thus make optimal decisions. Different from explicit communication, implicit communication does not rely on direct information exchange, but realizes the indirect transmission and sharing of information through the interaction between agents and the shared environment.

[0205] Each agent generates an action according to its own local observation information using the Actor network πθ(a|o). Agent m generates an action a according to the local observation O m and the policy network π m as shown in the formula (28): m π

[0206] π θ (a m |o m ) = Q m [f θm (o m )] (28);

[0207] In formula (28): π θ (a m |o m ) is the Actor network of agent m; Q m [f θm (o m )] is the value sampled from the probability distribution of the action a m of agent m; Q m is the response function of the agent to the current environment, used to sample the action a m from the probability distribution; f θm (o m ) is the action probability distribution obtained by the non-linear mapping of the local state through the Actor network; f θm is the non-linear mapping of the Actor network of agent m; θ is the parameter of the Actor network;

[0208] Since each agent m can only perceive its own local state O m , but this information is not sufficient to directly guide the global optimization control, it is necessary to perform feature extraction and sharing through the implicit communication module, as shown in formula (29):

[0209] z m = E φ (o m ) (29);

[0210] In Equation (29): z m is the feature embedding generated by the local state extraction network E φ ; φ are the Critic network parameters. The present invention needs to utilize the cooperation and optimization of each device, so the agents need to interact with each other, that is, the agent needs to discover the relationship between this agent and other agents. The present invention aggregates the local feature embeddings of all agents through the implicit communication module to form a globally shared feature to describe the global characteristics of the entire system, as shown in Equation (30):

[0211]

[0212] In Equation (30): z all is the globally shared feature; n are all agents.

[0213] Each agent m combines its own local state o m with the globally shared feature and assigns it to each agent through the set global reward function R, as shown in Equation (31), for guiding policy optimization; each agent performs reward distribution and adjusts the policy through the shared global reward signal to generate the final action decision, ensuring the optimality of the global performance, as shown in Equation (32):

[0214] R m = g(R, o m , a m ) (31);

[0215] π θ (a m | o m , z all ) = Q m [f θm (o m , z all )] (32);

[0216] In the formula: g() is the reward decomposition function; R m is the global reward function of agent m; π θ (a m | o m , z all ) is the action decision of agent m; f θm (o m , z all ) is the non-linear mapping of the optimized local state of agent m through the Actor network to generate the corresponding action probability distribution; Qm [f θm (o m ,z all )] is the value sampled from the probability distribution of the optimized action a of the agent m. 3.2. Algorithm construction: m of.

[0217] The Multi-Agent Proximal Policy Optimization (MAPPO) algorithm is a policy gradient-based reinforcement learning algorithm and an extended deep reinforcement learning method for multi-agent systems. MAPPO enables multiple agents to efficiently learn policies in cooperative or adversarial environments through centralized training and decentralized execution. Its core idea is to optimize the policy while restricting the amplitude of policy updates, thereby ensuring the stability and efficiency of optimization.

[0218] MAPPO also adopts the Actor-Critic architecture. The difference is that the Critic network learns a centralized value function. The present invention applies the PPO algorithm to optimize the Actor network and the Critic network in solving the implicit communication multi-agent "network-variable-load" voltage optimization control problem;

[0219] The loss function L CLIP (θ) of the Actor network is the clipped policy objective function as shown below:

[0220]

[0221] R t = r t + γr t+1 + γ 2 r t+2 +… γ n r t+n (36);

[0222] In the formula: is the expectation of the data distribution at time step t; clip(r t (θ), 1 - ε, 1 + ε) is to limit the amplitude of policy updates; r t (θ) is to measure the change in the action selection probability between the current policy and the old policy, as shown in formula (34); is to measure the advantage function of the current action, as shown in formula (35); R t is the total reward accumulated from time step t, that is, the global reward R all , as shown in formula (36); ε is the set clipping parameter; V φ (o t ) is the value estimate of the current state o t ; π θ (a t |o t ,zall ) is the probability of selecting action a for the current policy; π t ; π θ′ (a t |o t , z all ) is the probability of selecting action a for the new policy; r t ; r t is the reward at time step t; r t+1 is the reward at time step t + 1; r t+2 is the reward at time step t + 2; r t+n is the reward at time step t + n; γ is the discount factor.

[0223] The goal of the Critic network is to minimize the current state value, and its value function error is:

[0224]

[0225] In Equation (37), L VF (φ) is the value function error.

[0226] To encourage the exploration of the agent, the PPO algorithm adds a policy entropy regularization term to the loss function:

[0227]

[0228] In Equation (38), H(π θ ) is the policy entropy, representing the uncertainty of action selection; is the expected value, representing the expected calculation under all possible states and actions.

[0229] The total loss function L(θ, φ) of the PPO algorithm combines the clipped policy objective, the value function error, and the entropy regularization term, specifically:

[0230] L(θ, φ) = L CLIP (θ) - c1L VF (φ) + c2H(π θ ) (39);

[0231] In Equation (39), L CLIP (θ) is the optimized objective of the clipped policy; c1 and c2 are the weight coefficients of the value function error and the entropy regularization term, used to adjust the weights of different loss terms.

[0232] The PPO algorithm updates the parameters of the Actor network and the Critic network respectively through the gradient descent method;

[0233] The update of the Actor network parameters is as follows:

[0234]

[0235] In Equation (40), α is the learning rate, which is used to control the update step size; is the update direction of the policy network parameters; ← represents parameter update; θ is the parameter of the Actor network.

[0236] The update of the Critic network parameters is as follows:

[0237]

[0238] In Equation (41): β is the learning rate, which is used to control the update step size of the Critic network; φ is the parameter of the Critic network.

[0239] The MAPPO training process is as Figure 4 shown. First, the distribution network environment updates the state in real time, collects the action information of devices such as SOP, DG, OLTC, SVC, CBs, and the load end using the existing management system, and dynamically updates the state parameters such as the voltage amplitude, active load, and reactive load of each key node in the distribution network. Secondly, the agent samples the collected operation state data and stores it in the experience replay buffer, and conducts centralized training by randomly extracting the data in the buffer. Then, through the value decomposition MAPPO of implicit communication for training and solving, each agent makes a decision based on its own local observation information. Through the implicit communication mechanism, the agents indirectly interact in the shared environment and obtain the decision compensation feedback from the environment. Subsequently, the system comprehensively evaluates the behavior of the agents, generates a global reward according to the overall goal, and reasonably distributes it to each agent to prompt the individual to optimize its own strategy. In this stage, the agent continuously learns and optimizes through the Actor-Critic architecture of the PPO algorithm, and finally forms a trained action strategy. Finally, each agent cancels the implicit communication, and each agent uses a distributed method to select relevant actions for the trained strategy and transmits it to the security layer for correction to ensure that its actions are safe and reliable. The corrected action set is then sent to each device in the three layers for coordinated control to achieve the optimal control of the distribution network voltage.

[0240] Verification example:

[0241] The present invention uses a system in which the IEEE 33-node and IEEE 15-node are flexibly interconnected through SOP to verify the proposed distribution network voltage optimization control strategy for three-layer coordination of network-transformer-load. The topological structure of the distribution system and the installation positions of each device are as Figure 5 shown.

[0242] The parameter settings of the MAPPO algorithm used in the present invention are as follows. The structures of the agents are the same. The discount factor γ is 0.99, the clipping parameter ε is 0.2, the learning rate is 0.0001, the number of algorithm training times is 1000, each episode is 24, the size of the experience replay pool is 10000, and the batch size is 64.

[0243] To verify the effectiveness of the strategy proposed in the present invention, the MAPPO algorithm under implicit communication proposed in the present invention and the following several schemes are all run in the same power distribution system, and the algorithm parameter settings used are the same.

[0244] Scheme 1: The network-transformer-load voltage coordinated control strategy of value decomposition MAPPO with implicit communication proposed in the present invention.

[0245] Scheme 2: The original scheme without considering network-transformer-load coordinated control.

[0246] Scheme 3: The network-transformer-load voltage coordinated control strategy using MAPPO under normal communication conditions.

[0247] Scheme 4: The network-transformer-load voltage coordinated control strategy using the multi-agent Deep Deterministic policy gradient (MADDPG) algorithm.

[0248] Scheme 5: The network-transformer-load voltage coordinated control strategy using the Multi-Agent Soft Actor-Critic (MASAC) algorithm.

[0249] The convergence of the agent rewards of each algorithm obtained is as Figure 6 shown. It can be Figure 6 seen that the algorithm proposed in the present invention has a faster convergence speed, better training effect, is easier to achieve collaborative tasks, and the obtained voltage control strategy is better.

[0250] Under the conditions of Scheme 2, the voltage variation of 48 nodes in a day is as Figure 7 shown. It can be seen from Figure 7 that during the period with the strongest sunlight, that is, from 11 am to 3 pm, due to the high power output of the photovoltaic power generation system, the voltage amplitudes of multiple nodes increase significantly, reaching a maximum of 1.08 p.u. At the same time, during this period, low voltage phenomena also occur at some nodes, indicating that the voltage levels are unevenly distributed throughout the system. This uneven voltage distribution not only affects the power supply quality but also poses a potential threat to the stability of the power system.

[0251] When using Scheme 1 proposed in the present invention for voltage control, the voltage variation of 48 nodes in a day is asFigure 8 As shown. From Figure 8 it can be clearly observed that the voltage values of each node are basically stable at about 1.00 p.u. in different time periods. This result indicates that the proposed scheme can effectively regulate the voltage distribution while realizing the accommodation of new energy, significantly improving the voltage stability and operation reliability of the system.

[0252] Figure 9 shows the control effects of Scheme 1, Scheme 3, Scheme 4 and Scheme 5 on the voltage of Node 18 within 24 hours of a day. It can be seen from the curve distribution that there are certain differences in the voltage regulation capabilities of each scheme in different time periods. Generally speaking, each scheme can control the node voltage between 0.992 p.u. and 1.012 p.u., meeting the basic requirements of voltage stability. Among them, the voltage amplitudes of Scheme 4 and Scheme 5 are slightly higher than those of Scheme 1 and Scheme 3, reflecting that the PPO algorithm used in the present invention is more effective than other algorithms. In addition, the curve of Scheme 1 is relatively smooth, and the voltage is closer to 1.000 p.u., showing strong voltage control stability; while the fluctuations of Scheme 3, Scheme 4 and Scheme 5 are more significant in some time periods, but the overall change range is still within a reasonable range.

Claims

1. A network-variable-load voltage coordinated control method based on multi-agent deep reinforcement learning, characterized in that It includes the following steps: Step 1: Construct a three-layer distribution network environment on the distribution network side, transformer side, and load side, and set agents on each layer; Step 2: Establish a model for voltage optimization control of the distribution network, control the actions of various devices on the distribution network side, transformer side, and load side, take the minimization of voltage deviation, network loss, and action cost as the objective function, and set corresponding constraint conditions; Step 3: Model the voltage optimization control problem of the network-transformer-load three-layer coordination into a Markov decision process, and define the corresponding action space and state space; Step 4: Optimize the global optimal strategy of each agent using value decomposition with implicit communication; Step 5: Use the multi-agent proximal policy optimization algorithm to train and solve the agents, realize the collaborative optimization among the agents, and achieve global voltage collaborative control.

2. The network-variable-load voltage coordinated control method based on multi-agent deep reinforcement learning according to claim 1, wherein: In Step 1, on the distribution network side, consider the output of DGs and adopt a flexible interconnection strategy to reconstruct and perform reactive power regulation on the distribution network; the flexible interconnection strategy is to interconnect the distribution network using intelligent soft switches; The distribution network reconstruction model is as follows: Establish an objective function with the goal of minimizing network loss and node voltage deviation: Where: ε is the line set of the distribution network; R ij is the resistance of line ij; I ij is the current of line ij; N is the node set of the distribution network system; both i and j are nodes in the node set N; U i and U i,ref are the voltage magnitude and the rated voltage of node i, respectively; Considering the output of DGs and SOP control, the power flow equation of the distribution network can be expressed as: Where: P DG,i , Q DG,i are the active and reactive power outputs of DG at node i; P i,load , Q i,load are the load powers at node i respectively; P ij , Q ij are the active and reactive power flows between nodes i and j respectively; The constraint conditions are as follows: I ij ≤I ij,max Where: P sop , Q sop are the active power output and reactive power output of the SOP respectively; Q SOP,min , Q SOP,max are the minimum value and maximum value of the reactive power output of the SOP respectively; P SOP,min , P SOP,max are the minimum value and maximum value of the active power output of the SOP respectively; U i,min , U i,max are the minimum value and maximum value of the voltage of node i respectively; I ij,max is the maximum value of the current of line ij.

3. The network-variable-load voltage coordinated control method based on multi-agent deep reinforcement learning according to claim 2, characterized in that: On the transformer side, flexibly adjust the active and reactive power of distributed power sources, and combine with the OLTC adjustment strategy; finally, adjust the action of the SVC to quickly compensate reactive power to suppress voltage fluctuations; The OLTC adjustment strategy is rule-based control based on voltage set points, and it is adjusted by setting the upper and lower limits of the node voltage. If the voltage is lower than the lower limit, the tap is increased; if the voltage is higher than the upper limit, the tap is decreased; if the voltage is within a reasonable range, no adjustment is made; as shown in the following formula: Where: ΔT is the gear position of the OLTC change; U i is the voltage of the set node i; U min , U max are the upper and lower voltage limits respectively.

4. The network-variable-load voltage coordinated control method based on multi-agent deep reinforcement learning according to claim 3, wherein: On the load side, achieve fast response on the low-voltage side through load regulation and reactive power compensation of capacitor banks.

5. The network-variable-load voltage coordinated control method based on multi-agent deep reinforcement learning according to claim 1, characterized in that: In Step 2, the model for voltage optimization control of the distribution network specifically includes: Control the actions of various devices in the network-transformer-load three-layer to minimize voltage deviation, minimize network loss, and control the action cost. The constructed objective function is as shown in Equation (1): In Equation (1): F is the objective function; t is the number of decision-making times; w1, w2, and w3 are the coefficients of voltage deviation, network loss, and control action cost of each device respectively; U i,t and U i,ref are the voltage magnitude and rated voltage of node i at the t-th decision, respectively; T is 24 hours; N is the set of nodes in the distribution network system; M is the set of devices in the distribution network system; P loss,t is the network loss of the distribution network, as shown in Equation (2); A j,t is the operation cost of the j-th device at the t-th decision, as shown in Equations (3)-(6): where: r ij is the resistance between node i and node j; I ij,t is the current between nodes i and j at the t-th decision; A SOP,t and A OLTC,t and A SVC,t and A CBs,t are the action costs of SOP, OLTC, SVC, and CBs at the t-th decision respectively; δ SOP and α OLTC and γ SVC and λ SVC are the unit action costs of SOP, OLTC, SVC, and CBs respectively; M SOP and M OLTC and M SVC and M CBs are the numbers of SOP, OLTC, SVC, and CBs in the distribution network respectively; H s,t is the adjustment action of the s-th SOP at the t-th decision; T g,t is the tap position of the g-th OLTC at the t-th decision; C q,t is the tap position of the q-th SVC at the t-th decision; R e,t is the tap position of the e-th CBs at the t-th decision; H s,t-1 is the adjustment action of the s-th SOP at the (t - 1)-th decision; T g,t-1 is the tap position of the g-th OLTC at the (t - 1)-th decision; C q,t-1 is the tap position of the q-th SVC at the (t - 1)-th decision; R e,t-1 is the tap position of the e-th CBs at the (t - 1)-th decision; s is the number of the SOP in a certain decision; g is the number of the OLTC in a certain decision; q is the number of the SVC in a certain decision; e is the number of the CBs in a certain decision.

6. The network-variable-load voltage coordinated control method based on multi-agent deep reinforcement learning according to claim 5, characterized in that: The constraint conditions include power flow constraints, voltage constraints, current constraints, distributed power source power constraints, SOP capacity constraints, distribution transformer power constraints, SVC and CBs capacity constraints, and action adjustment constraints of SOP, OLTC, SVC, and CBs; 1) Power flow constraints: Where: P i and Q i are the active power and reactive power of node i respectively; U i and U j are the voltages at the beginning and end of line ij; G ij and B ij are the conductance and susceptance of line ij respectively; θ ij is the phase difference of line ij; 2) Voltage and current constraints: U i,min ≤U i,△t ≤U i,max (9); I ij ≤I ij,max (10); Where: U i,△t is the voltage of node i at time △t; U i,min , U i,max are the minimum and maximum values of the voltage of node i respectively; I ij,max is the maximum value of the current of line ij; 3) Distributed power source power constraints: In formula (11): P DG,i , Q DG,i are the active power output and reactive power output of the DG connected to node i respectively; P DG,max , P DG,min are the upper limit and lower limit of the active power output of the DG respectively; Q DG,max , Q DG,min are the upper limit and lower limit of the reactive power output of the DG respectively; 4) SOP capacity constraints: Where: P SOP,s , Q SOP,s are the active power output and reactive power output of the sth SOP respectively; P SOP,min , P SOP,max are the minimum and maximum values of the active power output of the SOP respectively; Q SOP,min , Q SOP,max are the minimum and maximum values of the reactive power output of the SOP respectively; S SOP,max is the maximum capacity of the SOP; 5) Distribution transformer power constraints: In formula (14): is the load rate of the g-th OLTC; is the maximum load rate of the OLTC; 6) SVC and CBs capacity constraints: where: is the reactive power output of the qth SVC; is the reactive power output of the eth CBs; are the minimum and maximum values of the reactive power output of the SVC, respectively; are the minimum and maximum values of the reactive power output of the CBs, respectively; 7) Action adjustment constraints of SOP, OLTC, SVC, and CBs: Where: Ω SOP,s,max , Ω OLTC,g,max , Ω SVC,q,max , Ω CBs,e,max are respectively the upper limits of the number of actions of the sth SOP, the gth OLTC, the qth SVC, and the eth CBs within a day; H s,t+1 is the adjustment action of the sth SOP at the (t + 1)th decision; T g,t+1 is the tap position of the gth OLTC at the (t + 1)th decision; C q,t+1 is the tap position of the qth SVC at the (t + 1)th decision; R e,t+1 is the tap position of the eth CBs at the (t + 1)th decision.

7. The network-variable-load voltage coordinated control method based on multi-agent deep reinforcement learning according to claim 1, characterized in that: In step 3, the voltage optimization control problem of the network-transformer-load three-layer coordinated distribution network is modeled as a Markov decision process, which is represented as an eight-tuple <η, S, A, P, Ω, O, R, γ>. Among them, η represents the number of agents; S represents the global state information; A represents the set of action spaces; P: S×A×S→[0,1] represents the state transition probability function; Ω represents the observation function of the initial state; O represents the set of observation space information; R represents the reward function; γ represents the discount factor. 1) Agent state and observation space: In the multi-agent reinforcement learning framework, the state space S is the global state of the system at each moment. To meet the requirements of distributed decision-making, each agent only observes the local state and neighbor information. The state set includes all the state information of the network-transformer-load three layers, including the voltage amplitude, active and reactive power of each node, the actions of each device, and the provided power. The observation space O of the agent = [X s , Y s , Z s , where X s , Y s , Z s are the status information of the grid side, transformer side, and load side of the distribution network respectively; the observation space of agent m is O m , and O m ∈ O; The observation space is as follows: In formula (21): U i,t , P i,t , Q i,t are respectively the voltage, active power, and reactive power of node i when the agent makes the t-th decision; Q DG,t , P DG,t are respectively the active and reactive powers output by the DG; P SOP,s,t , Q SOP,s,t , Ω SOP,s,t are respectively the active power, reactive power, and cumulative number of actions output by the s-th SOP; Y s = {U T,g,t , Q SVC,q,t , Ω OLTC,g,t , Ω SVC,q,t}} (22); In formula (22): U T,g,t , Q SVC,q,t , Ω OLTC,g,t , Ω SVC,q,t are respectively the output voltage of the g-th transformer, the reactive power provided by the q-th SVC, the cumulative action times of the g-th OLTC and the q-th SVC when the agent makes the t-th decision; Z s = {U Load,t , Q CBs,e,t , P Load,t , Q Load,t , Ω CBs,e,t}} (23); In formula (23): U Load,t , Q CBs,e,t , P Load,t , Q Load,t , Ω CBs,e,t are respectively the load terminal voltage at the t-th decision of the agent, the reactive power provided by the e-th CBs, the active power and reactive power at the load terminal, and the cumulative action times of the e-th CBs; 2) Agent action space: The action space A of the agent = [X A , Y A , Z A , and the action space a of the agent m m , and a m ∈A; where X A , Y A , Z A are the action space information of the grid side, transformer side, and load side of the distribution network respectively, and are specifically as follows: X A = [ΔQ DG,t , ΔP DG,t , ΔP SOP,s,t , ΔQ SOP,s,t , H s,t (24); In formula (24): △Q DG,t , △P DG,t are the active and reactive power outputs regulated by the DG during the t-th decision-making of the agent respectively; △P SOP,s,t , △Q SOP,s,t , H s,t are the active power output, reactive power output regulated by the s-th SOP and the action of the SOP during the t-th decision-making of the agent respectively; Y A = [△Q SVC,q,t , T g,t , C q,t (25); In Equation (25): ΔQ SVC,q,t , T g,t , C q,t are the reactive power output regulated by the q-th SVC during the t-th decision-making of the agent, the tap position of the g-th transformer OLTC, and the tap position of the q-th SVC; Z A = [△Q CBs,e,t , △P Load,t , △Q Load,t , R e,t (26); In formula (26): △Q CBs,e,t and R e,t are the reactive power output and tap position adjusted by the e-th CBs during the t-th decision-making of the agent, respectively; △P Load,t and △Q Load,t are the active power output and reactive power output adjusted at the load end during the t-th decision-making of the agent, respectively. 3) Agent reward function: The reward function R of the agent during the training process is as follows: In formula (27): A j is the operation cost of the j-th device; Φ is the penalty coefficient, and σ() is the judgment function; Ω SOP,s,t , Ω OLTC,g,t , Ω SVC,q,t , Ω CBs,e,t are the action times of the s-th SOP, the g-th OLTC, the q-th SVC, and the e-th CBs of the agent at the t-th decision-making, respectively; Ω SOP,s,max , Ω OLTC,m,max , Ω SVC,q,max , Ω CBs,e,max are the maximum action times of the s-th SOP, the g-th OLTC, the q-th SVC, and the e-th CBs, respectively.

8. The network-variable-load voltage coordinated control method based on multi-agent deep reinforcement learning according to claim 1, characterized in that: In step 4, each agent generates an action using the Actor network π θ (a|o) according to its own local observation information; agent m generates an action a m based on the local observation o θ,m and the policy network π m through decision-making, as shown in formula (28): π θ (a m |o m )=Q m [f θm (o m )](28); In formula (28): π θ (a m |o m ) is the Actor network of agent m; Q m [f θm (o m )] is the value sampled from the probability distribution of action a m of agent m; Q m is the response function of the agent to the current environment, used to sample action a m from the probability distribution; f θm (o m ) is the action probability distribution obtained by the non - linear mapping of the local state through the Actor network; f θm is the non - linear mapping of the Actor network of agent m; θ is the parameter of the Actor network; Each agent m can only perceive its own local state o m , and perform feature extraction and sharing through the implicit communication module, as shown in formula (29): z m = E φ (o m )(29); In Equation (29): z m is the feature embedding generated by the local state extraction network E φ ; φ are the Critic network parameters; The local feature embeddings of all agents are aggregated through the implicit communication module to form global shared features to describe the global characteristics of the entire system, as shown in formula (30): In formula (30): z all is the globally shared feature; n is all agents; Each agent m combines its own local state o m with the globally shared features and assigns them to each agent through the set global reward function R, as shown in Equation (31), to guide policy optimization; each agent performs reward allocation and adjusts its policy through the shared global reward signal to generate the final action decision, ensuring the optimality of the global performance, as shown in Equation (32): R m = g(R, o m , a m ) (31); π θ (a m |o m ,z all ) = Q m [f θm (o m ,z all )] (32); where: g() is the reward decomposition function; R m is the global reward function of agent m; π θ (a m |o m ,z all ) is the action decision of agent m; f θm (o m ,z all ) is the non-linear mapping of the optimized local state of agent m through the Actor network to generate the corresponding action probability distribution; Q m [f θm (o m ,z all )] is the value sampled from the probability distribution of the optimized action a m of agent m.

9. The network-variable-load voltage coordinated control method based on multi-agent deep reinforcement learning according to claim 1, characterized in that: In step 5, the multi-agent proximal policy optimization algorithm MAPPO adopts the Actor-Critic architecture, and the PPO algorithm is used to solve the optimization of the Actor network and the Critic network in the implicit communication multi-agent network-transformer-load voltage optimization control problem. The loss function \(L\) of the Actor network CLIP (\(\theta\)) is the clipping policy objective function as follows: R t = r t + γr t+1 + γ 2 r t+2 +…γ n r t+n (36) In the formula: denotes taking the expectation of the data distribution at time step t; clip(r t (θ), 1 - ε, 1 + ε) restricts the magnitude of the policy update; r t (θ) measures the change in the action selection probability between the current policy and the old policy, as shown in Equation (34); measures the advantage function of the current action, as shown in Equation (35); R t is the total reward accumulated starting from time step t, i.e., the global reward R all , as shown in Equation (36); ∈ is the set clip parameter; V φ (o t ) is the value estimate of the current state o t ; π θ (a t |o t , z all ) is the probability of the current policy selecting action a t ; π θ′ (a t |o t , z all ) is the probability of the new policy selecting action a t ; r t is the reward at time step t; r t+1 is the reward at time step t + 1; r t+2 is the reward at time step t + 2; r t+n is the reward at time step t + n; γ is the discount factor; The goal of the Critic network is to minimize the current state value, and its value function error is: In Equation (37), L VF (φ) is the value function error; To encourage the exploration of agents, the PPO algorithm adds a policy entropy regularization term to the loss function: In Equation (38), H(π θ ) is the policy entropy, representing the uncertainty of action selection; is the expected value, representing the expected calculation under all possible states and actions; The total loss function L(θ, φ) of the PPO algorithm combines the clipped policy objective, the value function error, and the entropy regularization term, specifically: L(θ, φ) = L CLIP (θ) - c1L VF (φ) + c2H(π θ )(39); In Equation (39), L CLIP (θ) is the optimized objective of the cropped policy; c1 and c2 are the weight coefficients of the value function error and the entropy regularization term, which are used to adjust the weights of different loss terms; The PPO algorithm updates the parameters of the Actor network and the Critic network respectively through the gradient descent method; The update of the Actor network parameters is as follows: In Equation (40), α is the learning rate, which is used to control the update step size; is the update direction of the policy network parameters; ← represents parameter update; θ is the parameter of the Actor network; The update of the Critic network parameters is as follows: In formula (41): β is the learning rate, which is used to control the update step of the Critic network; φ is the parameter of the Critic network.

10. The network-variable-load voltage coordinated control method based on multi-agent deep reinforcement learning according to claim 1, characterized in that: In step 5, the training process of the multi-agent proximal policy optimization algorithm MAPPO includes the following steps: First, the real-time state of the distribution network environment is updated, and the action information of devices such as SOP, DG, OLTC, SVC, CBs, and the load side is collected, including Ω SOP,s,t 、Ω OLTC,g,t 、Ω SVC,q,t 、Ω CBs,e,t 、△Q DG,t 、△P DG,t 、△P SOP,s,t 、△Q SOP,s,t 、H s,t 、△Q SVC,q,t 、T g,t 、C q,t 、△Q CBs,e,t 、R e,t 、△P Load,t 、△Q Load,t ; and the state parameters such as the voltage amplitude, active load, and reactive load of each key node in the distribution network are dynamically updated, including U i,t 、P i,t 、Q i,t 、Q DG,t 、P DG,t 、P SOP,s,t 、Q SOP,s,t 、U T,g,t 、Q SVC,q,t 、U Load,t 、Q CBs,e,t 、P Load,t 、Q Load,t ; Secondly, the agent samples the collected operation state data and stores it in the experience replay buffer, and conducts centralized training by randomly extracting the data in the buffer; Then, it is trained and solved through the value decomposition MAPPO of implicit communication. Each agent makes a decision based on its own local observation information. Through the implicit communication mechanism, the agents indirectly interact in the shared environment and obtain the decision compensation feedback by the environment; Subsequently, the system comprehensively evaluates the behavior of the agents, generates a global reward according to the overall goal, and reasonably distributes it to each agent to prompt the individual to optimize its own strategy. In this stage, the agents continuously learn and optimize through the Actor-Critic architecture of the PPO algorithm, and finally form a trained action strategy. Finally, each agent cancels implicit communication. Each agent uses a distributed method to perform relevant action selection on the trained policy and transmits it to the security layer for correction to ensure the safety and reliability of its actions. The corrected set of actions is then sent to each device in the three layers for collaborative control to achieve optimal voltage control of the distribution network.