Three-phase tidal current power optimization scheduling method and system suitable for soft switch

By constructing a security reinforcement learning algorithm based on the Lagrangian function method, combining the three-phase current constraints of the distribution network and the soft switch operation constraints, the three-phase power timing model of soft switches is optimized, and the unsafe problem of SOP optimization scheduling in renewable energy scenarios is solved, and efficient and reliable scheduling decisions are achieved.

CN120545979APending Publication Date: 2025-08-26STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510650872.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the uncertainty and volatility scenarios of renewable energy, the existing SOP optimization scheduling methods have problems with unsafe strategies. Traditional methods are difficult to find feasible solutions within a reasonable time and have low computational efficiency. Deep reinforcement learning algorithms have shortcomings in balancing the security and exploration of strategy.

Method used

A security reinforcement learning algorithm based on Lagrangian function method is adopted, combined with the three-phase current constraints of the distribution network and the soft switch operation constraints, a three-phase power timing model of soft switch is constructed, and optimized and updated through the Markov decision-making process, and cost constraints are introduced to ensure the safety and effectiveness of the strategy.

Benefits of technology

Adaptive adjustment and online real-time decision-making in renewable energy scenarios are realized, computing efficiency is improved, the accuracy and reliability of the scheduling scheme are ensured, and the emergence of unsafe strategies is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120545979A_ABST
    Figure CN120545979A_ABST
Patent Text Reader

Abstract

The invention discloses a three-phase power flow power optimization scheduling method and system suitable for a soft switch. The method comprises the following steps: acquiring a mathematical model based on a three-phase power control characteristic of the soft switch; constructing a soft switch three-phase power time sequence model for the mathematical model based on the soft switch three-phase power control characteristics in combination with a power distribution network three-phase power flow constraint, a power distribution network safe operation constraint and a soft switch operation constraint, and converting the soft switch three-phase power time sequence model into a Markov decision process based on cost constraint; obtaining a soft switching three-phase power time sequence optimization model; a safety reinforcement learning algorithm based on a Lagrange function method is used to optimize and update the soft switching three-phase power time sequence optimization model; and the optimized and updated soft switching three-phase power time sequence optimization model is applied to schedule the power of the power distribution network, so that the intelligent agent is effectively prevented from generating an unsafe scheduling strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of soft-switch power scheduling, and in particular to a method and system for optimizing scheduling of soft-switch three-phase power flow. Background Art

[0002] With the increasing popularity of intermittent distributed generation (DG), the prospects for energy decarbonization have improved. However, due to the uncertainty and volatility of renewable energy, the operation of distribution networks also faces major challenges, including more frequent voltage limit violations and increased network losses. Soft Open Point (SOP), as a flexible regulation device at the "grid" level in the distribution system, can change the traditional operation mode of the system, achieve continuous regulation of power flow between feeders, effectively deal with a series of new problems brought about by the access of distributed generation, and ensure safe, reliable and cost-effective operation. Therefore, it is very necessary to use SOP to optimize the operation of the distribution network.

[0003] Distribution networks incorporate a large number of distributed generation (DGs), new loads, and flexible resources. Traditional mathematical optimization methods must simultaneously consider modeling multiple complex factors, such as source and load uncertainty, three-phase power flow equations, and the regulation characteristics of flexible resources, and then solve them through convex optimization.

[0004] Although existing research has made some progress in optimizing scheduling strategies using operations research or heuristic algorithms, the limitations of such methods in practical engineering applications have become increasingly apparent as system scale continues to expand. In particular, the complex three-phase power flow equations coupled with other variables exponentially increase the dimensionality of the optimization problem, leading to the algorithm facing a "curse of dimensionality" and making it difficult to find a feasible solution within a reasonable timeframe. Furthermore, the linearization techniques introduced to ensure model solvability can lead to the accumulation of theoretical errors, causing deviations in the calculation of key indicators such as system losses and voltage deviation rates, thus compromising the engineering applicability of the scheduling strategy.

[0005] As the scale of distribution networks continues to expand and the complexity of optimization models continues to increase, traditional model-driven solution methods will face challenges such as a sharp increase in optimization difficulty, low solution efficiency, and difficulty in achieving real-time decision-making on scheduling plans.

[0006] Existing deep reinforcement learning (DRL) algorithms for SOP optimization scheduling have effectively overcome the limitations of traditional mathematical optimization methods, but they lack the ability to balance the safety and explorative nature of DRL algorithm strategies. Constraints in the model (such as voltage safety constraints) are typically added to the reward function in the form of penalty functions. While penalty-based modeling effectively implements DRL for constrained optimization problems, it relies heavily on the proper setting of penalty coefficients. Excessively large penalty coefficients can limit the agent's exploration capabilities, causing the strategy to fall into a local optimum. On the other hand, excessively small penalty coefficients can cause the agent to ignore safety constraints in order to obtain greater rewards, leading to unsafe scheduling strategies in certain scenarios.

[0007] Patent application publication number CN119444502A addresses the difficulties and time-consuming nature of integrated energy system planning. This invention proposes a method and system for optimizing integrated energy system planning based on deep reinforcement learning. The method establishes a mathematical model of the integrated energy system, determines the objective function and constraints of the integrated energy system, and formulates an initial dynamic programming optimization problem. However, the mathematical model is based on a single-phase power flow.

[0008] Patent application publication number CN119813402A discloses a method and system for improving the transmission capacity of AC / DC hybrid transmission corridors. Building on traditional maximum entropy reinforcement learning, this patent introduces safety constraints through safety reinforcement learning and transforms the constrained optimization problem into a Lagrangian form. This patent directly uses the Lagrangian function to transform the constraint dual into the objective function. The Q network training process requires simultaneous updates to both the entropy and the dual variable, meaning that entropy training is required. The additional fluctuations introduced by the policy entropy affect the stability of the optimization. Summary of the Invention

[0009] The technical problem to be solved by the present invention is to solve the problem of unsafe strategies in data-driven SOP optimization under the uncertainty and volatility of renewable energy.

[0010] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0011] A method for optimizing power scheduling of three-phase power flows in soft switching, comprising:

[0012] Obtain a mathematical model based on the soft-switching three-phase power control characteristics;

[0013] Based on the mathematical model of the soft-switching three-phase power control characteristics, combined with the three-phase power flow constraints of the distribution network, the distribution network safety operation constraints, and the soft-switching operation constraints, a soft-switching three-phase power timing model is constructed. This soft-switching three-phase power timing model is transformed into a Markov decision process based on cost constraints, and a soft-switching three-phase power timing optimization model is obtained.

[0014] The soft-switching three-phase power timing optimization model is optimized and updated using a secure reinforcement learning algorithm based on the Lagrangian function method. The optimized and updated soft-switching three-phase power timing optimization model is used to dispatch the power of the distribution network.

[0015] In this embodiment, the expression of the soft switching three-phase power timing model is:

[0016]

[0017] Where M SOP PDN is a mathematical model based on the soft switching three-phase power control characteristics. st are the three-phase power flow constraints of the distribution network, the safe operation constraints of the distribution network, and the soft switch operation constraints, st is the constraint, is the objective function, C OBJ For the goal, C LOSS is the total system loss, C UNB is the system three-phase voltage unbalance, κ1 and κ2 are the weight coefficients of the total system loss and the three-phase voltage unbalance in the objective function respectively.

[0018] In this embodiment, the three-phase power flow constraints of the distribution network are:

[0019]

[0020]

[0021]

[0022]

[0023] In the formula, ij represents that the starting point of the branch is node i and the end point is node j; ki represents that the starting point of the branch is node k and the end point is node i; S ij,t and S ki,t are the three-phase power matrices transmitted by branch ij and branch ki at time t; Z ij and Z ki are the three-phase impedance matrices of branch ij and branch ki respectively; and are the three-phase voltage vectors of node i and node j at time t; L ij,t and L ki,t are the matrices obtained by multiplying the three-phase current vectors of branch ij and branch ki at time t by their own conjugate transpose; is the column vector I ij,t The conjugate transpose of H is conjugate; S net,i,t is the three-phase injection power matrix of node i at time t; Snet,i,t,a 、S net,i,t,b 、S net,i,t,b are the injected power phasors of phase a, phase b and phase c of node i at time t respectively; is the node i at time t Phase injection power phasor; and They are respectively the node i at time t Phase load active and reactive demands; DG is at time t is the active power transmitted by the phase, DG is the distributed generation, j is the imaginary unit, and T is the transpose.

[0024] In this embodiment, the distribution network safe operation constraints are:

[0025]

[0026] Where V min and V max are the minimum and maximum values ​​of the node voltage, respectively. is the node i at time t Phase complex voltage.

[0027] In this embodiment, a mathematical model based on the soft switching three-phase power control characteristics is used as a condition for the soft switching operation constraint.

[0028] In this embodiment, the soft-switching three-phase power timing model is transformed into a cost-constrained Markov decision process, including:

[0029] The three-phase active power, three-phase reactive power and DG output of all nodes in the distribution network are defined as the system state s t ;

[0030] Define action a by inputting / outputting the three-phase active power and three-phase reactive power of the soft switch port. t ;

[0031] Reducing system losses and alleviating the three-phase imbalance of the system is defined as the reward R t ;

[0032] The violation of the safe operation constraint of the distribution network is defined as the cost function

[0033] In a given state s t and action a t Under this condition, the next state of the system will be transferred based on the probability transfer function Pr; after executing the action, it reaches the next state, and continuously iterates and updates the strategy π through the reward corresponding to the action.

[0034] In this embodiment, the expression of the soft switching three-phase power timing optimization model is:

[0035] π=argmax(J(π));

[0036]

[0037]

[0038]

[0039] Where J(π) is the cost corresponding to each reward under each strategy, J C (π) is the cost of each action under each strategy, and the corresponding cost is, is the upper bound of the corresponding cost under each strategy, τ is the state sequence, E is the expectation, and γ is the discount factor.

[0040] In this embodiment, the soft switching three-phase power timing optimization model is optimized and updated, including:

[0041] The Lagrangian function method's secure reinforcement learning algorithm includes a state-action value network, a state-action cost network, and a policy network;

[0042] Among them, the expression of the state-action-value network is:

[0043]

[0044] The expression of the state-action cost network is:

[0045]

[0046] The expression of the policy network is:

[0047]

[0048] Where, J Q (θ) is the loss of the state-action-value network, D is the experience replay pool, Q after parameter θ update θ network, θ is the updated Q network parameter, θ is the parameter of the Q network, is the loss of the state-action cost network, For parameters Updated Network, s t 、a t 、R t ,γ, E represents the state, action, reward, discount factor, cost function and expectation in the Markov decision process; is the loss of the policy network, x is the original variable, α and λ are the Lagrange multipliers of the temperature coefficient and entropy respectively, and are divided into dual variables, H is the entropy term, π φ (a t |s t ) is the action a under the parameter φ t , state s t For the strategy, is the upper bound of the corresponding cost under each strategy.

[0049] In this embodiment, in the policy network, the updating process of the original variables and the dual variables is:

[0050]

[0051]

[0052]

[0053] Where, δ x , δ α , δ λ are the update steps of different parameters, x k , α k ,λ k are the original variables and dual variables updated for the kth time, are the updated gradients of different parameters, respectively. + is the projection of a non-negative real number.

[0054] The present invention further provides a system for optimizing and dispatching three-phase power flow for soft switching, which applies the above-mentioned method for optimizing and dispatching three-phase power flow for soft switching, including:

[0055] Basic model module, used to obtain the mathematical model based on the soft switching three-phase power control characteristics;

[0056] The Markov decision module is used to construct a soft-switching three-phase power timing model based on the mathematical model of the soft-switching three-phase power control characteristics, combined with the distribution network three-phase power flow constraints, distribution network safe operation constraints, and soft-switching operation constraints. The soft-switching three-phase power timing model is converted into a Markov decision process based on cost constraints to obtain the soft-switching three-phase power timing optimization model;

[0057] The update optimization module is used to optimize and update the soft-switching three-phase power timing optimization model using the Lagrangian function method-based security reinforcement learning algorithm; the optimized and updated soft-switching three-phase power timing optimization model is used to dispatch the power of the distribution network.

[0058] Compared with the prior art, the present invention has the following beneficial effects:

[0059] This invention combines the advantages of data-driven approaches to achieve adaptive adjustment of source-load output uncertainty and online real-time decision-making. It also improves the offline training phase of the reinforcement learning algorithm, effectively preventing the intelligent agent from generating unsafe scheduling strategies. This improves computational efficiency while ensuring the accuracy and reliability of the scheduling solution.

[0060] In the three-phase power flow constraint of the distribution network, the original network is expanded into a virtual equivalent network, the dimensions of the voltage and current matrices are expanded, and virtual connections between and within nodes that did not originally exist are introduced, successfully expressing the implicit coupling relationship between the three phases of the distribution network explicitly.

[0061] In the soft-switching three-phase power sequence model, the three-phase active power and three-phase reactive power of each port of the SOP are used as decision variables, and the weighted minimization of the total system loss and the three-phase voltage imbalance is taken as the optimization objective.

[0062] The cost-constrained Markov decision process incorporates six core elements, compared to the five used in conventional Markov decision processes. The proposed cost-constrained Markov decision process incorporates constraints, requiring the agent to ensure that its trained action strategy does not violate established constraints while completing its mission objectives. Therefore, the agent's strategy training is no longer solely based on maximizing reward, but instead requires fully considering the constraints and seeking the optimal strategy while satisfying them.

[0063] In a cost-constrained Markov decision process, a cost function is associated with the agent's system state and the actions it performs. Only when the cumulative cost is below a set threshold can the agent's actions be guaranteed to be safe, meaning that the distribution network will not exceed voltage limits. If the agent performs an unsafe action in a given state, a cost penalty is imposed on its behavior. Furthermore, by considering the cumulative cost to define the safety of the entire trajectory, the long-term safety of the agent can be effectively assessed, rather than focusing solely on a single state or action.

[0064] Conventional SAC algorithms contain two neural networks and are only applicable to conventional Markov decision processes (MDPs). The present invention uses a secure reinforcement learning algorithm based on the Lagrangian function method to optimize and update the soft-switched three-phase power timing optimization model, which is a constrained optimization problem, namely, CMDP. For the CMDP problem, a common approach is to convert the constraints into penalty terms and add them to the reward function, thereby converting the problem into an MDP and then using the SAC algorithm to solve it. However, this conversion method makes it difficult to accurately measure the impact of constraint violations on scheduling results. At the same time, the design of the penalty term will directly affect the learning process. Inappropriate settings may lead to overly conservative strategies or even infeasible strategies. To overcome the limitations of the above methods, the present invention expands the SAC algorithm, introduces Lagrangian multipliers, and proposes a SAC algorithm based on the Lagrangian function method, namely, the core idea of ​​the AL-SAC algorithm is to introduce constraints into the objective function, dualize the constraints through Lagrangian multipliers, and embed them into the objective function.

[0065] In the AL-SAC algorithm, a Q-network is added to the original SAC algorithm to evaluate the cost function. Furthermore, the entropy regularization term is removed during the Q-network training process. This ensures that the state-action function accurately estimates the reward function, avoiding the additional fluctuations introduced by policy entropy that can affect optimization stability. Furthermore, the state-cost function is designed to strictly reflect the cumulative cost of the constraints. The inclusion of an entropy term can lead to biased constraint cost estimates, potentially interfering with the dynamic update of the Lagrange multiplier.

[0066] A Lagrangian-based safety reinforcement learning algorithm extends the traditional SAC algorithm with a Q-network. Multiple Q-networks are trained simultaneously to approximate the state-action benefits and costs of CMDP. The policy network training process is centered around the Lagrangian function method, aiming to reduce losses and alleviate three-phase imbalance while ensuring the system operates within a safe range. By iteratively updating the Lagrangian multiplier, the algorithm demonstrates improved convergence and stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 This is a flow chart of a method for optimizing power scheduling of three-phase power flow in soft switching according to an embodiment of the present invention.

[0068] Figure 2 Schematic diagram of a typical SOP access mode according to an embodiment of the present invention.

[0069] Figure 3 Schematic diagram of the SOP operation boundary of an embodiment of the present invention.

[0070] Figure 4Schematic diagram of the original network and the expanded equivalent network of nodes in three phases according to an embodiment of the present invention.

[0071] Figure 5 Schematic diagram of a Markov decision process based on cost constraints according to an embodiment of the present invention.

[0072] Figure 6 This is a schematic diagram of the IEEE 123 node system structure according to an embodiment of the present invention.

[0073] Figure 7 The figure is a schematic diagram of the iterative convergence process of each algorithm in the IEEE 123 node test case according to an embodiment of the present invention.

[0074] Figure 8 This figure is a schematic diagram of the number of unsafe scenarios that occur during the training process of each algorithm under the IEEE 123 node test case according to an embodiment of the present invention.

[0075] FIG9( a ) is a load characteristic curve diagram of an embodiment of the present invention.

[0076] Figure 9(b) to Figure 9(d) This is a graph showing the output characteristic of photovoltaic PV1-PV10 according to an embodiment of the present invention.

[0077] Figure 10(a) 、 10(b) This is a schematic diagram of active and reactive transmission of the SOP port a phase according to an embodiment of the present invention.

[0078] Figure 10(c) 、 10(d) Schematic diagram of active and reactive transmission of SOP port b phase in an embodiment of the present invention.

[0079] Figure 10(e) 、 10(f) This is a schematic diagram of active and reactive transmission of the SOP port phase C according to an embodiment of the present invention.

[0080] Figure 11 Schematic diagram of the three-phase unbalanced changes in the voltage at node 85 in various scenarios of an embodiment of the present invention.

[0081] Figure 12 Schematic diagram of the current amplitude of branches 75 to 82 in various scenarios of an embodiment of the present invention. DETAILED DESCRIPTION

[0082] To facilitate those skilled in the art to understand the technical solution of the present invention, the technical solution of the present invention is further described with reference to the accompanying drawings.

[0083] The terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, "plurality" means two or more, unless otherwise specifically defined.

[0084] It should be noted in advance that:

[0085] Soft switch: refers to a new type of intelligent power electronic device installed at the traditional tie switch, which can transform the distribution network from traditional closed-loop design open-loop operation to flexible closed-loop operation.

[0086] Distributed power source: refers to a power source that is connected to a power grid with a voltage level of 35kV or below, located near the user, and mainly consumed locally at a voltage level of 35kV or below.

[0087] Example 1

[0088] See also Figure 1 As shown, the present invention discloses a method for optimizing the power scheduling of a three-phase power flow applicable to soft switching, comprising:

[0089] S10, obtaining a mathematical model based on soft switching three-phase power control characteristics.

[0090] Soft switching three-phase power control characteristics:

[0091] See also Figure 2 、 3 As shown in FIG, in this embodiment, the SOP based on power electronic devices has excellent voltage control performance and three-phase power flow regulation capability. In the actual distribution network application scenario, each port of the SOP is connected to the three-phase AC feeder of the distribution network. This connection method enables different AC feeders to form an energy path through the SOP, thereby achieving continuous regulation of the three-phase active power and three-phase reactive power between feeders. Its typical access mode and three-phase model are as follows: Figure 2 shown.

[0092] Figure 2 The demonstration demonstrates a split-phase connection of two feeders via a SOP. Two voltage source converters (VSCs) are located on either side of the SOP, with energy exchanged between the two ports via a single DC bus. The VSCs at both ends utilize a modular multilevel architecture, with each phase consisting of a reactor and multiple submodules connected in series. The VSCs can inject or absorb three-phase active and reactive power into the interconnected feeder.

[0093] When the distribution network is in a three-phase unbalanced operation state, SOP can adjust the three-phase power flow through the VSC at both ends, thereby improving the overall operation efficiency of the system and alleviating the problems of voltage deviation and loss increase caused by three-phase imbalance. dc Q-PQ control mode, that is, a VSC i The three-phase power flow is regulated by controlling the active power P and reactive power Q. Another VSC j The control mode of the VSC port is V dc Q maintains the voltage of the DC bus. Where i and j are different feeders. In this control mode, the controllable variables of SOP include the three-phase active power flowing through SOP and the three-phase reactive power emitted by the two ports of SOP. Its operating boundaries are as follows: Figure 3 shown.

[0094] Mathematical model based on soft switching three-phase power control characteristics:

[0095] In the distribution network optimization operation, for the VSCs flowing through both ends of the SOP i and VSC j The three-phase active power must maintain input and output balance. Furthermore, as a power electronic device, the VSC inevitably incurs energy losses during power transmission. Therefore, the active power balance equations for the SOP are shown in Equations (1-1) through (1-3).

[0096]

[0097]

[0098]

[0099] Where, and They are VSC at both ends of SOP at time t i and VSC j exist Active power transmitted between phases; and They are VSC at both ends of SOP at time t i and VSC j exist Reactive power transmitted between phases; and They are VSC at both ends of SOP at time t i and VSC j exist Phase active power loss; A SOP is the active power loss coefficient of VSC.

[0100] To ensure the safe operation of SOP, the reactive power limit of SOP is shown in Equations (1-4) to (1-5).

[0101]

[0102]

[0103] Where, and VSC at both ends of SOP i and VSC j exist Minimum reactive power of phase transmission; and VSC at both ends of SOP i and VSC j exist The maximum reactive power transmitted by each phase.

[0104] When performing three-phase power regulation, the VSCs at both ends of the SOP must meet active power balance constraints. Reactive power can be compensated or absorbed based on the actual control requirements of the interconnected feeders. The reactive power transmitted by each phase must be regulated within the SOP capacity limits. The SOP capacity constraints are shown in Equations (1-6) to (1-7).

[0105]

[0106]

[0107] Where S SOP is the port capacity of SOP.

[0108] That is, equations (1) to (7) are mathematical models based on the soft switching three-phase power control characteristics.

[0109] S20, for the mathematical model based on the soft-switching three-phase power control characteristics, combined with the three-phase power flow constraints of the distribution network, the safe operation constraints of the distribution network and the soft-switching operation constraints, a soft-switching three-phase power timing model is constructed, and the soft-switching three-phase power timing model is converted into a Markov decision process based on cost constraints to obtain a soft-switching three-phase power timing optimization model.

[0110] Objective function of the soft-switching three-phase power timing model:

[0111] In this embodiment, the soft-switching three-phase power timing model is based on the previous day's source-load data, using the three-phase active power and three-phase reactive power at each SOP port as decision variables, with the optimization goal of minimizing system losses and three-phase imbalance. In distribution network operation, three-phase imbalance specifically includes two key indicators: three-phase voltage imbalance and three-phase current imbalance.

[0112] In this embodiment, the present invention is directed to a medium-voltage distribution network, focusing on the optimization of the system's operation under unbalanced three-phase connection of sources and loads. In this scenario, the three-phase imbalance of the node voltage determines whether the system is in a safe and economical operating range. Therefore, the present invention primarily focuses on the system's three-phase voltage imbalance, with the weighted minimization of the total system loss and the three-phase voltage imbalance as the optimization objective. The specific objective function, namely the soft-switching three-phase power timing model, is shown in Equations (2-1) to (2-3).

[0113] C OBJ =κ1C LOSS +κ2C UNB , (2-1);

[0114]

[0115]

[0116] Where C OBJ Target; C LOSS is the total system loss, including distribution network branch loss and SOP active power loss; C UNB The system three-phase voltage imbalance is represented by the difference between the average voltage of each phase and the complex voltage; is the branch ij at time t Phase current; For branch ij Phase resistance; is the node i at time t Phase complex voltage; V i,t,a 、V i,t,b 、V i,t,c are the voltages of phase a, phase b, and phase c at node i at time t; T is the total scheduling period; Ω bus is the set of distribution network nodes; Ω branch is the distribution network branch set; Ω SOP is the SOP set; Φ is the three-phase set of phase a, phase b, and phase c; κ1 and κ2 are the weight coefficients of the total system loss and the weight coefficient of the system three-phase voltage imbalance in the objective function, respectively, which can be calculated by the hierarchical analysis method.

[0117] In this embodiment, in addition to the above-mentioned objective function, the soft-switching three-phase power timing model constructed by the present invention must also fully consider various constraints that the distribution network must meet during actual operation, including the three-phase power flow constraints of the distribution network, the safe operation constraints of the distribution network, and the soft-switching operation constraints. The specific contents are described below.

[0118] Three-phase power flow constraints of distribution network:

[0119] In this embodiment, the three-phase power flow constraint of the distribution network ensures that the power flowing into / out of the node by the branch is always balanced with the power injected into the node, and at the same time characterizes the relationship between the branch voltage drop and the branch power flow. The power flow equation usually takes the node voltage method as a reference and models the three-phase power flow in the form of node injection, which includes the node voltage and the injected power of the node. Specifically, in a radial distribution network, the three-phase power flow equation established in the form of branch power is more intuitive and concise, and the established model involves multiple variables such as branch current, branch power, and node voltage. In addition, in order to consider the neutral line loss, it is necessary to combine the Kron (Kronecker Product) method to reduce the three-phase impedance matrix of the branch, and reduce the 4×4 matrix dimension to 3×3 to facilitate power flow calculation. In summary, the specific power flow equations are shown in Equations (2-4) to (2-8).

[0120]

[0121]

[0122]

[0123]

[0124]

[0125] In the formula, ij represents that the starting point of the branch is node i and the end point is node j; ki represents that the starting point of the branch is node k and the end point is node i; S ij,t and S ki,t are the three-phase power matrices transmitted by branch ij and branch ki at time t; Z ij and Z ki are the three-phase impedance matrices of branch ij and branch ki respectively; and are the three-phase voltage vectors of node i and node j at time t; L ij,t and L ki,t are the matrices obtained by multiplying the three-phase current vectors of branch ij and branch ki at time t by their own conjugate transpose; is the column vector I ij,t The conjugate transpose of H is conjugate; S net,i,t is the three-phase injection power matrix of node i at time t; S net,i,t,a 、S net,i,t,b 、S net,i,t,b are the injected power phasors of phase a, phase b and phase c of node i at time t respectively; is the node i at time t Phase injection power phasor; and They are respectively the node i at time t Phase load active and reactive demands; DG is at time t is the active power transmitted by the phase, DG is the distributed generation, j is the imaginary unit, and T is the transpose.

[0126] Considering the coupling relationship between the voltage, current and impedance of each phase in the three-phase unbalanced distribution network, the above three-phase power flow equation of the distribution network is expanded to reflect the implicit coupling relationship between each node and each phase. Define the matrix Perform conjugate multiplication on formula (2-6) to obtain formula (2-9).

[0127]

[0128] Where, For S ij,t The conjugate transpose of Z ij The conjugate transpose of , where W i,t 、W j,t It can be understood as the matrix obtained by multiplying the voltage vector by its own conjugate transpose.

[0129] At the same time, the relationship between node voltage, branch current and branch transmission power in the distribution network is further expressed as formula (2-10).

[0130]

[0131] According to the original tidal equation, that is, formula (2-4)-(2-8), the vector and I ij,t After right multiplying their respective conjugate transposed vectors, we get W i,t and L ij,t , which is equivalent to expanding the original network into a virtual equivalent network, such as Figure 4 As shown in Figure 2. Before expansion, the explicit structure of the i-th and j-th nodes indicates no direct connection between the three phases. At this point, the node voltage and current matrices contain only three elements, and the connections between the three phases are only reflected in the corresponding phases at different nodes. After expansion, the dimensions of the voltage and current matrices are expanded to 3×3, introducing virtual connections between and within nodes that did not originally exist, successfully explicitly expressing the implicit coupling relationship between the three phases.

[0132] Constraints on safe operation of distribution networks:

[0133] In this embodiment, when the distribution network operates normally, its node voltage deviation must meet the safe operation range and should not be too large or too small.

[0134] Therefore, the node voltage of the distribution network should meet the safe operation constraints, that is, the safe operation constraints of the distribution network, as shown in formula (2-11).

[0135]

[0136] Where V min and V max are the minimum and maximum values ​​of the node voltage, respectively.

[0137] Soft switching operation constraints:

[0138] In this embodiment, in addition to the above constraints, the operation constraints of SOP need to be considered when constructing the soft switch three-phase power timing model, as shown in Formulas (1-1) to (1-7). The constraints of SOP have been analyzed in detail in the previous section and will not be repeated here. In order to make the soft switch three-phase power timing model constructed by the present invention clearer, the above soft switch three-phase power timing model is summarized as follows, as shown in Formula (2-12).

[0139]

[0140] Where M SOP It is a mathematical model based on the soft switching three-phase power control characteristics, specifically formulas (1-1)-(1-7), PDN st are the three-phase power flow constraints of the distribution network, the safe operation constraints of the distribution network, and the soft switching operation constraints, specifically as shown in (2-4)-(2-11), st is the constraint, and P is the objective function.

[0141] The soft-switching three-phase power timing model is transformed into a Markov decision process based on cost constraints:

[0142] The SOP three-phase active / reactive power optimization scheduling problem constructed by the present invention aims to ensure that the distribution network operates within a safe range while maximizing the objective function. However, the traditional Markov Decision Process (MDP) is difficult to characterize the dynamic decision process with constraints. Therefore, the CMDP is constructed by expanding on the MDP.

[0143] Compared to traditional MDPs, the CMDP constructed in this step introduces constraints into the dynamic decision-making process. This means that while the agent completes the task objective, it must ensure that the trained action policy does not violate the established constraints. Therefore, the agent's policy training is no longer based solely on maximizing reward returns. Instead, it must fully consider the constraints and seek the optimal policy while satisfying them.

[0144] Based on the above analysis, CMDP is obtained by expanding MDP and modeling based on 6 core elements, as shown in formula (2-13).

[0145] <S,A,R,R C ,Pr,γ>,(2-13);

[0146] Where γ is the discount factor.

[0147] In this embodiment, the state S: The choice of state is a key factor affecting the algorithm's convergence performance. When constructing the state space, it is important to consider not only the technical nature of the state characteristics but also their necessity to reduce the negative impact of redundant information on algorithm performance. Therefore, considering the characteristics of the flexible interconnected distribution network, the three-phase active power, three-phase reactive power, and DG output of all nodes in the distribution network are defined as the system state, as shown in Equation (2-14).

[0148] s t =[P t ,Q t ,P t DG ], (2-14);

[0149] Where s t is the state at time t; P t is the three-phase active power vector of all nodes in the distribution network at time t; Q t is the three-phase reactive power vector of all nodes in the distribution network at time t; P t DG is the active power vector of the DG connected to all nodes in the distribution network at time t.

[0150] In this embodiment, Action A: The action space is the decision variable for the SOP three-phase active / reactive power optimization scheduling task, namely, the three-phase active power and three-phase reactive power input / output of the SOP port, as shown in Equation (2-15). Furthermore, the constructed action space must satisfy the SOP capacity constraint.

[0151]

[0152] Where a t is the action at time t; P t SOP Transmit three-phase active power vector for all SOPs in the distribution network at time t; It is the three-phase reactive power vector transmitted by all SOPs in the distribution network at time t.

[0153] In this embodiment, the reward function R: The goal of the optimized scheduling model is to reduce system losses and alleviate the system's three-phase imbalance. Therefore, when designing the agent's reward function, these two objectives must still be taken as goals. The reward function is designed to be the negative of the objective function, as shown in (2-16).

[0154] R t (s t ,a t )=-C OBJ , (2-16);

[0155] Where R t is the reward at time t.

[0156] In this embodiment, the cost function R C Based on the distribution network state space and actual operating conditions, the distribution network's operational constraints are safety constraints, meaning that node voltages must meet the system's permissible range. Therefore, the cost function is set to the amount of voltage constraint violation, as shown in Equation (2-17).

[0157]

[0158] Where, is the cost at time t, and is the cost of the action taken at time t.

[0159] In this embodiment, the probability transfer function Pr: given the state s t and action a t Under this condition, the next state of the system will be transferred based on the probability transfer function, as shown in formula (2-18).

[0160] Pr(s t ,s t+1 )=Pr(s t+1 |s t ,a t ), (2-18);

[0161] In summary, the interaction process between the agent and the environment in CMDP is as follows: Figure 5 shown.

[0162] The agent observes the current state, takes actions to reach the next state, and iteratively updates its strategy π based on the rewards corresponding to its actions. Unlike an MDP, strategy π must maximize the cumulative reward while also satisfying the system's operational constraints. The specific modeling formulas are shown in Equations (2-19) through (2-22).

[0163] π=argmax(J(π)), (2-19);

[0164]

[0165]

[0166]

[0167] Where J(π) is the cost corresponding to each reward under each strategy, JC (π) is the cost of each action under each strategy, and the corresponding cost is, is the upper bound of the corresponding cost under each strategy, τ is the state sequence extracted from the distribution of strategy π, and E is the expectation.

[0168] In this embodiment, the cost function corresponding to the policy in the CMDP is associated with the system state and the actions performed by the agent. Only when the cumulative cost is below a set threshold can the agent's actions be guaranteed to be safe, that is, the distribution network will not exceed the voltage limit. If the agent performs an unsafe action in a certain state, a cost penalty is imposed on its behavior. In addition, by considering the cumulative cost to define the safety concept of the entire trajectory, the long-term safety of the agent can be effectively evaluated, rather than focusing solely on a single state and action.

[0169] S30, and optimize and update the soft-switch three-phase power timing optimization model based on the Lagrangian function method security reinforcement learning algorithm; apply the optimized and updated soft-switch three-phase power timing optimization model to dispatch the power of the distribution network.

[0170] In this embodiment, the Lagrangian function method and the SAC (Soft Actor-Critic, SAC) algorithm are combined to propose a secure reinforcement learning algorithm based on the Lagrangian function method, and the implementation process of the algorithm is constructed.

[0171] Flexible action-evaluation algorithm:

[0172] In this embodiment, this step uses the SAC algorithm under the AC (Actor-Critic) framework and the Markov decision process constructed in step S20. SAC is a DRL (Deep Reinforcement Learning, DRL) algorithm based on the maximum entropy principle. Compared with the DDPG (deep deterministic policy gradient, DDPG) or TD3 (Twin Delayed Deep Deterministic policy gradient algorithm) algorithm, SAC further enhances the randomness of the strategy by introducing entropy, significantly improving the exploration ability of the agent. The core of the SAC algorithm lies in entropy regularization. The agent needs to maximize the reward value and entropy value at the same time during the training process. As a measure of randomness in the agent's strategy, entropy can effectively guide the agent to conduct more extensive or more limited exploration, thereby avoiding the strategy from converging to a local optimal solution.

[0173] Specifically, a higher entropy value means a more random strategy will choose actions, significantly improving the agent's exploration capabilities. By introducing entropy, the agent can not only select actions that maximize rewards during learning but also actively explore other potential options. As entropy increases, the agent can better balance exploring new strategies with leveraging known ones, enabling a more comprehensive decision-making process in complex distribution network environments. The formula for calculating entropy is shown in Equation (3-1).

[0174]

[0175] Where H is the entropy term, π(·|s t ) is the agent action strategy based on mapping the state space to the action space.

[0176] The reward obtained by the agent at each time step is proportional to the policy entropy. The reinforcement learning formula based on maximum entropy is shown in formula (3-2).

[0177]

[0178] Where α is the temperature coefficient, which determines the importance of entropy relative to the reward value, thereby ensuring the randomness of the optimal strategy, and π * The updated strategy.

[0179] The SAC algorithm consists of two neural networks: the policy network π φ And the value network (Q network). During the training process, SAC will also learn a policy network π φ and two Q functions Q θ1 and Q θ2 The Q network parameterized by θ is used to approximate the soft Q-function Q θ (s,a), while the policy network parameterized by φ outputs the mean and variance of the action probability distribution.

[0180] By calculating entropy and expected reward, the Bellman equation of the SAC algorithm is shown in Equation (3-3).

[0181]

[0182] The update of the policy network follows formula (3-4).

[0183]

[0184] Where D is the experience replay pool, Q(s t ,a t ) is the state action value function, J π (φ) is the loss function of the strategy, π φ (at |s t ) is the action a under the parameter φ t , state s t For the strategy, Q θ (s t ,a t ) is the state-action value function under the θ parameter.

[0185] The random strategy optimization characteristics of the SAC algorithm determine that its actions are not directly output by the strategy network, but the strategy is reparameterized through neural network transformation. Its action function is shown in formula (3-5).

[0186]

[0187] Where, ε t is a noise vector following a normal distribution; and are the mean and variance of the re-parameter sampling output, respectively, f φ (ε t ;s t ) is the parameterized function of the policy network with state t and noise as input, J π (φ) can be re-expressed as formula (3-6).

[0188]

[0189] The gradient strategy can be approximated using formula (3-7).

[0190]

[0191] Where, is the policy loss after gradient update, is the gradient symbol, is the direct gradient of the policy parameter θ with respect to the log probability of the action, is the gradient of the action with respect to the log probability of the policy, is the negative gradient of the Q network to the action, is the action of gradient update under the parameter φ.

[0192] The Q network is updated based on minimizing the Bellman residual. The specific formulas are Equations (3-8) to (3-9).

[0193]

[0194]

[0195] Where, is the updated target Q network parameter, J Q (θ) is the loss of the state-action-value network, is the state-action-value function after the θ parameter is updated.

[0196] The gradient strategy can be expressed using formula (3-10).

[0197]

[0198] The target Q network performs soft update, and its update formula is shown in the following formula (3-11).

[0199]

[0200] In summary, as a maximum entropy DRL algorithm, the SAC algorithm explores all possible optimal paths during the learning process by continuously updating the policy network and Q network. Furthermore, SAC employs technologies such as a dual Q network and a target network to further reduce the bias in Q-value estimation, making the model more stable during training and preventing excessively high or low Q-value estimates from impacting the quality of policy updates. As a result, the SAC algorithm demonstrates superior exploration performance in complex tasks and effectively avoids the problem of the algorithm being trapped in local optimal solutions.

[0201] Design of flexible action-evaluation algorithm based on Lagrangian function method:

[0202] In one embodiment of the present invention, the SAC algorithm proposed in the above steps is only applicable to the MDP model. For CMDP problems, a common approach is to convert the constraints into penalty terms and add them to the reward function, thereby converting the problem into an MDP and then solving it using the SAC algorithm. However, this conversion method makes it difficult to accurately measure the impact of constraint violations on scheduling results. Furthermore, the design of the penalty term directly affects the learning process; inappropriate settings may lead to overly conservative or even infeasible strategies.

[0203] To overcome the limitations of the aforementioned methods, this paper expands on the SAC algorithm, introduces Lagrange multipliers, and proposes a SAC algorithm based on the Lagrange function method. The core idea of ​​the secure reinforcement learning algorithm AL-SAC based on the Lagrange function method is to introduce constraints into the objective function, dualize the constraints using Lagrange multipliers, and embed them into the objective function.

[0204] Based on the constructed soft-switching three-phase power timing optimization model, the algorithm designed in this step requires that, under a given policy π, the cumulative expected cost be less than a safety threshold. If this condition is met, policy π is relatively safe; if the policy also maximizes the objective function J(π), it is considered optimal. In summary, expanding the constructed soft-switching three-phase power timing optimization model, we obtain Equation (3-12).

[0205] Formula (3-41) is the Lagrangian form of the original function, which is obtained by maximizing The optimal strategy π can be obtained.

[0206]

[0207] Where λ and α L are the Lagrange multipliers of constraint and entropy, respectively.

[0208] In the AL-SAC algorithm, both the neural network parameters and the Lagrange multipliers need to be updated simultaneously. Regarding the neural network, a Q-network is added to the original SAC algorithm to evaluate the cost function. Furthermore, the entropy regularization term is removed during the Q-network training process. This ensures that the state-action function accurately estimates the reward function, avoiding the additional fluctuations introduced by policy entropy that can affect optimization stability. Furthermore, the state cost function is designed to strictly reflect the cumulative cost of the constraints. The inclusion of an entropy term could lead to biased constraint cost estimates, thereby interfering with the dynamic update of the Lagrange multipliers. The specific training process for the state cost function is described below.

[0209] The state-action cost function is similar to the state-action value function and can also be approximated by the Bellman optimal equation. Therefore, the state-action value evaluation network and the state-action cost evaluation network are shown in Equations (3-13) to (3-14) respectively.

[0210]

[0211]

[0212] Where, is the parameter of the state-action cost evaluation network, is the loss of the value network, and J Q (θ) is the loss of the evaluation network.

[0213] The training objective of the policy network is to maximize the Lagrangian function. Therefore, to achieve a saddle point for the primal and dual variables, this step uses the primal-dual method to update the variables. Let x, α, and λ be the saddle points of the Lagrangian function, where x is the primal variable representing the neural network parameters; α and λ are both dual variables. Using the defined value function, Equation (3-12) is transformed into Equation (3-15).

[0214]

[0215] The update process of the original variable x and the dual variables α, λ is shown in equations (3-16) to (3-18)

[0216]

[0217]

[0218]

[0219] Where, δ x ,δ α ,δ λ is the update step size of the parameters; [] + is the projection of a non-negative real number, where δ x , δ α , δ λ are the update steps of different parameters, x k , α k ,λ k are the original variables and dual variables updated for the kth time, are the updated gradients of different parameters respectively.

[0220] The algorithm proposed in this step extends the traditional SAC algorithm by expanding the Q network and simultaneously training multiple Q networks to approximate the state-action benefits and costs of the CMDP. During policy network training, this algorithm utilizes a Lagrangian function approach to reduce losses and alleviate three-phase imbalance while ensuring the system operates within a safe range. By iteratively updating the Lagrangian multiplier, the algorithm demonstrates improved convergence and stability.

[0221] In summary, the pseudo code of the secure reinforcement learning algorithm AL-SAC based on the Lagrangian function method is shown in Table 1 below.

[0222] Table 1. Secure reinforcement learning algorithm based on Lagrangian function method

[0223]

[0224]

[0225] Example 2

[0226] The following further describes in detail the method for optimizing the power dispatch of three-phase power flow for soft switching according to the present invention with reference to the accompanying drawings and typical cases:

[0227] Step 1: The improved IEEE 123 node test system topology is as follows: Figure 6 As shown in the figure, the system has an active load of 3490kW, a reactive load of 1920kVar, and a rated voltage of 4.16kV. The safe interval of the system voltage is [0.95pu, 1.05pu]. In the objective function, the weight coefficients κ1 and κ2 are 0.68 and 0.32 respectively. Three two-port SOPs are connected to the distribution network. The port capacity of each SOP is 300kVA. The port loss coefficient of the SOP is A.SOP is 0.02, and its access position is as follows Figure 6 As a typical distributed power source, this step considers connecting 10 distributed photovoltaic units to the distribution network. Their connection locations and connection capacities are shown in Table 2. The training data comes from actual data from a city power supply company from 2019 to 2022, and the algorithm parameters are shown in Table 3.

[0228] Table 2 Distributed photovoltaic access location and parameters

[0229]

[0230] Table 3 Parameters of the secure reinforcement learning algorithm based on the Lagrangian function method

[0231]

[0232] In order to verify the effectiveness of the algorithm AL-SAC proposed in the present invention, this step will set up different DRL algorithms for comparative experiments. The algorithms used for comparison include the SAC algorithm and DDPG algorithm with the best performance at present. In order to apply these comparative algorithms to the SOP three-phase active / reactive optimization scheduling problem proposed in the present invention, it is necessary to modify the reward function. The specific approach is to introduce the constraint conditions into the reward function in the form of a penalty function, and the modified reward function is shown in Formula (3-19).

[0233]

[0234] Where, κ p is the weight coefficient of the penalty function, and its size is set to 1.5.

[0235] For the SAC algorithm, except for the difference in the Lagrange multiplier update, the other parameters remain the same as those of the AL-SAC algorithm. The neural network parameters of the DDPG algorithm are the same as those of the AL-SAC algorithm proposed in this invention.

[0236] Step 2: Algorithm performance evaluation and analysis

[0237] In this section, we will deeply analyze the actual effects of the proposed AL-SAC algorithm and two other typical DRL algorithms (SAC and DDPG) in the SOP three-phase active / reactive optimization scheduling task. All DRL algorithm training processes are set to 5000 rounds. The iterative convergence process of each DRL algorithm is as follows: Figure 7 shown.

[0238] Comparing the AL-SAC algorithm, the SAC algorithm, and the DDPG algorithm revealed that the AL-SAC algorithm converged the fastest between 2000 and 3000 rounds, reaching a stable reward value at approximately 2800 rounds. In contrast, the SAC algorithm and the DDPG algorithm still had not achieved convergence after 3000 rounds. Comparing the convergence performance of the three algorithms between 4000 and 5000 rounds revealed that all three converged to a stable reward value, but the AL-SAC algorithm achieved the highest stable reward value, indicating that the AL-SAC algorithm is more inclined to learn strategies that result in smaller system losses and lower three-phase imbalance.

[0239] In summary, traditional DDPG and SAC algorithms can only guide agents to avoid unsafe strategies through penalty functions, which not only increases the agent's learning cost but also affects the learning path, ultimately leading to poor convergence performance. In contrast, the AL-SAC algorithm proposed in this paper effectively avoids the training instability problems caused by traditional penalty functions. It can minimize system losses and three-phase imbalance while ensuring the safe operation of the distribution network, and avoids the negative impact of the algorithm caused by the unreasonable selection of penalty coefficients.

[0240] To further analyze the proposed AL-SAC, we analyzed the unsafe scenarios that occurred during the training of each algorithm. When the voltage at any node in the distribution network exceeds the limit, the distribution network is considered to be operating in an unsafe state. Figure 8 The unsafe scenarios that occurred when each algorithm was trained for 1500 rounds are shown.

[0241] Depend on Figure 8 It can be seen that the SAC algorithm and the DDPG algorithm failed to effectively avoid the occurrence of unsafe scenarios within 1500 rounds, that is, the trained agent strategy caused the distribution network to have voltage over-limit problems. In sharp contrast, the AL-SAC algorithm proposed in this invention was able to avoid unsafe scenarios within 300 rounds, significantly improving the safety of the strategy. In the early stages of training, the strategy of the AL-SAC algorithm agent may also cause unsafe scenarios, which is mainly due to the strategy exploration process of the algorithm in the early iterations. As the strategy gradually becomes better, the number of over-limit scenarios gradually decreases, and the executed strategy can eventually meet the system voltage constraints.

[0242] In order to verify the effectiveness of the present invention compared to the traditional model-driven method, this section sets up an SDP model for comparative experiments. Since the training process of the DRL algorithm is performed offline, this step focuses on comparing the online decision time of the two on the test set. The solution process uses the YALMIP optimization modeling toolbox and Gurobi optimization solver in MATLAB. The test hardware is the same as the test hardware of this patent algorithm. 10 typical test scenarios were randomly selected from the test scenario set. These scenarios do not overlap with the training scenarios at all. The test results are shown in Table 4.

[0243] Table 4 Computational performance test results of different algorithms

[0244]

[0245] As can be seen in Table 4, the data-driven AL-SAC algorithm proposed in this paper demonstrates significant advantages in computational efficiency. This is because the DRL algorithm only requires the neural network to perform feedforward operations during testing, while model-based methods require solving the overall optimization problem at runtime, which consumes a considerable amount of computation time. Furthermore, as the scale of distribution network topologies continues to increase, the DRL algorithm's advantage in decision speed will become even more pronounced.

[0246] Computational accuracy analysis shows that this method can achieve similar results to the SDP model, proving that this method can find the global optimal solution. In summary, the method proposed in this paper not only has computational accuracy comparable to traditional model-driven methods, but also significantly reduces decision-making time, making it more valuable for application.

[0247] Step 3: SOP control capability assessment and analysis

[0248] A set of data is randomly selected in the test set to evaluate the SOP control effect, such as Figure 9(a) to Figure 9(d) Taking SOP2 as an example, we analyze the regular trend of the active power and reactive power emitted by the two ports of SOP changing with time, as shown in the figure below. Figure 10(a) to Figure 10(f) shown.

[0249] During the periods of t = 11h-13h (midday) and t = 19h-21h (nighttime), PV output reaches its maximum and load demand peaks, respectively. During these periods, load imbalance between the phases of the distribution network is particularly prominent, necessitating active power transfer via the SOP to alleviate the three-phase load imbalance between the connected feeders and achieve optimal operation. Taking the midday period (t = 11h-13h) as an example, the active power of phase c of the VSC-55 is positive, while that of the VSC-93 is negative, indicating that the active power of phase c is transferred from node 55 to node 93. During this process, the excess power generated by the PV system is effectively absorbed, while optimizing the system's power flow distribution. During the nighttime period (t = 19h-21h), due to zero PV output and peak system net load, the active power output of the SOP port increases significantly, effectively balancing the active power distribution within the system, reducing system losses, and alleviating the three-phase voltage imbalance.

[0250] Furthermore, the SOP not only regulates three-phase active power but also outputs reactive power through each port, providing reactive power compensation for the system and thus increasing voltage amplitude. By providing local three-phase reactive power support, the SOP directly meets the reactive power requirements of the load, avoiding the limitations of remote reactive power transmission from the upper grid, effectively reducing system power losses and improving voltage distribution.

[0251] Step 4: Verification of the advantages of SOP applied to three-phase unbalanced distribution network

[0252] In this embodiment, two typical scenarios are set up for comparative analysis.

[0253] Scenario 1: No SOP in the system

[0254] Scenario 2: Three SOPs are connected to the system

[0255] The test results of network loss, SOP loss, voltage over-limit rate and maximum voltage imbalance in different scenarios are shown in Table 5.

[0256] Table 5 Optimization results under different scenarios

[0257]

[0258]

[0259] The test results in Table 5 show that the soft-switching three-phase power timing optimization model proposed in this paper can effectively alleviate the system losses and voltage offset problems in three-phase unbalanced distribution networks. This is because the integration of the SOP into the distribution network enables continuous power flow regulation, thereby effectively improving the overall operational economy of the system.

[0260] In order to further analyze the effect of SOP on alleviating the three-phase imbalance of the system, this step selected node 85 with significant imbalance characteristics (both a-phase load and b-phase load are 0) as the test object, and analyzed the changes in its three-phase voltage imbalance within 24 hours. The test results are as follows: Figure 11 shown.

[0261] Compared with Scenario 1, the voltage imbalance at node 85 in Scenario 2 is significantly alleviated due to the access of SOP. Within 24 hours, the voltage imbalance is always kept below 0.05. For example, during the night period (t=19h-21h), the loads of phases a and b at node 85 are 0, while the load of phase c reaches its peak, resulting in a serious imbalance in the three-phase load. When SOP is not connected, the system lacks effective control measures, and it is easy for the three-phase imbalance to exceed the limit, thereby increasing system losses and possibly damaging the distribution equipment. After the SOP is connected, the three-phase load of the node is effectively balanced through active power transmission and reactive power compensation, significantly alleviating the three-phase imbalance problem. In addition, the three-phase imbalance curve of node 85 becomes smoother after the SOP is connected, indicating that the system is in a more stable operating state.

[0262] Further analysis of the impact of SOP regulation on system flow distribution, Figure 12 The current distribution from branch 75 to branch 82 at t = 20h is shown.

[0263] It can be seen that compared with scenario 1, after scenario 2 is optimized using SOP, the three-phase active power transmitted between node 55 and node 93 is Figure 10(a) to Figure 10(f) As shown, this significantly reduces reliance on long-distance power transmission from the upstream power grid. Simultaneously, the current amplitude between Branch 75 and Branch 82 has been significantly reduced, effectively reducing network losses. Furthermore, through SOP power flow optimization, the system current distribution has been effectively balanced.

Claims

1. A method for optimizing the power flow scheduling of three-phase power flow in soft switching, characterized in that: include: Obtain a mathematical model based on soft-switching three-phase power control characteristics; Based on the mathematical model of the soft-switching three-phase power control characteristics, combined with the three-phase power flow constraints of the distribution network, the distribution network safety operation constraints, and the soft-switching operation constraints, a soft-switching three-phase power timing model is constructed. This soft-switching three-phase power timing model is transformed into a Markov decision process based on cost constraints, and a soft-switching three-phase power timing optimization model is obtained. The soft-switching three-phase power timing optimization model is optimized and updated using a secure reinforcement learning algorithm based on the Lagrangian function method. The optimized and updated soft-switching three-phase power timing optimization model is used to dispatch the power of the distribution network.

2. The method for optimizing the power scheduling of three-phase power flow with soft switching according to claim 1, characterized in that: The expression of the soft switching three-phase power timing model is: Where M SOP PDN is a mathematical model based on the soft switching three-phase power control characteristics. st are the three-phase power flow constraints of the distribution network, the safe operation constraints of the distribution network, and the soft switch operation constraints, st is the constraint, is the objective function, C OBJ For the goal, C LOSS is the total system loss, C UNB is the system three-phase voltage unbalance, κ1 and κ2 are the weight coefficients of the total system loss and the three-phase voltage unbalance in the objective function respectively.

3. The method for optimizing the power scheduling of three-phase power flow with soft switching according to claim 1, characterized in that: The three-phase power flow constraints of the distribution network are: In the formula, ij represents that the starting point of the branch is node i and the end point is node j; ki represents that the starting point of the branch is node k and the end point is node i; S ij,t and S ki,t are the three-phase power matrices transmitted by branch ij and branch ki at time t; Z ij and Z ki are the three-phase impedance matrices of branch ij and branch ki respectively; and are the three-phase voltage vectors of node i and node j at time t; L ij,t and L ki,t are the matrices obtained by multiplying the three-phase current vectors of branch ij and branch ki at time t by their own conjugate transpose; is the column vector I ij,t The conjugate transpose of H is conjugate; S net,i,t is the three-phase injection power matrix of node i at time t; S net,i,t,a 、S net,i,t,b 、S net,i,t,b are the injected power phasors of phase a, phase b and phase c of node i at time t respectively; is the node i at time t Phase injection power phasor; and They are respectively the node i at time t Phase load active and reactive demands; DG is at time t is the active power transmitted by the phase, DG is the distributed generation, j is the imaginary unit, and T is the transpose.

4. The method for optimizing the power scheduling of three-phase power flow with soft switching according to claim 1, characterized in that: The constraints for safe operation of the distribution network are: Where V min and V max are the minimum and maximum values ​​of the node voltage, respectively. is the node i at time t Phase complex voltage.

5. The method for optimizing the dispatching of three-phase power flow with soft switching according to claim 1, characterized in that: The mathematical model based on the soft switching three-phase power control characteristics is used as the condition of the soft switching operation constraint.

6. The method for optimizing and dispatching three-phase power flow for soft switching according to claim 1, characterized in that: The soft-switching three-phase power timing model is converted into a cost-constrained Markov decision process, including: The three-phase active power, three-phase reactive power and DG output of all nodes in the distribution network are defined as the system state s t ; Define action a by inputting / outputting the three-phase active power and three-phase reactive power of the soft switch port. t ; Reducing system losses and alleviating the three-phase imbalance of the system is defined as the reward R t ; The violation of the safe operation constraint of the distribution network is defined as the cost function In a given state s t and action a t Under this condition, the next state of the system will be transferred based on the probability transfer function Pr; after executing the action, it reaches the next state, and continuously iterates and updates the strategy π through the reward corresponding to the action.

7. The method for optimizing and dispatching three-phase power flow for soft switching according to claim 3, characterized in that: The expression of the soft switching three-phase power timing optimization model is: π=argmax(J(π)); Where J(π) is the cost corresponding to each reward under each strategy, J C (π) is the cost of each action under each strategy, and the corresponding cost is, is the upper bound of the corresponding cost under each strategy, τ is the state sequence, E is the expectation, and γ is the discount factor.

8. The method for optimizing and dispatching three-phase power flow for soft switching according to claim 1, characterized in that: The soft-switching three-phase power timing optimization model has been optimized and updated, including: The Lagrangian function method's secure reinforcement learning algorithm includes a state-action value network, a state-action cost network, and a policy network; Among them, the expression of the state-action-value network is: The expression of the state-action cost network is: The expression of the policy network is: Where, J Q (θ) is the loss of the state-action-value network, D is the experience replay pool, Q after parameter θ update θ network, is the updated Q network parameter, θ is the parameter of the Q network, is the loss of the state-action cost network, For parameters Updated Network, s t 、a t 、R t ,γ, E represents the state, action, reward, discount factor, cost function and expectation in the Markov decision process; is the loss of the policy network, x is the original variable, α and λ are the Lagrange multipliers of the temperature coefficient and entropy respectively, and are divided into dual variables, H is the entropy term, π φ (a t |s t ) is the action a under the parameter φ t , state s t For the strategy, is the upper bound of the corresponding cost under each strategy.

9. The method for optimizing and dispatching three-phase power flow for soft switching according to claim 5, characterized in that: In the policy network, the update process of the original variables and the dual variables is: Where, δ x , δ α , δ λ are the update steps of different parameters, x k , α k ,λ k are the original variables and dual variables updated for the kth time, are the updated gradients of different parameters, respectively. + is the projection of a non-negative real number.

10. A three-phase power flow optimization scheduling system suitable for soft switching, characterized in that: The method for optimizing and dispatching three-phase power flow for soft switching according to any one of claims 1 to 9 comprises: Basic model module, used to obtain the mathematical model based on the soft switching three-phase power control characteristics; The Markov decision module is used to construct a soft-switching three-phase power timing model based on the mathematical model of the soft-switching three-phase power control characteristics, combined with the distribution network three-phase power flow constraints, distribution network safe operation constraints, and soft-switching operation constraints. The soft-switching three-phase power timing model is converted into a Markov decision process based on cost constraints to obtain the soft-switching three-phase power timing optimization model; The update optimization module is used to optimize and update the soft-switching three-phase power timing optimization model using the Lagrangian function method-based security reinforcement learning algorithm; the optimized and updated soft-switching three-phase power timing optimization model is used to dispatch the power of the distribution network.

Citation Information

Patent Citations

  • Comprehensive energy system planning optimization method and system based on deep reinforcement learning

    CN119444502A

  • Method and system for improving transmission capacity of alternating-current and direct-current hybrid power transmission corridor

    CN119813402A