A method and system for three-phase imbalance treatment of energy storage power distribution network based on DSAC algorithm

CN122553264APending Publication Date: 2026-08-11STATE GRID JIANGSU ELECTRIC POWER CO LTD NANTONG POWER SUPPLY BRANCH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-11
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

由于分布式光伏出力具有显著的间歇性与随机性,且多为单相接入或不对称接入,叠加用户负荷时空分布的不均匀性,极易引发配电网严重的三相不平衡问题

Benefits of technology

(1) 本发明通过支持独立分相控制的储能系统,将三相有功和无功功率作为连续可调参数输入智能体,实现对各相电压偏差和不平衡度的独立快速调节,相比传统统一调节方式具有更高的控制精度和更强的补偿能力,可有效降低中性点电流、电压不平衡率及相间电压偏差,提高电能质量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122553264A_ABST
    Figure CN122553264A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of active distribution network operation control and artificial intelligence application technology, specifically involving a method and system for managing three-phase imbalance in energy storage distribution networks based on the DSAC algorithm. This invention proposes a DSAC algorithm with adaptive risk preference based on operating conditions, applied to Markov decision processes for training the agent. Data is extracted from an experience replay buffer to update the policy network and evaluation network for offline training of the agent. The risk preference coefficient during the policy update process is adaptively adjusted using an operating condition severity index. This invention deeply integrates the phase-decoupling adjustment capability of the energy storage system with the operating condition adaptive risk preference mechanism of the DSAC algorithm, enabling the algorithm to dynamically adjust the trade-off between benefits and risks based on photovoltaic fluctuations, the severity of three-phase imbalance, and the safety margin of energy storage SOC. This achieves rapid and robust management of three-phase imbalance in the distribution network, improving the power quality and operational stability of high-penetration photovoltaic distribution networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of active distribution network operation control and artificial intelligence application technology. Specifically, it relates to a three-phase imbalance management method for energy storage distribution networks based on the DSAC algorithm, which is applicable to power quality management and safe and stable operation of distribution networks under high-penetration distributed photovoltaic access. Background Technology

[0002] With the deepening implementation of the "dual-carbon" strategy, the penetration rate of distributed photovoltaic (PV) power in distribution networks is continuously increasing, and distribution networks are gradually transforming from passive networks to active distribution networks with uncertain source and load. Due to the significant intermittency and randomness of distributed PV power output, and the fact that it is mostly single-phase or asymmetrically connected, coupled with the uneven spatial and temporal distribution of user loads, severe three-phase imbalance problems in distribution networks are easily triggered. Three-phase imbalance not only leads to increased line and transformer losses and reduced equipment utilization, but may also cause neutral point voltage deviation and local voltage exceedances, seriously threatening the safe and stable operation of the distribution network. Traditional mitigation methods, such as switching capacitor banks and regulating transformer taps, are limited in the number of operations and have slow response speeds, making them difficult to cope with high-frequency fluctuations in PV power output. In contrast, energy storage systems possess millisecond-level response speeds and four-quadrant power regulation capabilities, achieving flexible power transfer between phases through independent phase control strategies, becoming an effective physical means of mitigating three-phase imbalance in high-penetration PV distribution networks. High-penetration photovoltaic distribution networks refer to a state in which the total amount of distributed photovoltaic power connected to the distribution network is so high that it has a substantial impact on the power flow direction, voltage stability, and protection configuration of the power grid.

[0003] However, existing methods for researching control strategies for energy storage systems still have significant shortcomings. Traditional optimization methods based on physical models heavily rely on accurate distribution network topology parameters, and their computational complexity increases rapidly with network size, making it difficult to meet the needs of online real-time control. While deep reinforcement learning methods, which have emerged in recent years, can achieve model-free real-time decision-making, mainstream algorithms such as DDPG (Deep Deterministic Policy Gradient) and SAC (Soft Actor-Critic) are mostly based on maximizing expected returns, focusing only on the average benefit of the strategy and failing to reflect the distributed risk of control results under random fluctuations in source and load. For example, Chinese invention patent CN120033728A discloses a method and system for generating energy storage charging and discharging strategies based on reinforcement learning. This patent selects the SAC algorithm as the core algorithm of reinforcement learning and dynamically adjusts the strategy to adapt to the real-time operating environment. It aims to optimize the charging and discharging behavior of energy storage systems through intelligent algorithms, smooth out fluctuations on the generation side, meet the dynamic demand on the consumption side, improve the stability of the power system, and maximize economic and environmental benefits.

[0004] No existing technologies applying the DSAC algorithm to energy storage system control strategy research were found. Even if existing technologies apply the Distributive Soft Actor-Critic (DSAC) algorithm to this scenario, due to the limitations of the algorithm itself, its risk-reward trade-off remains unchanged under different operating conditions. Therefore, even if applied to energy storage system control strategies, this algorithm is difficult to adapt to the constantly changing operating characteristics of high-penetration photovoltaic distribution networks, including photovoltaic fluctuations, the severity of three-phase imbalances, and the constantly changing state of charge (SOC) of energy storage. Especially when photovoltaic fluctuations are severe, voltage imbalances worsen, or the SOC of energy storage approaches its limit, a fixed-risk-preference control strategy may still produce excessive regulatory actions, leading to increased risk of exceeding limits or exacerbated control oscillations.

[0005] Therefore, there is an urgent need to develop a control method that can fully utilize the phase regulation potential of energy storage and dynamically adjust risk preferences according to the current operating conditions of the distribution network, so as to ensure safe and robust three-phase imbalance management of the distribution network under uncertain environments. Summary of the Invention

[0006] Objective: To overcome the shortcomings of existing technologies, this invention discloses a three-phase imbalance management method for energy storage distribution networks based on the DSAC algorithm. Addressing the severe three-phase voltage and power imbalance caused by random source-load fluctuations in high-penetration photovoltaic distribution networks, it proposes a phase-by-phase collaborative control framework for energy storage that considers operational risks. Unlike methods that directly apply the DSAC algorithm to distribution network regulation scenarios, this invention, considering the characteristics of high-penetration photovoltaic distribution networks—strong operating condition fluctuations, time-varying operational risks, and significant influence of SOC constraints on energy storage regulation capabilities—introduces an adaptive risk preference mechanism. This allows the algorithm to dynamically adjust the trade-off between benefits and risks based on the current operating state. Furthermore, this application also provides a three-phase imbalance management system for energy storage distribution networks based on the DSAC algorithm.

[0007] Technical Solution: In a first aspect, this invention provides a method for managing three-phase imbalance in energy storage distribution networks based on the DSAC algorithm. This method includes the following steps: A mathematical model for managing three-phase imbalance in a distribution network including distributed photovoltaic (PV) and energy storage systems is constructed. Using the three-phase voltage amplitudes, active and reactive power of the load, output of the distributed PV system, and the state of charge of the energy storage system at a certain moment in the mathematical model, a state space for a Markov decision process is constructed. The action space for the Markov decision process is constructed using the active and reactive power commands of the energy storage system at a certain moment in the mathematical model. The reward function for the Markov decision process is constructed based on the three-phase voltage amplitudes, three-phase voltage imbalance, power loss, and state of charge constraints of the energy storage system at the nodes in the model. An improved value distribution reinforcement learning algorithm is applied to the Markov decision process for training the agent. The improved value distribution reinforcement learning algorithm includes: constructing a condition severity index based on the current operating conditions of the distribution network, and using the condition severity index to adaptively adjust the risk preference coefficient in the policy network update process. The trained agent is deployed to the power distribution network for online optimization, and energy storage is scheduled according to the optimal action strategy to achieve three-phase imbalance management.

[0008] Furthermore, including: The reward function for constructing the Markov decision process based on the three-phase voltage amplitude, three-phase voltage imbalance, power loss, and state of charge constraints of the energy storage system at the nodes in the model includes: exist t The reward function at time t is expressed as: ; in, These represent the penalties for exceeding the limit of the node voltage amplitude. Three-phase voltage imbalance penalty Power loss penalty and the degree of over-limit of the state of charge of the energy storage system Weighting coefficients; The penalty for exceeding the limit of the node voltage amplitude Represented as: ; In the formula, N It is the collection of voltages at all three-phase nodes in the distribution network. Three-phase nodes i of The extent to which the phase voltage exceeds the upper and lower limits; express It belongs to phases A, B, and C, and ; in, Three-phase nodei of Phase voltage amplitude.

[0009] Furthermore, including: The three-phase voltage imbalance penalty Represented as: ; ; in, Indicates the negative sequence unbalance of three-phase voltage; These represent the zero-sequence, positive-sequence, and negative-sequence components of the three-phase voltage, respectively. These represent the voltage phasors of phases A, B, and C, respectively. and Indicates the rotation factor; The degree of state of charge exceeding the limit of the energy storage system Represented as: ; In the formula, Indicates the distribution network area number j Taiwan's energy storage system t time Phase state of charge, express t time Minimum constraint value for the charge state of a phase. express t time Maximum constraint value of phase charge state, m This indicates the number of energy storage systems in the distribution network.

[0010] Furthermore, including: The improved value distribution reinforcement learning algorithm is applied to the process of training the agent using the Markov decision process, including: Two identical evaluation networks and one policy network are established; the evaluation networks are used to output the Gaussian distribution parameters of the state-action pair reward values, and the policy network is used to output the Gaussian distribution parameters of the actions. When calculating the Bellman target value, an expected value substitution mechanism is adopted, which uses the expected Q value output by the target evaluation network to replace the random sampling return, and selects the one with the smaller mean of the two target evaluation networks as the benchmark to eliminate random noise and suppress Q value overestimation. The severity index of the operating conditions is constructed based on the current operating conditions of the distribution network. The severity index of the operating conditions is composed of the intensity of photovoltaic power output fluctuation, the severity of three-phase voltage imbalance, and the deviation of the safety margin of the state of charge of the energy storage system. The time-varying risk preference coefficient is calculated based on the severity index of the operating conditions. By utilizing the distribution information from the evaluation network feedback and the time-varying risk preference coefficient, while maximizing the cumulative expected return and strategy entropy, the strategy network is updated through reparameterized sampling. This allows the network to adaptively adjust the degree of suppression of return volatility risk under different operating conditions, thus tending to select high-yield, low-risk control actions that match the current operating conditions.

[0011] Furthermore, including: The degree of deviation of the state-of-charge safety margin of the energy storage system is expressed as follows: ; in, m This indicates the number of energy storage systems in the distribution network.

[0012] Furthermore, including: The calculation of the time-varying risk preference coefficient based on the severity index of the working condition includes: The time-varying risk preference coefficient is expressed as follows: ; in, The function is a numerical clipping function. Basic risk preference coefficient; These are the upper and lower limits of the time-varying risk preference coefficient, respectively. To adjust the gain, This is an indicator of the severity of the operating condition; The severity index of the operating condition is expressed as: ; in, The weighting coefficients for operating condition indicators. Indicates the intensity of photovoltaic fluctuations. As a penalty for three-phase voltage imbalance, the photovoltaic fluctuation intensity is expressed as: ; in, Indicates from Time's up t Photovoltaic power output sequence at any given time The function represents a function that directly calculates the variance of an array.

[0013] Furthermore, including: The method of utilizing the distribution information fed back by the evaluation network and the time-varying risk preference coefficient to maximize the cumulative expected return and policy entropy while updating the policy network through reparameterized sampling includes: The updated policy network is represented as: ; in, This represents the expectation operator, which is used to sample the state from the experience replay pool D. s and the policy network in that state Generated actions aThe corresponding objective function is averaged. For the first i The mean of the output of the evaluation network for each target is expressed as: ; in, and They represent the first i The mean and standard deviation of the state-action-reward distribution output of the target evaluation network.

[0014] On the other hand, the present invention also provides a three-phase imbalance mitigation system for energy storage distribution networks based on the DSAC algorithm, the system comprising: The model building module is used to construct a mathematical model for managing three-phase imbalance in a distribution network that includes distributed photovoltaic (PV) systems and energy storage systems. It utilizes the three-phase voltage amplitudes at each node of the distribution network, the active power and reactive power of the load, the output of the distributed PV systems, and the state of charge of the energy storage system at a certain moment in the mathematical model to construct the state space of a Markov decision process. It also utilizes the active and reactive power commands of the energy storage system at a certain moment in the mathematical model to construct the action space of the Markov decision process. Finally, it constructs the reward function of the Markov decision process based on the three-phase voltage amplitudes, three-phase voltage imbalance, power loss, and the state of charge constraints of the energy storage system at each node in the model. The agent training module is used to improve the value distribution reinforcement learning algorithm and apply it to the process of training the agent in the Markov decision process. The improved value distribution reinforcement learning algorithm includes: constructing a condition severity index based on the current distribution network operating conditions, and using the condition severity index to adaptively adjust the risk preference coefficient in the policy network update process. The online optimization module is used to deploy the trained agent to the distribution network for online optimization, and to schedule energy storage according to the optimal action strategy to achieve three-phase imbalance management.

[0015] Thirdly, the present invention also provides a computer device, comprising: one or more processors; The processor is used to store one or more programs; When the one or more programs are executed by the one or more processors, the three-phase imbalance management method for energy storage distribution networks based on the DSAC algorithm described above is implemented.

[0016] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed, it implements the three-phase imbalance management method for energy storage distribution networks based on the DSAC algorithm as described above.

[0017] Beneficial effects: Compared with the prior art, the present invention has the following advantages: (1) This invention uses an energy storage system that supports independent phase control to input three-phase active and reactive power as continuously adjustable parameters into an intelligent agent, thereby achieving independent and rapid adjustment of voltage deviation and unbalance of each phase. Compared with the traditional unified adjustment method, it has higher control accuracy and stronger compensation capability, and can effectively reduce neutral point current, voltage unbalance rate and phase-to-phase voltage deviation, thereby improving power quality.

[0018] (2) Unlike traditional optimization methods based on power flow equations, mathematical programming or sensitivity analysis, the DSAC scheduling strategy of this invention is obtained through interactive learning between the agent and the power grid environment. It does not rely on accurate distribution network modeling and can automatically converge to the optimal adjustment strategy under complex conditions such as random photovoltaic output, severe load fluctuations and continuous changes in power grid operating status, which significantly improves the engineering applicability and robustness to model uncertainty.

[0019] (3) This invention does not directly adopt the general DSAC algorithm with a fixed risk preference. Instead, it constructs a condition severity index based on the degree of photovoltaic fluctuation, the severity of three-phase imbalance and the safety margin of energy storage SOC, and adaptively adjusts the risk preference coefficient accordingly. This enables the algorithm to improve governance efficiency when the condition is stable and improve operational safety and control stability when the condition deteriorates, thereby significantly enhancing the adaptability of the DSAC algorithm to the three-phase imbalance governance scenario of high-penetration photovoltaic distribution network.

[0020] (4) The scheduling strategy obtained from training can be directly deployed to the actual distribution network for online operation, realizing the real-time tracking and rapid response of the energy storage system to the three-phase operating status. It does not require continuous solution of complex optimization problems, has a fast response speed, and can significantly improve the online operation control efficiency of the distribution network, and improve the three-phase voltage steady-state level and power supply reliability. Attached Figure Description

[0021] Figure 1 A flowchart of a three-phase imbalance control method for energy storage phase-controlled distribution networks based on the operating condition adaptive DSAC algorithm provided in an embodiment of the present invention; Figure 2 This is a physical topology diagram of a phase-controlled energy storage system provided in an embodiment of the present invention; Figure 3 A schematic diagram of the interaction principle of the Markov decision process provided in the embodiments of the present invention; Figure 4 A schematic diagram of the policy network and evaluation network structure based on the working condition adaptive DSAC algorithm provided in an embodiment of the present invention; Figure 5 A model diagram of an energy storage access IEEE 33-node distribution network system provided for embodiments of the present invention; Figure 6 A convergence graph of the average reward value per round during the training process of the DSAC algorithm provided in this embodiment of the invention; Figure 7 A comparison diagram of the three-phase voltage imbalance before and after treatment, provided for an embodiment of the present invention. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] This invention proposes a method for managing three-phase imbalance in a distribution network based on an adaptive DSAC algorithm for operating conditions, the steps of which are as follows: S1: Construct a mathematical model for the three-phase imbalance management of a distribution network that includes distributed photovoltaic and energy storage systems, including: distribution network safety operation constraints and energy storage system physical constraints; S2: Transform the three-phase imbalance management problem of high-penetration photovoltaic distribution network into a mathematical Markov decision process (MDP), which includes state space, action space, state transition probability function and reward function; S3: A DSAC algorithm with adaptive risk preference based on operating conditions is proposed and applied to Markov decision process to train the agent. Data is extracted from the experience replay buffer to update the policy network (Actor) and evaluation network (Critic) for offline training of the agent. Among them, the operating condition severity index is constructed according to the current operating conditions of the distribution network, and the risk preference coefficient in the policy update process is adaptively adjusted using the operating condition severity index. S4: Deploy the trained agent to the distribution network for online optimization, and schedule energy storage according to the optimal action strategy to achieve three-phase imbalance management and verify its effectiveness.

[0024] Furthermore, step S1, which involves constructing the mathematical model and defining constraints, specifically includes: S11: Establish a power flow calculation model for a three-phase four-wire distribution network, treat the output fluctuation of the photovoltaic inverter as a random state disturbance, and model the energy storage system as a three-phase independent adjustable active and reactive power source (A, B, and C). S12: Establish the constraints for safe operation of the distribution network and the physical constraints for the energy storage system, including node voltage amplitude constraints, three-phase voltage imbalance constraints, energy storage state of charge (SOC), and energy storage converter capacity constraints. These constraints will serve as the basis for the design of subsequent reward function penalty terms.

[0025] Furthermore, the energy storage model in step S11 is expressed as follows: (1) In the formula, This indicates independent phase control. for t time Phase charge state; For charge and discharge efficiency; For energy storage active power; For the rated capacity of energy storage, this model ensures that the energy storage system can achieve physical power transfer between phases through differentiated throughput of active power in each phase.

[0026] Furthermore, the formulas for the power distribution network safety operation constraints and the energy storage system physical constraints in step S12 are as follows: (1) Node voltage amplitude constraints (2) (2) State of charge constraints and capacity constraints of energy storage systems

[0027] In the formula, the upper limit of the node voltage is... Set to 1.05, lower limit Set to 0.95. This indicates a distribution network node; the maximum and minimum allowable values ​​for energy storage SOC are 0.8 and 0.2, respectively. The SOC value of the stored energy at the beginning of the day. The SOC value of the stored energy at the end of the day.

[0028] Furthermore, the Markov decision process in step S2 is defined as follows: S21: State Space S t Included t Real-time data on the three-phase voltage amplitude at each node of the distribution network, active / reactive power of the load, output of distributed photovoltaic systems, and state of charge of energy storage systems are used. express; S22: Motion Space Defined as in t Active power commands of the energy storage system on phases A, B, and C at any time and reactive power command ,use It means that, among them, m This indicates the number of energy storage systems.

[0029] S23: The reward function comprehensively considers node voltage amplitude, three-phase voltage imbalance, power loss, and the energy storage system's state of charge (SOC) constraint, i.e., the energy storage SOC constraint. t The reward function at time step is: (6) In the formula, These represent the penalties for exceeding the limit of the node voltage amplitude. Three-phase voltage imbalance penalty Power loss penalty and SOC limit exceedance The weighting coefficients.

[0030] Furthermore, this embodiment designs a penalty item for exceeding the node voltage amplitude limit based on the voltage exceeding the limit. : (7) In the formula, N It is the collection of voltages at all three-phase nodes in the distribution network. Three-phase nodes i of The extent to which the phase voltage exceeds the upper and lower limits; (8) Pressure imbalance penalty item The formula is as follows: (9) (10) In the formula, Indicates the negative sequence unbalance of three-phase voltage; These represent the zero-sequence, positive-sequence, and negative-sequence components of the three-phase voltage, respectively. In this embodiment, when using the above calculation method, it is stipulated that: under normal operation of the power grid, the negative-sequence voltage imbalance at the point of common coupling of the power system shall not exceed 2%, and shall not exceed 4% for short periods.

[0031] SOC exceeding limit R SOC The calculation formula is as follows: (11) In the formula, Indicates the distribution network area number j Taiwan Energy Storage t time The state of charge of a phase.

[0032] Power loss penalty Represented as: (12) in, For nodes i With nodes j The conductance of the line admittance between them. for t Time Node i exist Phase voltage amplitude, for t Time Node i With nodes In phase The voltage phase angle difference.

[0033] Furthermore, the DSAC agent network construction in step S3 specifically includes: constructing two evaluation networks with identical structures. and To fit the reward distribution of state-action pairs, each evaluation network outputs the mean of a Gaussian distribution of the reward. and standard deviation The corresponding return distribution is denoted as ; Constructing a policy network Both the evaluation network and the policy network are implemented using deep neural networks. The input layer receives the state vector, the hidden layer uses a fully connected layer and an activation function, and the output layer is mapped to the corresponding distribution parameters.

[0034] Furthermore, the core mechanism of the improved DSAC algorithm in step S3 is that the evaluation network (Critic) does not output a single Q-value, but instead outputs a reward value. Z ( s,a The probability distribution parameters of ) follow a Gaussian distribution. ,in Represents the average expected return. This represents the risk of return volatility caused by the inherent operational uncertainties of the distribution network and the randomness of photovoltaics; when updating, the policy network (Actor) does not adopt a fixed risk appetite, but instead constructs a condition severity index based on the current operating conditions. Based on this, a time-varying risk preference coefficient is generated. ,by As a comprehensive evaluation metric for returns and risks, it adaptively adjusts the intensity of suppression of return volatility risk while maximizing the expected cumulative return and strategy entropy. This allows the agent to tend to improve governance effectiveness when operating conditions are stable and to tend to improve operational safety and control stability when operating conditions deteriorate.

[0035] Furthermore, the DSAC agent construction and training process in step S3 specifically includes: S31: Establish two identical evaluation networks and one policy network; the evaluation networks output the Gaussian distribution parameters of the state-action pair reward values, i.e., the mean. with standard deviation The policy network outputs Gaussian distribution parameters for the actions; S32: When calculating the Bellman target value, the expected value substitution mechanism is adopted. The expected Q value output by the target evaluation network is used to replace the random sampling return, and the one with the smaller mean of the two target evaluation networks is selected as the benchmark to eliminate random noise and suppress Q value overestimation. That is, in this embodiment, according to the next state The next action is obtained by sampling from the policy network. Then, the state-action pair The input is fed into two target evaluation networks, and the Gaussian parameters, i.e., the mean, of their corresponding reward distributions are obtained. Standard deviation , The corresponding representation is: Since the expected value of a Gaussian distribution equals its mean, the mean of the output distribution of the target evaluation network is used as the expected Q value. The smaller of the two target evaluation network output means is selected as the target value estimate to reduce the risk of overestimating the Q value due to function approximation error. Therefore, the final Bellman target value is expressed as: (13) in, As a reward, Discount factor, For the next state, Next move, The mean value of the network output is used as the target evaluation value, which achieves smoothing of photovoltaic power output fluctuation noise. This is an entropy regularization term that encourages strategies to maintain randomness while pursuing high rewards, thereby enhancing exploration capabilities. The temperature parameter is used to balance the importance of reward and entropy.

[0036] Furthermore, in this embodiment, the evaluation network (Critic) in step S3 updates the network parameters by minimizing the agent's loss function. The specific formula for calculating the loss function is as follows: (14) Here, E[...] is the expectation operator, used to calculate the average value of a random variable.

[0037] S33: Construct a condition severity index based on the current operating conditions of the distribution network. The severity index of the operating condition is composed of the intensity of photovoltaic power output fluctuation, the severity of three-phase voltage imbalance, and the deviation of the energy storage SOC safety margin. A time-varying risk preference coefficient is calculated based on the severity index of the operating condition. .

[0038] In this embodiment, the operating condition severity index and the time-varying risk preference coefficient in step S33 are constructed as follows:

[0039] In the formula, This indicates the average deviation of the State of Charge (SOC) of each phase of the energy storage system from the median of the safe range at the current moment; It represents the variance of photovoltaic power output within a preset time window, which is used to characterize the fluctuation intensity of photovoltaic power output; Indicates from Time's up t The photovoltaic power output sequence at any given time is derived from historical photovoltaic power generation operation data as input to the photovoltaic power output sequence, and is used to characterize the actual fluctuation characteristics of photovoltaic power output. These represent the intensity of photovoltaic fluctuations, the severity of three-phase voltage imbalance, and the degree of deviation from the SOC safety margin, respectively. The weighting coefficients for operating condition indicators; Basic risk preference coefficient; To adjust the gain; These are the upper and lower limits of the time-varying risk preference coefficient, respectively.

[0040] S34: Utilize the distribution information from the target evaluation network feedback and the time-varying risk preference coefficient While maximizing the expected cumulative return and strategy entropy, the strategy network is updated by reparameterized sampling, enabling it to adaptively adjust the degree of suppression of return volatility risk under different operating conditions, thus tending to select high-yield, low-risk control actions that match the current operating conditions.

[0041] The update formula for the policy network (Actor) in step S34 is:

[0042] As can be seen from the above, The mean of the network output is used to evaluate the target, that is... For the first i The mean of the network output for evaluating each target. and They represent the first i The mean and standard deviation of the state-action pair reward distribution output of the evaluation network; This is the time-varying risk preference coefficient determined by the current operating conditions.

[0043] That is, in a preferred embodiment of this invention, when photovoltaic fluctuations increase, three-phase imbalance increases, or the energy storage SOC approaches its limit, Increase the policy network's ability to mitigate risk volatility. The suppression weights; when the system is running smoothly, Reduce, the strategy network improves the expected return term Optimized weights.

[0044] The further step S4, the online execution and scheduling process, is as follows: the policy network parameters that have been converged through offline training are deployed to the distribution network controller, the distribution network status is collected in real time, the phase control actions of the energy storage system are generated through the policy network, and hard constraints such as the state of charge are introduced to verify the actions in real time. Finally, the power command is sent to the energy storage converter, and the three-phase imbalance of the distribution network is managed in real time by independently processing active and reactive power of each phase.

[0045] Furthermore, the method in this embodiment deeply couples the distributed risk perception capability of the DSAC algorithm with the phased power transfer capability of energy storage, and adaptively adjusts the risk preference coefficient through the operating condition severity index, thus solving the problems of computational complexity of traditional optimization methods, fixed risk preference of conventional reinforcement learning methods, and insufficient safety and adaptability under complex operating conditions.

[0046] The invention will now be further explained with reference to the accompanying drawings.

[0047] Figure 1 This is a flowchart of a three-phase imbalance mitigation method for energy storage-based phase-separated control distribution networks, based on the operating condition adaptive DSAC algorithm, provided by this invention. This method integrates an energy storage phase-separated decoupling model, Markov decision process construction, and the DSAC algorithm into a risk-aware decision framework. Through the updating and training of the policy network and evaluation network, and the adjustment of operating condition-adaptive risk preferences, it achieves rapid and robust mitigation of three-phase imbalance under random source-load fluctuations in high-penetration photovoltaic distribution networks. The method steps are as follows: S1: Construct a mathematical model for the three-phase imbalance management of a distribution network that includes distributed photovoltaic and energy storage systems, including: distribution network safety operation constraints and energy storage system physical constraints; In this invention, a power flow calculation model for a three-phase four-wire distribution network is established, and the output fluctuation of the photovoltaic inverter is treated as a random state disturbance. The energy storage system is modeled as a three-phase independent adjustable active and reactive power source with phases A, B, and C.

[0048] Figure 2 This invention provides a physical topology diagram of a phase-controlled energy storage system. The system is connected to the distribution network using a three-phase four-wire connection. Each of the three phases (A, B, and C) has an independent power control channel, allowing each phase to output active and reactive power. The neutral line (N line) provides a path for unbalanced current, thereby compensating for the zero-sequence current in the distribution network and fundamentally addressing the three-phase imbalance.

[0049] The energy storage model is expressed as: ;

[0050] In the formula, This indicates independent phase control. for t time Phase charge state; For charge and discharge efficiency; For energy storage active power; This refers to the rated capacity of the energy storage.

[0051] In this invention, the formulas for the safety operation constraints of the power distribution network and the physical constraints of the energy storage system are divided into the following two categories: (1) Node voltage amplitude constraints ; (2) State of charge constraints and capacity constraints of energy storage systems

[0052] In the formula, the upper limit of the node voltage is... Set to 1.05, lower limit Set to 0.95. This indicates a distribution network node; the maximum and minimum allowable SOC values ​​for energy storage are 0.8 and 0.2 respectively, to prevent overcharging and over-discharging of energy storage caused by SOC exceeding the threshold range; The SOC value of the stored energy at the beginning of the day. This refers to the SOC value of the energy storage at the end of the day. During the daily training process, the energy storage daily clearing principle must be met, that is, the SOC value at the beginning of the training is equal to the value at the end, and the control range of the active and reactive power of the energy storage must meet its capacity constraints.

[0053] S2: The mathematical transformation of the three-phase imbalance problem in high-penetration photovoltaic distribution networks into a Markov decision process (MDP), which includes a state space, action space, state transition probability function, and reward function. Figure 3 This is the interactive principle diagram of the Markov decision process provided by the present invention, at each time step. t Inside, the agent observes the current state. Then according to Make an action Finally, a reward value is obtained. And the state at the next moment is obtained based on the state transition probability. The state space Included t Real-time data on the three-phase voltage amplitude at each node of the distribution network, active / reactive power of the load, output of distributed photovoltaic systems, and state of charge of energy storage systems are used. Representation; Action Space Defined as int Active power commands of the energy storage system on phases A, B, and C at any time and reactive power command ,use It means that, among them, m This indicates the number of energy storage systems. Reward function. Taking into account node voltage amplitude, three-phase voltage imbalance, power loss, and energy storage SOC constraints, in t The reward function at time step is: ; In the formula, These represent the penalties for exceeding the limit of the node voltage amplitude. Three-phase voltage imbalance penalty Power loss penalty and SOC limit exceedance The weighting coefficients.

[0054] S3: A DSAC algorithm with adaptive risk preference based on operating conditions is proposed and applied to Markov decision process to train the agent. Data is extracted from the experience replay buffer to update the policy network (Actor) and evaluation network (Critic) for offline training of the agent. Among them, the operating condition severity index is constructed according to the current operating conditions of the distribution network, and the risk preference coefficient in the policy update process is adaptively adjusted using the operating condition severity index. In this embodiment, according to the next state The next action is obtained by sampling from the policy network. Then, the state-action pair The input is fed into two target evaluation networks, and the Gaussian parameters, i.e., the mean, of their corresponding reward distributions are obtained. with standard deviation , The corresponding representation is: Since the expected value of a Gaussian distribution equals its mean, the mean of the output distribution of the target evaluation network is used as the expected Q value. The smaller of the two target evaluation network output means is selected as the target value estimate to reduce the risk of overestimating the Q value due to function approximation error. Therefore, the final Bellman target value is expressed as: ; in, As a reward, As a discount factor, For the next state, Next move, The mean value of the network output is used as the target evaluation value, which achieves smoothing of photovoltaic power output fluctuation noise. α This is the temperature coefficient.

[0055] Figure 4 This is a schematic diagram of the policy network and evaluation network structure based on the working condition adaptive DSAC algorithm provided in this embodiment; the evaluation network (Critic) updates the network parameters by minimizing the agent's loss function, and the specific formula for calculating the loss function is as follows: ; Here, E[...] is the expectation operator, used to calculate the average value of a random variable.

[0056] In this embodiment, the operating condition severity index and the time-varying risk preference coefficient in step S33 are constructed as follows: ;

[0057] The update formula for the policy network (Actor) in step S34 is: ; ; Furthermore, in this embodiment, the training steps of the improved DSAC algorithm are as follows: (1) Input parameters: Evaluation network parameters θ 1 , θ 2 Policy network parameters Temperature coefficient Evaluation of network learning rate Temperature coefficient learning rate Target network soft update coefficient Weighting coefficients of operating condition indicators Basic risk appetite coefficient Adjusting the gain and the upper and lower limits of the time-varying risk preference coefficient ; (2) Initialize the target network parameters: ; (3) For each iteration and each sampling step: calculate the severity index of the operating condition and the corresponding risk preference coefficient based on the current operating status of the distribution network, and then apply the results to the current strategy network. Generate actions ; (4) The action a Applying to the power distribution network environment to obtain rewards and benefits r and the state at the next moment ; (5) Store in experience replay buffer pool D; (6) For each update step: randomly sample a batch of data from the experience replay buffer pool D; (7) Update the evaluation network, update the policy network, update the temperature coefficient, and soft update the objective function; S4: The trained agent is deployed to the distribution network for online optimization. Based on the optimal action strategy, the energy storage is scheduled to achieve three-phase imbalance management and verify its effectiveness.

[0058] Figure 5 This embodiment presents a model diagram of an energy storage system integrated into an IEEE 33-node distribution network. The effectiveness of the proposed method was verified using an IEEE 33-node test system. During the simulation phase, four 150kW / 300kWh energy storage units were installed at nodes 18, 22, 25, and 33, respectively. Photovoltaic and load data from a prefecture-level city over half a month were selected, and the data was divided into a training set and a test set. Fifteen days of data were used as the training set, and after training, typical daily data was selected as the test set.

[0059] Figure 6 This diagram shows the convergence of the average reward value per round during the training of the DSAC algorithm provided in this invention. During agent training, the learning rate was set to 0.0005, the discount factor to 0.99, the maximum number of training rounds to 400, the number of neurons in the two-layer neural network to 256, the experience replay pool size to 10,000, and the soft update coefficient to 0.01. In the early stages of training, the reward value was low and fluctuated significantly. This was because the agent was in an unknown environment without prior knowledge. Without interaction experience, the agent's action network was not fully optimized, and the resulting control strategy could not guarantee a good mitigation effect on the node voltage and three-phase imbalance of the distribution network. As the number of training rounds gradually increased to 80, the agent gradually adapted to the complex and changing distribution network environment, the action network continuously optimized, the voltage control and three-phase imbalance mitigation effects improved, and the reward value gradually converged.

[0060] Figure 7The comparison chart of the maximum voltage imbalance of the system before and after optimization provided in this embodiment shows the trend of the maximum voltage imbalance of the system within one day during simulation testing. The black curve represents the result before optimization, the blue curve represents the result after optimization using the SAC algorithm, and the red curve represents the result after optimization using the method proposed in this invention. As can be seen from the figure, the maximum voltage imbalance of the system before optimization was generally at a high level, basically maintained above 2.5%; after optimization using the SAC algorithm, the maximum voltage imbalance of the system decreased significantly, but still approached or even exceeded 2% at some times; while after optimization using the method of this invention, the maximum voltage imbalance of the system was further reduced, generally controlled below 2%, and significantly lower than that of the SAC algorithm at most times. This indicates that the method proposed in this embodiment can more effectively suppress the degree of three-phase imbalance in the distribution network, with better optimization effect and better operational stability, demonstrating the effectiveness and superiority of this intelligent agent in the operation control of three-phase unbalanced distribution networks.

[0061] As can be seen from the above technical solution, the beneficial effects of this embodiment are as follows: First, by using an energy storage system that supports independent phase-by-phase control, the three-phase active and reactive power are input as continuously adjustable parameters into the intelligent agent, enabling independent and rapid adjustment of voltage deviation and imbalance of each phase. Compared with the traditional unified adjustment method, it has higher control accuracy and stronger compensation capability, effectively reducing neutral point current, voltage imbalance rate, and phase-to-phase voltage deviation, thus improving power quality. Second, unlike traditional optimization methods based on power flow equations, mathematical programming, or sensitivity analysis, the DSAC scheduling strategy of this invention is obtained through interactive learning between the intelligent agent and the power grid environment. It does not rely on precise distribution network modeling and can automatically converge to the optimal adjustment strategy under complex operating conditions such as random photovoltaic output, severe load fluctuations, and continuous changes in power grid operating status, significantly improving the engineering applicability of the method. Third, this invention does not directly adopt the general DSAC algorithm with a fixed risk preference. Instead, it constructs a condition severity index based on the degree of photovoltaic fluctuation, the severity of three-phase imbalance, and the safety margin of energy storage SOC. Based on this, the risk preference coefficient is adaptively adjusted, which improves the governance efficiency when the operating conditions are stable and improves the operational safety and control stability when the operating conditions deteriorate. This significantly enhances the adaptability of the DSAC algorithm to the three-phase imbalance governance scenario of high-penetration photovoltaic distribution networks. Fourth, the scheduling strategy obtained from the training can be directly deployed to the actual distribution network for online operation, realizing real-time tracking and rapid response of the energy storage system to the three-phase operating status. It does not require continuous solving of complex optimization problems, has a fast response speed, and can significantly improve the online operation control efficiency of the distribution network, improve the steady-state level of three-phase voltage, and enhance the reliability of power supply.

[0062] On the other hand, this embodiment also provides a three-phase imbalance management system for energy storage distribution networks based on the DSAC algorithm, the system comprising: The model building module is used to construct a mathematical model for managing three-phase imbalance in a distribution network that includes distributed photovoltaic (PV) systems and energy storage systems. It utilizes the three-phase voltage amplitudes at each node of the distribution network, the active power and reactive power of the load, the output of the distributed PV systems, and the state of charge of the energy storage system at a certain moment in the mathematical model to construct the state space of a Markov decision process. It also utilizes the active and reactive power commands of the energy storage system at a certain moment in the mathematical model to construct the action space of the Markov decision process. Finally, it constructs the reward function of the Markov decision process based on the three-phase voltage amplitudes, three-phase voltage imbalance, power loss, and the state of charge constraints of the energy storage system at each node in the model. The agent training module is used to improve the value distribution reinforcement learning algorithm and apply it to the process of training the agent in the Markov decision process. The improved value distribution reinforcement learning algorithm includes: constructing a condition severity index based on the current distribution network operating conditions, and using the condition severity index to adaptively adjust the risk preference coefficient in the policy network update process. The online optimization module is used to deploy the trained agent to the distribution network for online optimization, and to schedule energy storage according to the optimal action strategy to achieve three-phase imbalance management.

[0063] The other technical features of the three-phase imbalance mitigation system for energy storage distribution networks based on the DSAC algorithm described in this embodiment are similar to those of the three-phase imbalance mitigation method for energy storage distribution networks based on the DSAC algorithm, and will not be repeated here.

[0064] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. "A plurality of" means two or more, unless otherwise explicitly specified.

[0065] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0066] In this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can mean that the first feature is in direct contact with the second feature, or that the first feature is in indirect contact with the second feature through an intermediate medium. Furthermore, "above," "over," and "on top" of the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.

[0067] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0068] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.

[0069] Therefore, this embodiment discloses a computer device.

[0070] The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the embodiments of the airport baggage system fault diagnosis intelligent agent construction method described above. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the embodiments of the above systems.

[0071] For example, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the airport baggage system fault diagnosis agent.

[0072] The computer device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device may include, but is not limited to, a processor and memory.

[0073] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0074] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0075] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0076] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0077] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for managing three-phase imbalance in energy storage distribution networks based on the DSAC algorithm, characterized in that, The method includes: A mathematical model for managing three-phase imbalance in a distribution network including distributed photovoltaic (PV) and energy storage systems is constructed. Using the three-phase voltage amplitudes, active and reactive power of the load, output of the distributed PV system, and the state of charge of the energy storage system at a certain moment in the mathematical model, a state space for a Markov decision process is constructed. The action space for the Markov decision process is constructed using the active and reactive power commands of the energy storage system at a certain moment in the mathematical model. The reward function for the Markov decision process is constructed based on the three-phase voltage amplitudes, three-phase voltage imbalance, power loss, and state of charge constraints of the energy storage system at the nodes in the model. An improved value distribution reinforcement learning algorithm is applied to the Markov decision process for training the agent. The improved value distribution reinforcement learning algorithm includes: constructing a condition severity index based on the current operating conditions of the distribution network, and using the condition severity index to adaptively adjust the risk preference coefficient in the policy network update process. The trained agent is deployed to the power distribution network for online optimization, and energy storage is scheduled according to the optimal action strategy to achieve three-phase imbalance management.

2. The method for managing three-phase imbalance in energy storage distribution networks based on the DSAC algorithm according to claim 1, characterized in that, The reward function for constructing the Markov decision process based on the three-phase voltage amplitude, three-phase voltage imbalance, power loss, and state of charge constraints of the energy storage system at the nodes in the model includes: exist t The reward function at time t is expressed as ; in, These represent the penalties for exceeding the limit of the node voltage amplitude. Three-phase voltage imbalance penalty Power loss penalty and the degree of over-limit of the state of charge of the energy storage system Weighting coefficients; The penalty for exceeding the limit of the node voltage amplitude Represented as: ; In the formula, N It is the collection of voltages at all three-phase nodes in the distribution network. Three-phase nodes i of The extent to which the phase voltage exceeds the upper and lower limits; express It belongs to phases A, B, and C, and ; in, Three-phase node i of Phase voltage amplitude.

3. The method for managing three-phase imbalance in energy storage distribution networks based on the DSAC algorithm according to claim 2, characterized in that, The three-phase voltage imbalance penalty Represented as: ; ; in, Indicates the negative sequence unbalance of three-phase voltage; These represent the zero-sequence, positive-sequence, and negative-sequence components of the three-phase voltage, respectively. These represent the voltage phasors of phases A, B, and C, respectively. and Indicates the rotation factor; The degree of state of charge exceeding the limit of the energy storage system Represented as: ; In the formula, Indicates the distribution network area number j Taiwan's energy storage system t time Phase state of charge, express t time Minimum constraint value for the charge state of a phase. express t time Maximum constraint value of phase charge state, m This indicates the number of energy storage systems in the distribution network.

4. The method for managing three-phase imbalance in energy storage distribution networks based on the DSAC algorithm according to claim 3, characterized in that, The improved value distribution reinforcement learning algorithm is applied to the process of training the agent using the Markov decision process, including: Two identical evaluation networks and one policy network are established; the evaluation networks are used to output the Gaussian distribution parameters of the state-action pair reward values, and the policy network is used to output the Gaussian distribution parameters of the actions. When calculating the Bellman target value, an expected value substitution mechanism is adopted, which uses the expected Q value output by the target evaluation network to replace the random sampling return, and selects the one with the smaller mean of the two target evaluation networks as the benchmark to eliminate random noise and suppress Q value overestimation. The severity index of the operating conditions is constructed based on the current operating conditions of the distribution network. The severity index of the operating conditions is composed of the intensity of photovoltaic power output fluctuation, the severity of three-phase voltage imbalance, and the deviation of the safety margin of the state of charge of the energy storage system. The time-varying risk preference coefficient is calculated based on the severity index of the operating conditions. By utilizing the distribution information from the evaluation network feedback and the time-varying risk preference coefficient, while maximizing the cumulative expected return and strategy entropy, the strategy network is updated through reparameterized sampling. This allows the network to adaptively adjust the degree of suppression of return volatility risk under different operating conditions, thus tending to select high-yield, low-risk control actions that match the current operating conditions.

5. The method for managing three-phase imbalance in energy storage distribution networks based on the DSAC algorithm according to claim 4, characterized in that, The degree of deviation of the state-of-charge safety margin of the energy storage system is expressed as follows: ; in, m This indicates the number of energy storage systems in the distribution network.

6. The method for managing three-phase imbalance in energy storage distribution networks based on the DSAC algorithm according to claim 5, characterized in that, The calculation of the time-varying risk preference coefficient based on the severity index of the working condition includes: The time-varying risk preference coefficient is expressed as follows: ; in, The function is a numerical clipping function. Basic risk preference coefficient; These are the upper and lower limits of the time-varying risk preference coefficient, respectively. To adjust the gain, This is an indicator of the severity of the operating condition; The severity index of the operating condition is expressed as: ; in, The weighting coefficients for operating condition indicators. Indicates the intensity of photovoltaic fluctuations. As a penalty for three-phase voltage imbalance, the photovoltaic fluctuation intensity is expressed as: ; in, Indicates from Time's up t Photovoltaic power output sequence at any given time The function represents a function that directly calculates the variance of an array.

7. The method for managing three-phase imbalance in energy storage distribution networks based on the DSAC algorithm according to claim 6, characterized in that, The method of utilizing the distribution information fed back by the evaluation network and the time-varying risk preference coefficient to maximize the cumulative expected return and policy entropy while updating the policy network through reparameterized sampling includes: The updated policy network is represented as: ; in, This represents the expectation operator, which is used to sample the state from the experience replay pool D. s and the policy network in that state Generated actions a The corresponding objective function is averaged. For the first i The mean of the output of the evaluation network for each target is expressed as: ; in, and They represent the first i The mean and standard deviation of the state-action-reward distribution output of the target evaluation network.

8. A three-phase imbalance mitigation system for energy storage distribution networks based on the DSAC algorithm, characterized in that, The system includes: The model building module is used to construct a mathematical model for managing three-phase imbalance in a distribution network that includes distributed photovoltaic (PV) systems and energy storage systems. It utilizes the three-phase voltage amplitudes at each node of the distribution network, the active power and reactive power of the load, the output of the distributed PV systems, and the state of charge of the energy storage system at a certain moment in the mathematical model to construct the state space of a Markov decision process. It also utilizes the active and reactive power commands of the energy storage system at a certain moment in the mathematical model to construct the action space of the Markov decision process. Finally, it constructs the reward function of the Markov decision process based on the three-phase voltage amplitudes, three-phase voltage imbalance, power loss, and the state of charge constraints of the energy storage system at each node in the model. The agent training module is used to improve the value distribution reinforcement learning algorithm and apply it to the process of training the agent in the Markov decision process. The improved value distribution reinforcement learning algorithm includes: constructing a condition severity index based on the current distribution network operating conditions, and using the condition severity index to adaptively adjust the risk preference coefficient in the policy network update process. The online optimization module is used to deploy the trained agent to the distribution network for online optimization, and to schedule energy storage according to the optimal action strategy to achieve three-phase imbalance management.

9. A computer device, characterized in that, include: One or more processors; The processor is used to store one or more programs; When the one or more programs are executed by the one or more processors, the three-phase imbalance management method for energy storage distribution networks based on the DSAC algorithm as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that, It contains a computer program, which, when executed, implements the three-phase imbalance management method for energy storage distribution networks based on the DSAC algorithm as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Energy storage charging and discharging strategy generation method and system based on reinforcement learning

    CN120033728A