An active distribution network optimization method and system

By constructing an energy storage lifetime degradation model and using deep reinforcement learning methods, the control strategies of photovoltaic inverters and energy storage systems are optimized, solving the problem of accelerated degradation of energy storage systems under frequent adjustments, and improving the stability and voltage regulation capability of active distribution networks.

CN122092402APending Publication Date: 2026-05-26STATE GRID JIANGSU ELECTRIC POWER CO LTD NANTONG POWER SUPPLY BRANCH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID JIANGSU ELECTRIC POWER CO LTD NANTONG POWER SUPPLY BRANCH
Filing Date
2026-04-24
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing voltage/reactive power optimization methods fail to effectively consider the lifespan degradation of energy storage systems, leading to accelerated degradation of energy storage devices during frequent adjustments, which affects the economy and lifespan of active distribution networks.

Method used

A degradation model for energy storage lifespan is constructed. Combined with deep reinforcement learning methods, the control strategy of photovoltaic inverter and energy storage system is optimized through the soft actor-critic SAC algorithm. The calendar aging and cyclic aging characteristics of the energy storage system are introduced to form a Markov decision process to optimize voltage and reactive power control.

Benefits of technology

It enables simultaneous management of voltage deviation and energy storage health, improves the operational reliability and stability of the active distribution network, extends the service life of energy storage devices, and enhances voltage regulation capability and control adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122092402A_ABST
    Figure CN122092402A_ABST
Patent Text Reader

Abstract

This application discloses an active distribution network optimization method and system that considers energy storage degradation and deep reinforcement learning in the field of active distribution network operation optimization technology. The optimization method first constructs an active distribution network model including a photovoltaic inverter and an energy storage system, and sets its active and reactive power regulation capabilities. Then, it introduces the state of charge, calendar aging, and cyclic aging of the energy storage system to establish a degradation model reflecting battery life loss. Based on this, the voltage and reactive power optimization process is represented as a reinforcement learning decision structure, achieving joint regulation of the photovoltaic inverter and energy storage system by setting states, actions, and rewards. The SAC algorithm is used to train the control strategy, forming a regulator that can be used for online operation. Finally, control commands are generated based on real-time measurements to achieve voltage stability and degradation suppression in the active distribution network. This application solves the technical problem of accelerated degradation of energy storage in traditional regulation methods, achieving a technical effect that balances voltage control objectives and energy storage life management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of active distribution network operation optimization technology, and is particularly applicable to an active distribution network optimization method and system based on an energy storage lifetime degradation model and deep reinforcement learning. Background Technology

[0002] With the continuous increase in distributed photovoltaic (PV), energy storage, and electric vehicle (EV) loads in power distribution systems, the operation of active distribution networks is increasingly exhibiting high volatility and multi-source coupling characteristics. Different types of distributed energy sources can cause frequent changes in node voltage within a short period, making it difficult for traditional voltage control methods relying on slow-speed regulation equipment to maintain voltage stability under dynamic conditions. To improve the system's adaptability to rapidly changing conditions, voltage / reactive power optimization technology based on inverter reactive power regulation is gradually becoming an indispensable means in the operation of active distribution networks.

[0003] Existing voltage / reactive power optimization methods mostly rely on linearized models or heuristic searches, which are limited in performance when facing complex power flow relationships and uncertainties in active distribution networks. Meanwhile, although deep reinforcement learning methods that have emerged in recent years can achieve policy learning in high-dimensional state spaces, they usually do not include calendar aging and cyclic aging of energy storage systems in the optimization scope, which may lead to accelerated degradation of energy storage during frequent adjustments, affecting economic efficiency and service life. Summary of the Invention

[0004] In view of this, and in response to the problems of the existing technology, this application provides an active distribution network optimization method based on an energy storage lifetime degradation model, which aims to solve the technical problem that energy storage in traditional regulation methods is prone to accelerated degradation.

[0005] To achieve the above objectives, this application provides an active distribution network optimization method based on an energy storage lifetime degradation model and deep reinforcement learning, comprising the following steps: Step 1: Construct an active distribution network optimization model to optimize voltage and reactive power, define the active and reactive power regulation capabilities of photovoltaic inverters and energy storage systems, and establish the operational relationship between node voltage and power injection; Step 2: Construct a degradation model for the energy storage system, introducing the state of charge (SOC), calendar aging amount (CAL), and cycle aging amount (CYC) to form degradation indices that characterize the loss of energy storage lifespan. Step 3: Construct the voltage and reactive power control process as a Markov decision process, define the state space consisting of node voltage, photovoltaic output and energy storage degradation state, define the action space consisting of energy storage and photovoltaic regulation, and set a reward function that includes voltage deviation and degradation. Step 4: Construct a control policy based on the soft actor-critic SAC reinforcement learning algorithm, and train the parameters through the update mechanism of the policy network, value network and target network to obtain a controller that can be used for online operation; Step 5: Deploy the trained control strategy to the active distribution network dispatching system, and output the active / reactive setpoints of the energy storage system and the reactive setpoints of the photovoltaic inverter based on real-time measurement data.

[0006] Furthermore, the active distribution network optimization model includes photovoltaic inverter constraints, which are implemented using the formula... Definition, where The active power output is connected to the photovoltaic inverter node. To provide reactive power for connection to photovoltaic inverter nodes, This refers to the rated apparent power of the photovoltaic inverter.

[0007] Furthermore, the active distribution network optimization model includes energy storage system constraints, which are expressed using the formula... Definition, where The active power of the energy storage system converter. The reactive power of the energy storage system converter. This represents the rated apparent power of the converter in the energy storage system.

[0008] Furthermore, the degradation index is characterized by cumulative calendar aging and cumulative cycle aging, wherein the cumulative calendar aging is expressed using the formula... Definition, where For energy storage units The battery from the initial moment to time The cumulative calendar aging level, For energy storage units The battery from the initial moment to time The cumulative calendar aging level, For energy storage units At time step The calendar aging amount, the cumulative cycle aging amount is calculated using the formula... Definition, where For the battery from the initial moment to time The cumulative cyclic aging level, For the battery from the initial moment to time The cumulative cyclic aging level, For energy storage units At time step The amount of cyclic aging.

[0009] Furthermore, the calendar aging amount is calculated using the formula... Definition, where For energy storage units At time step Calendar aging amount, For model scaling coefficients, For the influence coefficient of state of charge, Temperature influence coefficient, Let be the time exponent parameter. Indicates the battery's time temperature, For energy storage units in time The state of charge, This refers to the expected capacity loss of the battery at the end of its life.

[0010] Furthermore, the cyclic aging amount is calculated using the formula... Definition, where For energy storage units At time step The amount of cyclic aging; For energy storage units At time step The cycle life correction factor is calculated using the formula. Definition, where This is the maximum value of the correction factor. The amplitude parameters of the cycle life versus depth of discharge curves in the battery accelerated aging experiment are given. This represents the rate at which the cycle life correction factor decreases with increasing charge / discharge rate. For energy storage systems At time step The charge / discharge rate is determined using the formula. Definition, where For energy storage units At time step The charging and discharging power, For energy storage units Battery terminal voltage, For energy storage units Maximum capacity; For energy storage units At time step The number of cycles a battery can withstand, using the formula... Definition, where This represents the maximum cycle life. This refers to the maximum number of battery cycles. This is an empirical coefficient. For energy storage units At time step The depth of discharge is determined using the formula. Calculation, where For energy storage units In time The state of charge, For energy storage units In time The state of charge, using the formula Calculation, where For energy storage units The rated capacity of the battery, For energy storage units At time step The change in energy, using the formula Definition, where For energy storage units In time The charging and discharging power, For energy storage units The discharge efficiency, For energy storage units The charging efficiency.

[0011] Furthermore, the state space uses the system at time... state vector express, ,in For a moment node voltage amplitude, For a moment Photovoltaic inverter k The active output, For energy storage systems At the present moment The cumulative calendar aging amount, For energy storage systems At the present moment The cumulative cyclic aging amount; the action space uses the system at any time Output action vector express, ,in For a moment Energy storage unit The power output setting, For a moment Energy storage unit The setting of no reactive power, For a moment Photovoltaic inverter The reactive power regulation amount; the reward function uses Expression, in which For a moment The reward For a moment node Voltage deviation, The weighting parameters are used to reflect the voltage stability requirements. Weighted parameters that reflect the need for battery life protection.

[0012] Furthermore, the control strategy is solved through the following process: Establish the soft Bellman equation ,in To evaluate in a given state Next action The parameter for long-term expected return is The A commentator on the Q network, For temperature parameters, As a discount factor, For the first The parameters for training a commentator network, For the first Parameters of a target critic network, For trainable policy parameter set The determined actor strategy network; the optimization objective of the actor strategy network is... ,in The entropy of the policy distribution; the target commentator network update rules are used. ,in The coefficients are soft update coefficients; the actor strategy network uses reparameter sampling, and the formula uses... ,in, In the parameter set and random noise Below, regarding the state The generated follow-up action value, and They are respectively states The following is a set of trainable policy parameters. The mean and standard deviation of the network output actions determined by the policy. It is random noise.

[0013] Furthermore, the real-time measurement data mentioned in step 5 includes node voltage, photovoltaic output, and energy storage status.

[0014] Based on the same inventive concept, this application also provides an active distribution network optimization system for implementing the aforementioned active distribution network optimization method. The system includes a data acquisition module and a strategy execution module. The data acquisition module is used to acquire real-time measurement data including node voltage, photovoltaic output, and energy storage status. The strategy execution module stores a strategy network model trained using the soft actor-commentator SAC algorithm, and is used to output active / reactive power setpoints for energy storage and reactive power setpoints for the photovoltaic inverter based on the real-time measurement data acquired by the data acquisition module.

[0015] The beneficial effects of this application are as follows: The active distribution network optimization method based on the energy storage lifetime degradation model disclosed in this application is a voltage and reactive power optimization method that can simultaneously manage voltage deviation and degradation losses, taking into account both voltage quality and energy storage health. It solves the problem of insufficient regulation reliability caused by not considering energy storage degradation, and improves the operational reliability of the active distribution network. By combining the energy storage aging characteristics with the learning-based control strategy, the voltage deviation suppression and battery lifetime management can be achieved simultaneously within the same framework, improving the operational stability and control adaptability of the active distribution network. This method integrates photovoltaic inverters and energy storage systems into a unified voltage and reactive power optimization framework, which can improve the voltage regulation capability of active distribution networks. By introducing the calendar aging and cyclic aging characteristics of the energy storage system into the control model, the degradation rate of batteries can be effectively reduced and the service life of energy storage devices can be extended. A structured modeling method based on Markov decision processes is adopted, combined with a soft actor-commentator (SAC) strategy solving mechanism, so that the control strategy remains stable and adaptive under different operating conditions. The strategy is trained based on historical data, the control model has a clear structure, low parameter requirements, and can be implemented and deployed in active distribution networks. Attached Figure Description

[0016] To illustrate the objectives and technical solutions of this invention, the following figures are provided: Figure 1 This is a flowchart illustrating an embodiment of the method of this application. Detailed Implementation

[0017] To make the objectives and technical solutions of this application clearer, the application will be described in detail below with reference to the accompanying drawings and embodiments.

[0018] An active distribution network optimization method in this embodiment includes the following steps: Step 1: Construct an active distribution network voltage and reactive power optimization model including photovoltaic inverters and energy storage systems to describe the basic operating structure of the system. The active distribution network is based on the nodal admittance model, adopting the standard power flow relationship between nodal voltage and nodal power, and using the admittance matrix... Describe the electrical connection characteristics of the feeder. In this model, the first... The voltage magnitude and phase angle of each node are denoted as follows: and The node power is determined by both voltage and admittance parameters.

[0019] For nodes with integrated photovoltaic inverters and energy storage systems, their active and reactive power injections are jointly generated by loads and distributed generation. (The last sentence appears to be incomplete and possibly refers to a different topic.) Taking a node as an example, its net injected power can be expressed as an algebraic relationship between the load power and the output of the photovoltaic inverter and energy storage system, which is used to construct the node power balance constraint.

[0020] A photovoltaic inverter operates by limiting its active and reactive power output to its rated apparent power. Let the rated apparent power of the inverter be... Then it is meritorious. With no merit Constraints: The reactive power output is limited based on equipment configuration to support regional voltage regulation. The energy storage system is connected to the distribution network via a converter, and its active and reactive power outputs are also constrained by the rated apparent power. Let the rated apparent power of the energy storage system converter be... Then its output satisfies:

[0021] At the same time, a lower limit for charging power and an upper limit for discharging power are given for energy storage to limit its operating range during the optimization process.

[0022] To characterize the node voltage deviation, a reference voltage is selected. and define the first The voltage deviation at each node is:

[0023] This deviation is used to measure whether the node voltage meets the operating requirements. Using the active and reactive power outputs of the photovoltaic inverter and energy storage system as control variables, and under the conditions of node power balance relationship, inverter and energy storage system capacity constraints and node voltage deviation limits, a voltage / reactive power optimization model for the active distribution network is constructed, providing a unified physical basis for subsequent energy storage degradation modeling and reinforcement learning control.

[0024] Step 2: In this step, to describe the lifespan degradation mechanism of the energy storage system, two aging modes are introduced: calendar aging and cycle aging. Calendar aging of lithium-ion batteries is mainly caused by mechanisms such as electrolyte decomposition and the formation of passivation films on the metal surface. Its degradation rate is significantly affected by battery temperature and state of charge (SOC). High temperatures and high SOC levels typically accelerate battery capacity decay.

[0025] Based on existing battery life models, energy storage units At time step Calendar aging Defined as:

[0026] The entire formula represents the normalized calendar lifetime loss increment at a single time step, using the normalized lifetime loss share representation. In the formula: , is the model scaling coefficient, dimensionless, used to determine the baseline level of calendar aging; This is the temperature influence coefficient, in K, used to characterize the effect of temperature on the aging rate. The time power exponent parameter is dimensionless and is used to characterize the nonlinear relationship between aging and time. The state-of-charge (SOC) influence coefficient is dimensionless and used to characterize the effect of SOC on the aging rate. The above parameters were obtained through curve fitting based on aging experimental data of lithium-ion batteries under different SOC, temperature, and storage time conditions. The corresponding parameters in the relevant reference model are shown in this embodiment. 165400 can be obtained. 0.33 is acceptable. 4148K and 0.5 is acceptable. Indicates the battery's time The temperature, which is an absolute temperature, is measured in Kelvin (K). For energy storage units in time The state of charge, dimensionless. This represents the expected capacity loss of the battery at the end of its life, expressed as a percentage.

[0027] The battery SOC update relationship is as follows:

[0028] In the formula: For energy storage units In time The state of charge decreases during discharge and increases during charging. For energy storage units The rated capacity of the battery, For energy storage units At time step The change in energy, characterizing the physical energy output of a battery, is defined as:

[0029] In the formula: For energy storage units In time The charging and discharging power, with discharging being positive and charging being negative. and Energy storage units The discharge and charge efficiency.

[0030] The recursive formula for cumulative calendar aging is:

[0031] in For energy storage units The battery from the initial moment to time The cumulative calendar aging level, For the battery from the initial moment to time The cumulative calendar aging level.

[0032] During the charging and discharging process of an energy storage system, the battery undergoes multiple cycles at different charge rates and depths of discharge (DOD). These cycle conditions directly affect the battery's cycle life. To describe the impact of charge / discharge rate and depth of discharge on lifespan degradation, a cycle aging model is introduced. The battery's charge / discharge rates are as follows, used to characterize the energy storage system. At time step Operating intensity:

[0033] In the formula: For energy storage units At time step The charging and discharging power, For energy storage units Battery terminal voltage, For energy storage units Maximum capacity. Energy storage unit. At time step The depth of discharge is calculated from the change in SOC between two adjacent time points:

[0034] in For energy storage units In time The state of charge, For energy storage units In time State of charge Under different DOD conditions, the number of cycles the battery can withstand (CTF) can be calculated:

[0035] In the formula: For energy storage units At time step The number of cycles the battery can withstand. This represents the maximum cycle life; in this embodiment, it is set to 40000. This represents the maximum number of battery cycles; in this embodiment, the value is set to 2000. The rate at which cycle life decreases with increasing depth of discharge is represented by a value of 0.99 in this embodiment. All parameters are based on lithium-ion battery accelerated aging tests, obtaining lifespan loss data under different depths of discharge (DOD) and rates, and then establishing CTF and CLC models through fitting. To further describe the impact of rate capability on lifespan, an energy storage unit is introduced. At time step Cycle life correction factor The calculation method is as follows:

[0036] In the formula, This is the maximum value of the correction factor; in this embodiment, it is set to 4. In this embodiment, the value of 1.041 is used as the amplitude parameter for the cycle life versus depth of discharge curve in the battery accelerated aging experiment. The rate at which the cycle life correction factor decreases with increasing charge / discharge rate is 0.445 in this embodiment. This value is also based on parameters obtained from accelerated aging tests of lithium-ion batteries and is used to characterize the impact of rate changes on battery life. Combining the two life factors mentioned above, the cycle life loss of the battery under the current DOD and Crate conditions is equivalent to the energy storage unit's cycle life loss. At time step Cyclic aging amount Defined as:

[0037] The cumulative amount of cyclic aging is calculated using a recursive method:

[0038] in For the battery from the initial moment to time The cumulative cyclic aging level, For the battery from the initial moment to time The cumulative cyclic aging level, For energy storage units At time step The amount of cyclic aging.

[0039] To ensure that the lifespan loss of the energy storage system is within acceptable limits, this step further imposes constraints on the intraday cycle aging amount and the intraday calendar aging amount:

[0040] In the formula: and These represent the maximum allowable cycle aging and calendar aging of the battery within a day, respectively. In this embodiment, the day is calculated using 24 hours, and the proportion of the total cycle life (e.g., 2000 cycles) and calendar life (e.g., 10 years) to the allowable daily consumption is less than 1.

[0041] Step 3: In the voltage and reactive power optimization problem, to ensure the control process can obtain policy outputs across multiple time-series operation scenarios, the system is represented as a Markov decision process. At any given time... The system state is composed of grid operating parameters and energy storage degradation information. The controller generates corresponding action values ​​based on this state and, based on environmental feedback, transitions to the state at the next time step, while simultaneously receiving a reward value. The system at time... The state vector is represented as follows:

[0042] In the formula: Indicates time node voltage amplitude, For a moment Photovoltaic inverter The active output, and They represent energy storage systems. At the present moment The cumulative calendar aging and cumulative cycle aging are recorded. This status information together describes the distribution network voltage status, distributed generation output level, and energy storage system health status.

[0043] Controller at time Output action vector The active and reactive power setpoints of the energy storage system and the reactive power setpoint of the photovoltaic inverter are expressed as follows:

[0044] In the formula: and Corresponding time Energy storage unit The setting of active and passive effort output, For a moment Photovoltaic inverter The reactive power regulation is used to perform joint control of voltage and reactive power. To ensure that the control strategy can simultaneously suppress node voltage deviation and limit battery degradation during operation, this step constructs a reward function based on voltage deviation and degradation, expressed as follows:

[0045] In the formula: For a moment The reward For a moment node The voltage deviation is normalized to a per-unit value. and The corresponding weight parameters reflect the voltage stability requirements and battery life protection requirements, respectively. In this embodiment, both parameters are set to 1 to ensure that the relative importance of voltage deviation penalty and battery life loss penalty is consistent. Through this reward function, the control process can complete policy evaluation under multi-objective conditions and use it to guide subsequent policy training and updates.

[0046] Step 4: After modeling the state space, action space, and reward function, a reinforcement learning method based on soft actor-critic SAC is used to solve the policy in order to obtain a control policy that can adapt to voltage and reactive power regulation requirements. This method belongs to the off-policy algorithm with entropy regularization. It improves the exploration ability of action selection by adding an entropy term to the policy optimization, while using experience replay and the target network to ensure the stability of policy updates. Under this control framework, a stochastic policy... With two groups function , Value assessment and strategy optimization are achieved through alternating training. The value update satisfies the soft Bellman equation. The Markov decision process in step 3 is solved based on the soft actor-commentator SAC reinforcement learning structure. The control policy is formed through the updates of the value network, policy network, and target network. Both the actor and commentator use two fully connected hidden layers, each with 256 neurons. The activation function is a non-linear ReLU activation function. The learning rates for both the actor and commentator are set to a certain value. Its mathematical form is established in the following order.

[0047] Soft Bellman equation:

[0048] In the formula: To evaluate in a given state Next action The parameter for long-term expected return is The A value network of online, untrained critics. The temperature parameter is used to measure the influence of entropy on the strategy. In this embodiment, the initial value is 0.2, and an automatic adjustment mechanism is employed. The learning rate is also set to... , For trainable policy parameter set The network of actor strategies determined at any given moment When in a state The random strategy to be adopted at that time, which is based on As input, the actual action is obtained by analyzing the probability distribution in the randomly sampled action space. ; As a discount factor, calculate Minimum value in the value network of a target critic; and These correspond to the parameters of the training network and the target network, respectively. For a moment state Next action The reward. The function updates by minimizing the above objective, thereby estimating the value of the action in the current state.

[0049] The training objective of the policy network is to obtain a high expected reward at each time step while maintaining a certain entropy value to enhance action diversity. Its optimization objective is:

[0050] in The entropy of the policy distribution. The same temperature parameters as mentioned above, For the state The information entropy of the conditional probability distribution. In strategy Next moment From 0 to state and actions The cumulative expected value.

[0051] To improve training stability, the target The network uses an exponentially weighted approach for updates, with the following update rules:

[0052] In the formula: This is the soft update coefficient, used to control the update speed of the target network. In this embodiment, it is set to 0.005.

[0053] To achieve the differentiable sampling process of actions, the policy network generates continuous actions based on the reparameterization technique, and its sampling form is as follows:

[0054] In the formula: and They are respectively states The following is a set of trainable policy parameters. The mean and standard deviation of the network output actions determined by the policy. As random noise, the actions generated in this form maintain gradient availability during backpropagation, thus enabling efficient training of the actor-policy network. This step forms the complete soft actor-critic SAC control solution structure, used to obtain the final control policy under the joint optimization objective of voltage deviation and battery degradation.

[0055] During the control strategy training phase, an interactive environment is constructed based on the aforementioned active distribution network model and energy storage degradation model, enabling the controller to form a complete training loop through state input, action output, and reward feedback. The soft actor-critic (SAC) utilizes experience replay and target network update mechanisms for strategy optimization, gradually converging the control process to a stable strategy under the dual objectives of voltage deviation and energy storage degradation. In this embodiment, the experience replay buffer size is... The training consisted of 500 rounds, with daily scheduling based on hourly time resolution. Each round's episode contained 24 steps, resulting in a total of 12,000 interaction samples. Step 5: Deploy the trained policy network in the active distribution network operation control stage. Based on operational information such as node voltage, photovoltaic output, and energy storage status, output the active / reactive power setpoints for energy storage and the reactive power setpoints for the photovoltaic inverter to achieve active distribution network voltage regulation and energy storage degradation suppression. Training can be completed on a conventional GPU platform, while online execution only requires forward policy calculations, meeting the real-time requirements of actual operation.

[0056] Based on the same inventive concept, this application also provides an active distribution network optimization system for implementing the aforementioned active distribution network optimization method. The system includes a data acquisition module and a strategy execution module. The data acquisition module is used to acquire real-time measurement data including node voltage, photovoltaic output, and energy storage status. The strategy execution module stores a strategy network model trained using the soft actor-commentator SAC algorithm, and outputs active / reactive power setpoints for the energy storage system and reactive power setpoints for the photovoltaic inverter based on the real-time measurement data acquired by the data acquisition module.

[0057] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made to it in form and detail without departing from the scope defined by the claims of the present invention.

Claims

1. An active distribution network optimization method, characterized in that, Includes the following steps: Step 1: Construct an active distribution network optimization model to optimize voltage and reactive power, define the active and reactive power regulation capabilities of photovoltaic inverters and energy storage systems, and establish the operational relationship between node voltage and power injection; Step 2: Construct a degradation model for the energy storage system, introducing the state of charge (SOC), calendar aging amount (CAL), and cycle aging amount (CYC) to form degradation indices that characterize the loss of energy storage lifespan. Step 3: Construct the voltage and reactive power control process as a Markov decision process, define the state space consisting of node voltage, photovoltaic output and energy storage degradation state, define the action space consisting of energy storage and photovoltaic regulation, and set a reward function that includes voltage deviation and degradation. Step 4: Construct a control policy based on the soft actor-critic SAC reinforcement learning algorithm, and train the parameters through the update mechanism of the policy network, value network and target network to obtain a controller that can be used for online operation; Step 5: Deploy the trained control strategy to the active distribution network dispatching system, and output the active / reactive setpoints of the energy storage system and the reactive setpoints of the photovoltaic inverter based on real-time measurement data.

2. The active distribution network optimization method according to claim 1, characterized in that: The active distribution network optimization model includes photovoltaic inverter constraints, which are expressed using the formula... Definition, where The active power output is for connection to the photovoltaic inverter node. To provide reactive power for connection to photovoltaic inverter nodes, This refers to the rated apparent power of the photovoltaic inverter.

3. The active distribution network optimization method according to claim 2, characterized in that: The active distribution network optimization model includes energy storage system constraints, which are expressed using the formula... Definition, where The active power of the energy storage system converter. The reactive power of the energy storage system converter. This represents the rated apparent power of the converter in the energy storage system.

4. The active distribution network optimization method according to claim 3, characterized in that: The degradation index is characterized by cumulative calendar aging and cumulative cycle aging, wherein the cumulative calendar aging is expressed using the formula... Definition, where For energy storage units The battery from the initial moment to time The cumulative calendar aging level, For energy storage units The battery from the initial moment to time The cumulative calendar aging level, For energy storage units At time step The calendar aging amount, the cumulative cycle aging amount is calculated using the formula... Definition, where For the battery from the initial moment to time The cumulative cyclic aging level, For the battery from the initial moment to time The cumulative cyclic aging level, For energy storage units At time step The amount of cyclic aging.

5. The active distribution network optimization method according to claim 4, characterized in that: The calendar aging amount uses the formula Definition, where For energy storage units At time step Calendar aging amount, For model scaling coefficients, For the influence coefficient of state of charge, Temperature influence coefficient, For time exponent parameters, Indicates the battery's time temperature, For energy storage units in time The state of charge, This refers to the expected capacity loss of the battery at the end of its life.

6. The active distribution network optimization method according to claim 4, characterized in that: The cyclic aging amount uses the formula Definition, where For energy storage units At time step The amount of cyclic aging; For energy storage units At time step The cycle life correction factor is calculated using the formula. Definition, where This is the maximum value of the correction factor. The amplitude parameters of the cycle life versus depth of discharge curves in the battery accelerated aging experiment are given. This represents the rate at which the cycle life correction factor decreases with increasing charge / discharge rate. For energy storage systems At time step The charge / discharge rate is determined using the formula. Definition, where For energy storage units At time step The charging and discharging power, For energy storage units Battery terminal voltage, For energy storage units Maximum capacity; For energy storage units At time step The number of cycles a battery can withstand, using the formula... Definition, where This represents the maximum cycle life. This refers to the maximum number of battery cycles. This represents the rate at which cycle lifetime decays with increasing depth of discharge. For energy storage units At time step The depth of discharge is determined using the formula. Calculation, where For energy storage units In time The state of charge, For energy storage units In time The state of charge, using the formula Calculation, where For energy storage units The rated capacity of the battery, For energy storage units At time step The change in energy, using the formula Definition, where For energy storage units In time The charging and discharging power, For energy storage units The discharge efficiency, For energy storage units The charging efficiency.

7. The active distribution network optimization method according to claim 6, characterized in that: The state space uses the system at time... state vector express, ,in For a moment node voltage amplitude, For a moment Photovoltaic inverter The active output, For energy storage systems At the present moment The cumulative calendar aging amount, For energy storage systems At the present moment The cumulative cyclic aging amount; the action space uses the system at any time Output action vector express, ,in For a moment Energy storage unit The power output setting, For a moment Energy storage unit The setting of no reactive power, For a moment Photovoltaic inverter The reactive power regulation amount; the reward function uses Expression, in which For a moment The reward For a moment node Voltage deviation, The weighting parameters are used to reflect the voltage stability requirements. Weighted parameters that reflect the need for battery life protection.

8. The active distribution network optimization method according to claim 7, characterized in that, The control strategy is solved through the following process: establishing the soft Bellman equation. ,in To evaluate in a given state Next action The parameter for long-term expected return is The A commentator on the Q network, For temperature parameters, As a discount factor, For the first The parameters for training a commentator network, For the first Parameters of a target critic network, For trainable policy parameter set The determined actor strategy network; the optimization objective of the actor strategy network is... ,in The entropy of the policy distribution; the target commentator network update rules are used. ,in The coefficients are soft update coefficients; the actor strategy network uses reparameter sampling, and the formula uses... ,in, In the parameter set and random noise Below, regarding the state The generated follow-up action value, and They are respectively states The following is a set of trainable policy parameters. The mean and standard deviation of the network output actions determined by the policy. It is random noise.

9. The active distribution network optimization method according to claim 1, characterized in that: The real-time measurement data mentioned in step 5 includes node voltage, photovoltaic output, and energy storage status.

10. An active distribution network optimization system, characterized in that: The active distribution network optimization method according to any one of claims 1 to 9 includes a data acquisition module and a strategy execution module. The data acquisition module is used to acquire real-time measurement data including node voltage, photovoltaic output, and energy storage status. The strategy execution module stores a strategy network model trained by the soft actor-commentator SAC algorithm and is used to output the active / reactive setpoints of the energy storage system and the reactive setpoints of the photovoltaic inverter based on the real-time measurement data acquired by the data acquisition module.