A method and system for supporting the flexibility of distribution networks based on improved security reinforcement learning

By using a data-driven deep neural network evaluation module and an improved security reinforcement learning algorithm, combined with an expert correction module and an online adaptation module, the problem of ignoring the flexibility requirements of the main power grid and high-dimensional non-convex constraints in the main-distribution coordination strategy is solved. This enables flexibility assessment and dispatch instruction decomposition under intraday fluctuations of distributed resources, thereby improving the operational reliability and economy of the distribution network.

CN121440803BActive Publication Date: 2026-04-21TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
Filing Date
2025-12-26
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing main-distribution coordination strategies ignore the flexibility requirements of the main power grid, traditional models are difficult to handle high-dimensional non-convex and nonlinear constraints, and the generalization ability of security reinforcement learning methods based on day-ahead training decreases under the real-time fluctuation of distributed resources within the day.

Method used

A data-driven deep neural network evaluation module is used to map the flexibility range. An improved security reinforcement learning algorithm and an expert correction module are combined to decompose scheduling instructions. Parameters are fine-tuned through an online adaptation module to ensure decision-making performance is maintained under intraday fluctuations in distributed resources.

Benefits of technology

It provides efficient, safe, and economical support for distribution network flexibility assessment, improves the reliability and economy of power grid operation, and alleviates the problem of declining generalization capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121440803B_ABST
    Figure CN121440803B_ABST
Patent Text Reader

Abstract

This invention discloses a distribution network flexibility support method and system based on improved security reinforcement learning. The method includes: S1, mapping state vectors to flexibility intervals via a data-driven evaluation module: mapping state vectors containing distributed resource characteristics to the flexibility interval at the interface between the main grid and the distribution network via the data-driven evaluation module; S2, performing secure and economic decomposition of dispatching instructions via an intelligent decision-making module; S3, fine-tuning the policy network parameters of the intelligent decision-making module online based on the deviation between real-time environment and historical training data, and the importance ranking of policy network parameters via an online adaptation module. This achieves fast, secure, economical, and robust decomposition of main grid dispatching instructions under complex and uncertain environments, significantly improving the reliability, economy, and flexibility of grid operation under high-proportion renewable energy integration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system operation and control technology, specifically to a method and system for assessing and scheduling distribution network flexibility resources applicable to scenarios with a high proportion of distributed energy access. Background Technology

[0002] The high proportion of distributed resources access enhances the potential of the distribution network to meet the flexibility requirements of the main power grid, but also poses challenges to the safe and economical operation of the distribution network.

[0003] However, existing master-distribution coordination strategies have the following drawbacks: First, most of them treat the master grid as an infinitely large system, determining the injected power at interface nodes only after the distribution network's economic regulation ends, based on the distribution network's calculation results, ignoring the potential flexibility requirements of the master grid. Second, the randomness and volatility of massive heterogeneous distributed resources subject traditional model optimization methods to more high-dimensional non-convex and nonlinear constraints, increasing solution time and making it prone to getting trapped in local optima, thus hindering their application in large-scale distribution networks. Third, although safety reinforcement learning algorithms, such as those that impose safety limit penalties in the reward function, near-end policy optimization algorithms based on the master-dual method, constrained policy optimization methods, and safety layer projection, can effectively regulate agent decision-making and improve the efficiency of distribution network decision-making in complex environments, existing methods mostly use them for day-ahead training of agents in distribution networks or distributed resources. This means that during intraday deployment, if data in the environment differs significantly from the original training set (e.g., real-time intraday fluctuations in net load), the agent's decision-making performance (generalization ability) will decline, or even lose the ability to guarantee decision safety.

[0004] To address these issues, new strategies for supporting distribution network flexibility are needed, including a data-driven distribution network flexibility assessment method, a distribution network dispatch command security decomposition method based on improved security reinforcement learning, and an online fine-tuning method for intelligent agent neural networks to cope with intraday real-time fluctuations in distributed resources. Specifically, the distribution network flexibility assessment is based on a deep neural network (DNN), employing only 2-3 layers of neurons, and uses a modified linear unit (ReLU) activation function to enable the neural network to learn nonlinear mappings. The distribution network agent achieves secure decomposition of dispatch commands through an improved soft actor-critic algorithm. This algorithm not only uses safety constraints such as voltage, power flow, and generator output exceeding limits as components of the reward signal, but also embeds a model-driven expert module to perform safety correction on the agent's original action vector, ensuring strict satisfaction of constraints such as generator ramping. The online fine-tuning method determines the intraday fine-tuning target of the distribution network agent by comparing the real-time environment with all previous day training quintuples, thereby mitigating the problem of declining generalization ability.

[0005] Existing master-distribution coordination strategies have the following drawbacks: First, most of them treat the master grid as an infinitely large system, determining the injected power at interface nodes only based on the economic regulation results of the distribution network, ignoring the potential flexibility requirements of the master grid; Second, the introduction of massive heterogeneous distributed resources introduces more high-dimensional non-convex and nonlinear constraints, increasing the solution time of traditional model-driven methods and the possibility of getting trapped in local optima, which is not conducive to large-scale distribution network deployment; Third, most existing methods based on security reinforcement learning are based on day-ahead training, which is not conducive to the agent maintaining generalization ability under the intraday real-time fluctuations of distributed resources. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention proposes a distribution network flexibility support method based on improved security reinforcement learning. The core technical challenge this method addresses is: how to efficiently evaluate the flexibility support capability of a distribution network with a high proportion of distributed resource access; and, based on this, how to safely and economically decompose the main grid's dispatching instructions to each adjustable unit within the distribution network, while ensuring that the dispatching strategy maintains good generalization performance and security in the face of intraday real-time fluctuations in distributed resources.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] A distribution network flexibility support method based on improved security reinforcement learning, the method comprising:

[0009] S1. Mapping the state vector to a flexibility range through a data-driven evaluation module: The data-driven evaluation module maps the state vector containing distributed resource characteristics to the flexibility range at the interface between the main grid and the distribution network. The data-driven evaluation module includes a deep neural network model. S2. Performing safe and economic decomposition of dispatch instructions through an intelligent decision-making module: The intelligent decision-making module models the continuous decision-making process of the distribution system operator at different time periods as a Markov decision process. Based on a safe reinforcement learning algorithm integrating an expert correction module, it performs safe and economic decomposition of the dispatch instructions issued by the main grid. The expert correction module is used to perform rule-based safe correction on the original action vector output by the intelligent decision-making module. S3. Fine-tuning the strategy network parameters of the intelligent decision-making module online based on the deviation between the real-time environment and historical training data, and the importance ranking of the strategy network parameters, through an online adaptation module, to maintain its decision-making performance under daily real-time fluctuations in distributed resources.

[0010] In some embodiments, at least one of the following technical means is also included:

[0011] In step S1, the state vector is mapped to a flexibility range through a data-driven evaluation module, including: S11, constructing a training dataset, which is generated based on a traditional economic dispatch model and is used to characterize the mapping relationship between the distribution network state vector and the dispatchable range of the main-distribution interface; S12, fitting the mapping relationship using the deep neural network model, wherein the upper and lower limits of the adjustable range at the main-distribution interface are each independently fitted using a deep neural network model with the same structure.

[0012] The deep neural network model includes an input layer, at least one hidden layer, and an output layer, with data transfer between adjacent layers via a modified linear unit activation function.

[0013] In the deep neural network module, there are 3 hidden layers, and the number of neurons in each layer ranges from [128, 256, 256].

[0014] S2 involves a smart decision-making module that performs a safe and economical decomposition of the dispatching instructions, including: S21, defining the state vector of the Markov decision-making process, which includes dispatching instructions from the main grid, net active and reactive load vectors of each node in the distribution network, and vectors composed of the inherent flexibility ranges of each adjustable unit; S22, defining the action vector of the Markov decision-making process, which is composed of the dispatching coefficients of all adjustable units; S23, defining the reward function of the Markov decision-making process, which includes components for penalizing voltage exceedances, power flow exceedances, and instruction decomposition deviations, to guide decisions to meet the safety and economic requirements of the distribution network; S24, using the safe reinforcement learning algorithm to solve for the optimal strategy of the Markov decision-making process, and processing the original action vector by the expert correction module before outputting the final action vector.

[0015] The expert correction module processes the original action vector, including: S241, constructing an instruction decomposition sequence containing all adjustable units; S242, for the i-th adjustable unit in the decomposition sequence, calculating its externally required adjustable range, and then combining it with its inherent flexibility range to obtain the actual adjustable range, and determining its baseline scheduling instruction; S243, after determining the baseline scheduling instructions for all adjustable units, determining the scheduling instructions for the balancing node units through power flow calculation; S244, redistributing the over-limit amount according to the over-limit situation of the balancing node unit scheduling instructions and the upward or downward capacity ratio weight of all adjustable units in the sequence.

[0016] The construction rule for the decomposition sequence is as follows: electric vehicle charging stations are arranged first, followed by distributed units.

[0017] S3 involves online fine-tuning of the policy network parameters via an online adaptation module, including: S31, observing the real-time environment and obtaining initial action vectors and their value assessments based on the offline-trained policy network and value network; S32, calculating the offset between the real-time environment state and historical data based on the replay buffer of historical training; S33, if the offset exceeds a preset threshold, ranking the parameters of each layer in the network according to the gradient information of the action output by the policy network; S34, selecting parameters with importance lower than a predetermined standard as fine-tuning targets, calculating the gradient of the deviation between the real-time action, the value function, and the historical benchmark on the selected parameters, and updating the parameters accordingly.

[0018] In step S34, the 'historical baseline' is defined as the average action value Q(s_t_hist, a_t_hist) calculated by the value network for all historical states s_t_hist during the same time period in the offline training phase. The 'bias' is calculated using mean squared error (MSE). The goal of fine-tuning is to minimize the following loss function L: L = ω1 * MSE(current value network output, historical baseline value) + ω2 * MSE(current policy network output action, historical baseline action), where ω1 and ω2 are weighting coefficients balancing the two losses. By calculating the gradient of the loss function L with respect to the selected, less important policy network parameters, gradient descent (such as the Adam optimizer) is used to update the parameters with a very small learning rate (such as 1e-6), so that when the agent's decision drifts, it can be appropriately 'pulled back' to the vicinity of decisions that performed well in historically similar situations, rather than blindly adapting to potentially noisy real-time data.

[0019] This invention also includes the following technical solutions:

[0020] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, performs the steps of any of the methods described above.

[0021] Furthermore,

[0022] A computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements the steps of any of the methods described above.

[0023] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0024] This invention achieves efficient, safe, and economical support for distribution network flexibility through the synergistic effect of a data-driven evaluation module, an intelligent decision-making module integrating an expert correction module, and an online adaptation module. First, the data-driven evaluation module uses a deep neural network to accurately fit the nonlinear mapping from complex distributed resource states to the flexibility range of the main distribution network interface. This overcomes the shortcomings of traditional model-driven methods, such as difficulty in solving high-dimensional non-convex problems and susceptibility to local optima, providing fast and accurate flexibility boundary information for subsequent decision-making. Building on this, the intelligent decision-making module models the scheduling process as a Markov decision process and innovatively integrates a rule-based interpretable expert module into a safe reinforcement learning framework. The expert module performs safety correction on the agent's original actions, ensuring that the decomposition of scheduling instructions strictly meets physical constraints such as unit ramping, thus compensating for the shortcomings of pure data-driven methods in ensuring strict safety. Meanwhile, safe reinforcement learning guides the economic efficiency of decision-making through a reward function. The combination of these two approaches achieves the safe and economical decomposition of scheduling instructions. Furthermore, the online adaptation module monitors the deviation between real-time environment and historical data, and selectively fine-tunes the policy network based on parameter importance ranking. This enables the agent to adapt to intraday fluctuations in distributed resources, effectively mitigating the decline in generalization ability caused by environmental changes in traditional offline training models. Ultimately, these three modules work in a progressive and coordinated manner to achieve rapid, safe, economical, and robust decomposition of main grid dispatch commands under complex and uncertain environments, significantly improving the reliability, economy, and flexibility of grid operation with a high proportion of renewable energy integration.

[0025] Further beneficial effects of this invention include: employing a symmetric DNN structure with independent fitting upper and lower limits improves the accuracy of flexibility interval evaluation; providing clear goals and constraints for agent learning through clearly defined MDP states, action spaces, and reward functions; constructing specific decomposition sequences (such as EVCS first) in the expert correction module optimizes the efficiency and effectiveness of instruction decomposition; and implementing a gradient-based importance ranking fine-tuning strategy achieves precise adaptation to real-time fluctuations, avoids over-tuning, and ensures computational efficiency. Attached Figure Description

[0026] Figure 1 This is a flowchart of an embodiment of the present invention for improving the security reinforcement learning method to support the flexibility of distribution networks.

[0027] Figure 2 This is an overall flowchart of a distribution network flexibility support method based on improved security reinforcement learning, provided as an embodiment of the present invention.

[0028] Figure 3A , Figure 3B , Figure 3C , Figure 3DThis is a partial simulation result of the improved security reinforcement learning method used in this invention (improved IEEE 33-node system).

[0029] Figure 4A , Figure 4B , Figure 4C , Figure 4D The comparison results of the online fine-tuning module before and after the introduction of the present invention under different load fluctuation scenarios (improved IEEE 33-node system).

[0030] Figure 5A , Figure 5B , Figure 5C , Figure 5D This is a partial simulation result of the improved security reinforcement learning method used in this invention (improved IEEE 141-node system).

[0031] Figure 6A , Figure 6B , Figure 6C , Figure 6D The comparison results of the online fine-tuning module before and after the introduction of the present invention under different load fluctuation scenarios (improved IEEE 141-node system). Detailed Implementation

[0032] The embodiments of the present invention will be described in detail below. It should be emphasized that the following description is merely exemplary and not intended to limit the scope and application of the present invention.

[0033] It should be noted that when a component is referred to as "fixed to" or "set on" another component, it can be directly on or indirectly on that other component. When a component is referred to as "connected to" another component, it can be directly connected to or indirectly connected to that other component. Furthermore, a connection can be used for both fixing and circuit / signal connectivity.

[0034] It should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", and "outer" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing the embodiments of the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the present invention.

[0035] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of the present invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0036] The core of this invention lies in solving the flexibility support problem of high-proportion distributed energy distribution networks through the synergy of three core modules. First, a data-driven deep neural network is used to replace the traditional optimization model to quickly assess the flexibility boundary of the distribution network. Second, a physical rule-based expert module is embedded in the reinforcement learning agent to ensure the safety of dispatch command decomposition. Finally, an online fine-tuning mechanism is designed to enable the agent to adapt to real-time fluctuations within the day and maintain decision-making performance.

[0037] The basic concept of this invention is as follows:

[0038] This invention proposes a distribution network flexibility support strategy based on improved security reinforcement learning. First, a deep neural network (DNN) is used to assess distribution network flexibility, mapping the state vectors, which incorporate the complex characteristics of distributed resources, to the flexibility range of the main-distribution interface. Then, the Distributed System Operator (DSO) issues scheduling instructions to adjustable units for the current period based on factors such as wind turbine (WT), photovoltaic (PV) unit output, real-time load at each node, and real-time flexibility range of each adjustable unit. The continuous decision-making process of these instructions across different time periods is modeled as a Markov Decision Process (MDP). A security reinforcement learning algorithm integrating improved expert modules is used to achieve the safe and economical decomposition of main grid scheduling instructions. During intraday deployment, a gradient-based real-time action and value function deviation measurement module are employed to fine-tune the agent policy network online based on parameter importance ranking. The process is as follows: Figure 1 , Figure 2 As shown.

[0039] Example 1

[0040] like Figure 2 As shown, this technology mainly includes two parts: training and testing of DSO agents, which correspond to the flowcharts on the left and right sides of the figure, respectively.

[0041] First, DSO day-ahead training is performed: In each time period, the trained DNN first propagates forward based on the current state of distributed resources in the distribution network (including the net load of all nodes and the real-time flexibility range of all adjustable units) to obtain the real-time adjustable range at the main-distribution network interface, which serves as a reference for the main grid to issue dispatch instructions. After receiving the main grid instructions, DSO issues dispatch instructions for each adjustable unit in the current time period based on the improved SAC algorithm framework (with an expert module embedded at the end to realize the safety correction of the agent's actions). Subsequently, the state of the entire network can be obtained from power flow calculation, and the flexibility range of each adjustable unit in the next cycle can be recursively derived. After training for nearly 5000 rounds, the DSO agent converges.

[0042] Subsequently, the trained DSO agent is deployed intraday. At each time period, the DSO observes real-time environmental characteristics and makes preliminary decisions based on the policy network parameters trained the previous day. To balance generalization ability and ease of operation, the policy network is first sorted according to the gradient norm of each fully connected layer of the current initial action, with those having smaller norms being selected for subsequent online fine-tuning (a smaller gradient norm means that iterating them at the same time interval allows for more subtle output adjustments). Then, backpropagation is performed based on the difference between the current action, the action value, and the average value of the corresponding time period during previous training, enabling online fine-tuning of the policy network parameters.

[0043] The flowchart involves three key modules: DNN-based flexibility assessment, scheduling instruction decomposition based on an improved safety reinforcement learning algorithm, and online fine-tuning of the DSO agent. Their specific implementation steps are described below:

[0044] (1) Flexibility assessment of distribution network

[0045] The power distribution network contains various distributed resources, such as wind power, photovoltaic power, electric vehicles, user loads, and distributed generators (DG). Electric vehicles, through the orderly charging and discharging control of EVCS (Electric Vehicle Charging Stations), can provide regulation capabilities at distribution network nodes. DG can also provide a certain level of regulation capability under the condition of satisfying ramping constraints. Both are defined as adjustable units in this invention. Without considering external influences, the active power adjustable range of each adjustable unit in a certain time period is called the "inherent flexibility range." However, to meet system safety requirements, the DSO also requires the output of the adjustable unit to fall within a certain range, i.e., the "externally required adjustable range." The intersection of the "inherent range" and the "externally required range" constitutes the flexibility range of each adjustable unit. It should be noted that in the processing flow of the expert correction module (especially corresponding to step S242), the proportional allocation method used to calculate the "externally required adjustable range" is a rule-based heuristic algorithm. Its main advantages lie in its high computational efficiency, clear logic, and strict guarantee that the instructions of each unit do not exceed their inherent physical constraints, thus ensuring the safety of the correction process. Although this allocation method may not be globally optimal in terms of economy, as a safe correction step after the output of the reinforcement learning agent, its core objective is to compensate for the shortcomings of data-driven methods in terms of security, and this objective is fully achieved. Further optimization of economy is mainly achieved by the front-end reinforcement learning agent through learning the reward function.

[0046] This patent generates a dataset mapping distribution network state vectors to the schedulable intervals of the main-distribution interface based on the following traditional economic dispatch model:

[0047]

[0048]

[0049]

[0050]

[0051]

[0052]

[0053]

[0054]

[0055]

[0056]

[0057]

[0058]

[0059]

[0060]

[0061]

[0062]

[0063] in: These are indexes for time periods, distributed units, electric vehicle charging stations (EVCS), nodes, and branches. These are sets of nodes, branches, DGs, and EVCSs, respectively. It is the sum of the costs of all DG; It is the cost or benefit coefficient of power interaction between the main power grid and the distribution network during time period t, when the interaction power When the time is positive, the value is negative; otherwise, the value is positive. These are the coefficients of the quadratic and linear terms of the cost curve for the k-th DG unit; These are the active and reactive power outputs of the k-th generating unit during time period t; These are the lower and upper limits of its inherent flexibility range during time period t (referring to the adjustable output range of a distributed unit (such as DG, EVCS) without violating its own physical constraints (such as ramp rate, capacity limit), taking into account upward or downward ramp rate). Maximum active power Contributing to the previous period ; It is its minimum and maximum reactive power output; This is the active load of the m-th EVCS during time period t, and its inherent flexibility range upper and lower limits are respectively... ; These are the voltage and phase angle of node i during time period t, respectively. These represent the active and reactive power flows of branch (i,j), respectively. Its capacity limit, This is the control coefficient; For node i, the active and reactive power injections and loads; It is the output of wind turbines and solar power; It is a matrix that reflects network information.

[0064] After obtaining the dataset, the training and test sets were divided in a 4:1 ratio, and a DNN-based method was used to fit the mapping relationship. :

[0065]

[0066] in: It is a non-linear function derived from a DNN. It's important to note that for the upper and lower bounds of the adjustable range at the master-slave interface, a separate DNN is used for fitting, and each DNN has the exact same configuration: three fully connected layers, with data passed between each other via ReLU (Modified Linear Unit) activation; the input layer dimension is equal to a vector. The dimensions are (total number of nodes + number of EVCS × 2 + number of DG × 2), the number of neurons in the hidden layer are 128, 256, and 256 respectively, and the dimension of the output layer is 1. During training, a step size of 0.5 × 1e-4 is used, the number of samples per batch is 32, and the training takes 1000 rounds to converge.

[0067] (2) Security decomposition of distribution network dispatching instructions based on improved security reinforcement learning

[0068] This patent models the DSO scheduling process as an MDP, with the following variables.

[0069] 1) State vector: t The state vector for the time period is

[0070]

[0071] in It is a dispatch instruction from the main power grid; It is the net active and reactive load vector of each node in the distribution network. The net active load includes the output of WT and PV in advance. It is a vector composed of the inherent flexibility ranges of each EVCS in the distribution network;

[0072] 2) Action vector:

[0073]

[0074] in: It consists of the scheduling coefficients of all EVCS; It consists of the active scheduling coefficients of all DGs (except for the last DG in the expert module (expert module: a set of rule-based instruction decomposition logic designed in this invention to further ensure that instruction decomposition meets strict physical constraints and has interpretability after the initial decision of the reinforcement learning agent) and the DG connected to the balance node). It consists of the reactive power dispatch coefficients of all generating units. Each component of the action vector should be constrained to... Above. After obtaining the action vectors, the active and reactive power outputs of each schedulable unit, as well as the total active power commands to be decomposed, can be estimated using the following formulas:

[0075]

[0076]

[0077]

[0078]

[0079] 3) State transition: Transition probability Influenced by numerous factors, such as the electric vehicle status and the correlation of DSO scheduling instructions across different time periods, this is considered to be more realistic. Unknown. For details on the implementation of state transitions, please refer to the description in the expert module.

[0080] 4) Reward function: Used to guide the DSO to make decisions that meet the safety and economic requirements of the distribution network, and consists of multiple parts:

[0081]

[0082]

[0083]

[0084]

[0085]

[0086]

[0087]

[0088]

[0089]

[0090] in arrive It is the weighting coefficient (its value is positive). The voltage and power flow of nodes and branches are limited; among all DGs, one unit is set to be connected to the slack node, and its active power output is limited. And the rest The station is located in the decomposition sequence of the expert module, and the active power scheduling instruction of the last DG in the sequence. Use the previous one directly The excess power caused by the decomposition of the generator set is... .

[0091] 5) Optimal Strategy and Solution Method: This patent models the continuous decision-making process of DSO as an MDP, with the goal of finding the strategy π that maximizes the discounted return, i.e.:

[0092]

[0093] in It is a strategy that generates actions based on the current state vector. This corresponds to the trajectory. It is worth noting that this patent adds a model-driven expert module to the traditional Soft Actor-Critic (SAC) algorithm framework to perform safety correction on the original action vector of the DSO agent, ensuring that dispatch command decomposition is achieved while satisfying constraints such as unit ramping. Following the traditional SAC algorithm approach, it uses the exceedance limits of safety constraints such as voltage, power flow, and command decomposition deviation as components of the reward signal, achieving safe and economical DSO dispatch through penalties. Clearly, this method improves the SAC algorithm framework and falls within the scope of safety reinforcement learning algorithms. It is also applicable to other reinforcement learning algorithms. Any modifications or partial substitutions that do not depart from the spirit and scope of this invention should be covered by the claims of this invention.

[0094] DSO initially generates action vectors using MDP. Then, the output estimates of each schedulable unit and the total instructions to be decomposed are obtained. To ensure strict satisfaction of the ramping constraints of each DG and to guarantee the interpretability of the decomposition process, the following expert module is used to determine the DGs other than those located at the equilibrium node. The scheduling instructions for each schedulable unit. The pseudocode for the expert module algorithm is shown below:

[0095] Algorithm: Improved Expert Instruction Decomposition Module

[0096] Input: Inherent flexibility range of EVCS and DG Initial action vector of the agent Total power estimation required for scheduling Distribution network node and topology information.

[0097] Output: Actual scheduling instructions for EVCS and DG Node voltage, branch power flow, the last DG in the decomposed sequence, and the output limit of the DG located at the slack node.

[0098] 1. Construct a system with N cs EVCS and N DG - A decomposition sequence of one DG;

[0099] 2. For adjustable units

[0100] 3. Update remaining total power The inherent flexibility range of all adjustable units from the first adjustable unit onwards is accumulated. ;

[0101] 4. If

[0102] 5. Calculate the adjustment range required by external factors for adjustable unit i. ;

[0103] 6. Calculate the actual adjustment range of adjustable unit i. ;

[0104] 7. Determine the reference scheduling command for the i-th adjustable unit based on the actual adjustment range. ;

[0105] 8. end for

[0106] 9. First power flow calculation to obtain the over-limit value of DG scheduling instructions at the balancing node. ;

[0107] 10. If ,then

[0108] 11. for adjustable units do

[0109] 12. Calculate the proportion of the downward / upward adjustment capacity of adjustable unit i in the total capacity of all adjustable units. ;

[0110] 13. Determine the increment based on the proportion. The reference signal used to correct the adjustable unit i ;

[0111] 14. end for

[0112] 15.end if

[0113] 16. The second power flow calculation obtains the voltage overruns at each node of the entire network, the power flow overruns at each branch, the active power output overruns at the last DG in the sequence, and the DG scheduling command overruns at the balancing node. .

[0114] The following provides supplementary explanations of the core steps in the Improvement Expert module. After the specified information is entered, this module executes the following sequentially:

[0115] Step 1: Construct a structure containing EVCS, The decomposition sequence of Taiwan DG generally follows EVCS first, then DG, according to... arrive The instructions are decomposed in the order they appear.

[0116] Step 2: For the first Each schedulable unit, first in Before the deduction The scheduling instructions for the first unit have been determined, and the calculation starts from the first unit. The sum of the inherent flexibility range up to the last schedulable unit :

[0117]

[0118]

[0119] Step 3: Calculate the... The external requirements of each schedulable unit are adjustable within a certain range. This allows us to obtain the actual adjustable range and determine the baseline scheduling instruction. :

[0120]

[0121]

[0122]

[0123] Step 4: Baseline scheduling instructions for all schedulable units Once all parameters are determined, power flow calculations are performed to determine the DG scheduling instructions for the balancing node:

[0124]

[0125] Step 5: Based on the upward or downward limit exceedance situation of the DG scheduling instructions of the balance node, and the upward / downward capacity adjustment ratio weight of all schedulable units in the decomposition sequence. or To redistribute beyond the limits:

[0126]

[0127]

[0128]

[0129]

[0130] Step 6: After redistribution, the final power flow distribution and limit exceedance of the distribution network are completed. All can be determined. Substitute them into the reward function expression mentioned above to participate in the value assessment of DSO decision-making in this period.

[0131] Repeat the above steps to complete the day-ahead training of the DSO agent and save the information of the DSO agent (the core of which is the policy neural network that reflects the decision-making of the DSO agent, the value neural network that evaluates the quality of the initial actions of the DSO agent, and the MDP quintuple information for each of the 96 time periods of the day in the entire training cycle).

[0132] (3) Online fine-tuning method for agents based on gradient information and importance ranking

[0133] When a DSO agent trained offline is deployed for intraday use, the net node load may fluctuate beyond the original training set, leading to a decrease in the model's generalization ability. To address this issue, this patent proposes an online fine-tuning module based on gradient-based real-time action / value function bias measurement and parameter importance ranking. This module aims to fine-tune some parameters of the DSO policy network online based on offline training to adapt to real-time scenarios. The algorithm pseudocode is shown below:

[0134] Algorithm: Online Fine-tuning Strategy

[0135] Input: The policy network and value network of the DSO agent, which were previously trained offline. , 96 replay buffers The number of fully connected layers that need to be adjusted Offset threshold Other hyperparameters related to online fine-tuning (such as step size, number of training epochs, etc.).

[0136] Output: Improved DSO initial actions .

[0137] 1. Observe the state vector of the current time period from the real-time environment. .

[0138] 2. Input the original policy network ,get Input the value network to obtain the action value. .

[0139] 3. From the original policy network Select A fully connected layer of relatively minor importance.

[0140] 4. If There are real-time fluctuations, then

[0141] 5. Measure real-time offset or .

[0142] 6. then

[0143] 7. Calculate the loss item And assign weights .

[0144] 8. Calculate the loss item And assign weights .

[0145] 9. Calculate the loss item And assign weights .

[0146] 10.else

[0147] 11. Measure real-time offset .

[0148] 12. If ,then

[0149] 13. Calculate the loss item .

[0150] 14. Select the fully connected layer Fine-tune a certain number of rounds.

[0151] 15.end if

[0152] 16. Use (St) substitution This serves as the output of the policy network.

[0153] The core steps of the online fine-tuning strategy are explained in detail below.

[0154] Step 1: Observation Real-time environment of the time period And make initial actions based on the policy network trained offline. Value networks trained offline Value assessment ;

[0155] Step 2: If there are load fluctuations during this period, the real-time offset is measured based on the offline training replay buffer using the following formula:

[0156]

[0157]

[0158] Step 3: If the real-time offset exceeds a set threshold, the importance of the layers is ranked based on the absolute values ​​of the gradients of the output actions of the policy network. It's important to note that for a given fully connected layer, its importance metric is equal to the sum of the absolute values ​​of the gradients of all its parameters with respect to the output.

[0159]

[0160] Step 4: A higher importance index for a fully connected layer means that iterating the parameters of that layer once with the same time interval will have a significant impact on the overall network output. To achieve "fine-tuning" of the output, it is obvious that fully connected layers with lower importance should be selected as the iteration targets, and the offset should be calculated. , and Their gradients:

[0161]

[0162] in: This is the weighted sum of the three offsets above.

[0163] Experimental verification and results:

[0164] The method proposed in this invention is implemented in PyCharm 2024.1.7 software on the Windows 11 operating system. The simulation device is a laptop computer with a 32-core CPU, 32GB of memory, and a 1TB hard drive.

[0165] By inputting the distribution network node voltage array (including constraints), phase angle array (including constraints), branch impedance information array, and various distributed resource characteristic information into the software's input module, the program can automatically perform preprocessing operations such as cleaning, per-unitization, and time series normalization on the raw data, and then send the data to the data processing module according to the specified format. In the distribution network flexibility assessment module, a training set is constructed using the Gurobi solver to generate economic dispatch results for the distribution network under different conditions. A neural network is then used to fit the mapping relationship between the complex state vectors of each DERs and the adjustable range of the master-distribution interface, providing a reference for issuing flexibility dispatch instructions to the master network in practical application scenarios. In the DSO agent day-ahead training module, a safe reinforcement learning algorithm integrating and improving expert modules is used to safely and economically decompose the dispatch instructions issued by the master network to each dispatchable unit in the distribution network. In the DSO agent parameter online fine-tuning module, a gradient importance-based ranking method is used to select suitable fully connected layers in the DSO agent policy network for online fine-tuning to ensure the agent's generalization ability under real-time fluctuation environments. Finally, the final dispatch results are output through the output module.

[0166] In practical data scenarios based on improved IEEE 33-node and IEEE 141-node power distribution systems and the EVCS in Nanshan District, Shenzhen, the DSO agent trained with improved security reinforcement learning exhibits good scalability and reliability. Compared to other classic security reinforcement learning algorithms, such as PDPPO (Penalized Distributional Projection Policy Optimization), CPO (Constrained Policy Optimization), and security layer projection methods, it offers higher security assurance and cost-effectiveness: the average system operating cost is reduced by 14.29%, and the average overruns of voltage, power flow, and balancing node units over 96 time periods are only 0.01 pu, 0 pu, and 0.048 pu, respectively. Furthermore, in data-driven flexibility assessment methods, the DNN-based method and other classic fitting methods, such as Long Short Term Memory (LSTM) networks and LSTM networks fused with Convolutional Neural Networks (CNN), reduce the training time by an average of 74.57% and the fitting error by an average of 0.49%.

[0167] The software's performance on the improved IEEE 33-node system is shown in Table 1 and Figure 2. Figures 3A-3D Table 1, Table 2, Figure Figures 4A-4D As shown.

[0168] Figure 3A For the flexibility range of the TG-DS interface; Figure 3B This refers to the actual scheduling instructions for DG4; Figure 3C Voltage exceeded limit; Figure 3D For branch line tidal current exceeding the limit.

[0169] Figure 4A For increased load scenarios: no fine-tuning; Figure 4B For cases of increased load: minor adjustments are made; Figure 4C For load reduction scenarios: no fine-tuning; Figure 4D For load reduction scenarios: minor adjustments are made.

[0170] Table 1. Performance of different deep learning methods in distribution network flexibility assessment (improved IEEE 33-bus system)

[0171] Table 2 Performance comparison of different security reinforcement learning algorithms (improved IEEE 33-node system)

[0172] The results of the method proposed in this patent on the improved IEEE 141-node system are shown in Tables 3 and 4, and Figures 3 and 4. Figures 5A-5D , Figures 6A-6D As shown.

[0173] Figure 5A For the flexibility range of the TG-DS interface; Figure 5B This refers to the actual scheduling instructions for DG4; Figure 5C Voltage exceeded limit; Figure 5D For branch line tidal current exceeding the limit.

[0174] Figure 6A For increased load scenarios: no fine-tuning; Figure 6B For cases of increased load: minor adjustments are made; Figure 6C For load reduction scenarios: no fine-tuning; Figure 6D For load reduction scenarios: minor adjustments are made.

[0175] Table 3. Performance of different deep learning methods in distribution network flexibility assessment (improved IEEE 141-bus system)

[0176] Table 4 Performance comparison of different security reinforcement learning algorithms (improved IEEE 141-node system)

[0177] Example 2

[0178] This embodiment provides a distribution network flexibility support method based on improved security reinforcement learning. The method includes the following steps: mapping state vectors to flexibility ranges through a data-driven evaluation module, performing safe and economic decomposition of dispatch instructions through an intelligent decision-making module, and fine-tuning policy network parameters online through an online adaptation module. Specifically, the "data-driven evaluation module" in this embodiment is a deep neural network evaluation program built using Python and the TensorFlow framework. The "state vector" is specifically a multidimensional array containing the net active and reactive power loads of all nodes in the improved IEEE 33-node distribution system, the inherent flexibility range upper and lower limits of all four distributed generation units (DGs), and the inherent flexibility range upper and lower limits of all three electric vehicle charging stations (EVCSs). The "flexibility range" specifically refers to the adjustable range of active power at the connection interface between the main grid and the distribution network (usually the root node), for example, [-2.5MW, +3.0MW]. The "intelligent decision-making module" is specifically a distribution system operator (DSO) intelligent agent program developed based on an improved soft actor-critic (SAC) algorithm framework. The "Expert Correction Module" is a sub-function within the program that receives the agent's raw action output and executes a series of rule-based safety checks and corrections. The "Online Adaptation Module" is a separate online learning program that monitors real-time data and compares it with historical training data (replay buffer) stored on disk, triggering fine-tuning logic.

[0179] Implementation Environment and Data Preparation: This embodiment was implemented in the PyCharm 2024.1.7 development environment under the Windows 11 operating system. The distribution network model used in the simulation is an improved IEEE 33-node system. This system includes 33 nodes, 32 branches, 4 distributed generation units (DGs), and 3 electric vehicle charging stations (EVCSs). Historical and real-time data of all distributed resources (including loads, load waves, PVs, and EVCSs) and network parameters (impedance matrix) were read from the database and preprocessed, including data cleaning, per-unit normalization, and time series alignment.

[0180] The specific implementation of S1 (distribution network flexibility assessment): First, based on the traditional economic dispatch model and the Gurobi solver, a large number of samples under different operating scenarios are generated. The input (state vector s) of each sample includes the net load of each node, the inherent flexibility range of each DG and EVCS under that scenario, and the output (label) is the upper and lower limits of the adjustable active power range of the main-distribution interface obtained through optimization calculation under that scenario. A total of 100,000 valid samples are generated and divided into training and test sets in a 4:1 ratio. Subsequently, two deep neural networks (DNNs) with identical structures are constructed to fit the upper and lower limits, respectively. The input layer dimension of each DNN is: 33 (number of nodes) * 2 (active, reactive power) + 4 (number of DGs) * 2 (upper and lower limits) + 3 (number of EVCSs) * 2 (upper and lower limits) = 80. The hidden layers are set to 3 layers, with 128, 256, and 256 neurons respectively. The output layer dimension is 1. The ReLU activation function is used between adjacent layers. During training, the Adam optimizer was used with a learning rate of 0.00005 and a batch size of 32, for a total of 1000 epochs. After training, the network parameters were saved for use in stages S2 and S3.

[0181] Specific implementation of S2 (safe and economical decomposition of scheduling instructions): This step involves offline training of the DSO agent.

[0182] Define MDP:

[0183] State vector s_t: Specifically includes: dispatch instructions of the main power grid (1D) Active and reactive net load vectors of each node in the distribution network (Where WT and PV outputs are pre-combined; the dimension of each vector is the number of nodes in the distribution network), the vector is composed of the inherent flexibility range of each EVCS in the distribution network. (The dimension is twice the number of EVCS).

[0184] Action vector a_t: Composed of the scheduling coefficients of all adjustable units. Specifically, it includes: scheduling coefficients of 3 EVCSs (3-dimensional); active power scheduling coefficients of 3 DGs (excluding the balancing node DG) (3-dimensional); and reactive power scheduling coefficients of all 4 DGs (4-dimensional). Each component is constrained to the interval [0,1].

[0185] The reward function r has the following specific form: ,in , , , ,in arrive It is the weighting coefficient (its value is positive). The voltage and power flow of nodes and branches are limited. Among all distributed generation (DG) systems, one unit is connected to the slack node, and its active power output is limited. And the rest The station is located in the decomposition sequence of the expert module, and the active power scheduling instruction of the last DG in the sequence. Use the previous one directly The excess power caused by the decomposition of the generator set is... .

[0186] Agent training: The SAC algorithm with integrated expert modules is used for training. The agent's policy network and value network also adopt a DNN structure. At each time step, the agent outputs the raw action a_t_raw based on the current state s_t.

[0187] Expert correction: The original action a_t_raw is first input into the expert correction module for safety processing.

[0188] Step 1 (S241): Construct the decomposition sequence. According to the rule "EVCS first, DG last", the sequence is: [EVCS1,EVCS2, EVCS3, DG1, DG2, DG3] (assuming DG4 is the balancing node unit).

[0189] Step 2 (S242): Calculate the actual adjustable range and baseline scheduling command for each unit in sequence.

[0190] Step 3 (S243): After all unit instructions in the sequence are determined, power flow calculation is performed to obtain the output P_slack required by the slack node unit DG4.

[0191] Step 4 (S244): Check if P_slack exceeds the limit. If it does, redistribute the exceeded amount proportionally based on the remaining up / down capacity of each unit. The expert module outputs the corrected safety action a_t_safe.

[0192] Environment Interaction and Learning: Execute `a_t_safe`, calculate the reward `r_t`, and observe the next state `s_{t+1}`. Store the experience tuple `(s_t, a_t_safe, r_t, s_{t+1})` in the replay buffer. The agent samples data from the buffer to learn and update the network parameters. After training for approximately 5000 epochs, the system converges, and the agent parameters and historical training data are saved.

[0193] Specific implementation of S3 (online fine-tuning): The trained agent is deployed for daily operation. During each decision-making period t:

[0194] Step 1 (S31): Observe the real-time environment state s_t_real. Based on the offline-trained policy network, output the initial action a_t_real_init. Based on the offline-trained value network, evaluate the value of this action Q_t_real_init.

[0195] Step 2 (S32): Extract all state data {s_t_hist}, action data {a_t_hist}, and action value function {Q(s_t_hist, a_t_hist)} from the historical replay buffer for the same time period during offline training. Calculate the Euclidean distance between the real-time action, the real-time action value, and the historical values, and then take the average or minimum distance as the offset metric D.

[0196] Step 3 (S33): Set a threshold D_threshold (e.g., 0.1). If D > D_threshold, fine-tuning is triggered. Calculate the gradient of each layer parameter of the current policy network with respect to the output action a_t_real_init, and calculate the sum of the gradient norms of each layer parameter as the importance index of that layer.

[0197] Step 4 (S34): Select the K layers with the lowest importance index (e.g., the least important layer) as the fine-tuning target. Calculate the loss function L, such as the deviation between real-time actions and historical average actions, and the deviation between real-time action value and historical average action value. Calculate the gradient of L with respect to the parameters of the selected layer, and update these parameters with a very small learning rate (e.g., 1e-6).

[0198] Experimental verification:

[0199] Comparison objects:

[0200] Comparative Example 1 (Traditional Optimization): The safety layer projection method based on traditional quadratic programming is used for intraday decision-making in the distribution network.

[0201] Comparative Example 2 (SAC + Expert Module): The SAC algorithm with integrated expert correction module proposed in this invention is used.

[0202] Comparative Example 3 (PDPPO): The main-dual near-end strategy optimization algorithm is used for intraday decision-making in the distribution network.

[0203] Comparative Example 4 (CPO): Constrained strategy optimization algorithm is used for intraday decision-making in the distribution network.

[0204] Performance metrics:

[0205] Average system operating cost ($)

[0206] Average number of DG scheduling instructions exceeding the limit (pu) at the balanced node

[0207] Voltage over-limit rate (%) or average over-limit amount (pu)

[0208] Branch flow over-limit rate (%) or average over-limit amount (pu)

[0209] Experimental results:

[0210]

[0211] Example 3

[0212] The difference between this embodiment and Embodiment 2 is that its algorithm architecture is based on the classic near-end policy optimization framework. During the iteration process, the policy network parameters tend to search for the optimal result in the nearby small neighborhood as a reference for selecting the distribution network dispatch instructions for this period.

[0213] Example 4

[0214] The difference between this embodiment and embodiment 2 is that its algorithm architecture is based on the classic proximal policy optimization framework, but directly applies constraints to each Lagrangian function, rather than constructing a unified Lagrangian function to influence the training of the value network as in embodiment 3.

[0215] Example 5

[0216] This embodiment provides an electronic device. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it performs the steps of the method as described in any one of Embodiments 1 to 3. The electronic device may be a server, workstation, or dedicated computing device of a distribution system operator (DSO).

[0217] Example 6

[0218] This embodiment provides a computer-readable storage medium. The computer-readable storage medium (such as an SSD, USB flash drive, optical disc, etc.) stores a computer program (such as the Python program described in Embodiment 2). When the computer program is executed by a processor (such as the device processor in Embodiment 5), it implements the steps of the method described in any one of Embodiments 2 to 4.

[0219] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A distribution network flexibility support method based on improved security reinforcement learning, characterized in that, The method includes: S1. Mapping the state vector to a flexibility range through the data-driven evaluation module: The data-driven evaluation module maps the state vector containing distributed resource characteristics to the flexibility range at the interface between the main power grid and the distribution network. The data-driven evaluation module includes a deep neural network model. The flexibility range refers to the adjustable range of active power at the connection interface between the main power grid and the distribution network. S2. Safe and economic decomposition of dispatch instructions through intelligent decision-making module: The intelligent decision-making module of the distribution system operator intelligent agent program developed based on the improved soft actor-commentator algorithm framework models the continuous decision-making process of the distribution system operator in different time periods as a Markov decision process, and performs safe and economic decomposition of dispatch instructions issued by the main grid based on a safe reinforcement learning algorithm with an integrated expert correction module. The expert correction module is used to perform rule-based safe correction on the original action vector output by the intelligent decision-making module. S3. Through the online adaptation module, based on the deviation between the real-time environment and historical training data, and the importance ranking of the strategy network parameters, the strategy network parameters of the intelligent decision-making module are fine-tuned online to maintain its decision-making performance under the daily real-time fluctuations of distributed resources. S2 involves a smart decision-making module that performs a safe and economical decomposition of the dispatching instructions, including: S21, defining the state vector of the Markov decision-making process, which includes dispatching instructions from the main grid, net active and reactive load vectors of each node in the distribution network, and vectors composed of the inherent flexibility ranges of each adjustable unit; S22, defining the action vector of the Markov decision-making process, which is composed of the dispatching coefficients of all adjustable units; S23, defining the reward function of the Markov decision-making process, which includes sub-items for penalizing voltage exceedances, power flow exceedances, and instruction decomposition deviations, to guide decisions to meet the safety and economic requirements of the distribution network; and S24, using the safe reinforcement learning algorithm to solve for the optimal strategy of the Markov decision-making process, and processing the original action vector by the expert correction module before outputting the final action vector.

2. The method according to claim 1, characterized in that, In step S1, the state vector is mapped to a flexibility range through a data-driven evaluation module, including: S11, constructing a training dataset, which is generated based on a traditional economic dispatch model and is used to characterize the mapping relationship between the distribution network state vector and the schedulable range of the interface between the main power grid and the distribution network; S12, fitting the mapping relationship using the deep neural network model, wherein the upper and lower limits of the adjustable range at the interface between the main power grid and the distribution network are each independently fitted using a deep neural network model with the same structure.

3. The method according to claim 2, characterized in that, The deep neural network model includes an input layer, at least one hidden layer, and an output layer, with data transfer between adjacent layers via a modified linear unit activation function.

4. The method as described in claim 3, characterized in that, In the deep neural network model, there are 3 hidden layers, and the number of neurons in each layer ranges from [128, 256, 256].

5. The method according to claim 1, characterized in that, The expert correction module processes the original action vector, including: S241, constructing an instruction decomposition sequence containing all adjustable units; S242, for the i-th adjustable unit in the decomposition sequence, calculating its externally required adjustable range, and then combining it with its inherent flexibility range to obtain the actual adjustable range, and determining its baseline scheduling instruction; S243, after determining the baseline scheduling instructions for all adjustable units, determining the scheduling instructions for the balancing node units through power flow calculation; S244, redistributing the over-limit amount according to the over-limit situation of the balancing node unit scheduling instructions and the upward or downward capacity ratio weight of all adjustable units in the sequence; the inherent flexibility range refers to the active power adjustable range of each adjustable unit in a certain period of time without considering external influences.

6. The method as described in claim 5, characterized in that, The construction rule for the decomposition sequence is as follows: electric vehicle charging stations are arranged first, followed by distributed units.

7. The method according to claim 1, characterized in that, S3 involves online fine-tuning of the policy network parameters via an online adaptation module, including: S31, observing the real-time environment and obtaining initial action vectors and their value assessments based on the offline-trained policy network and value network; S32, calculating the offset between the real-time environment state and historical data based on the replay buffer of historical training; S33, if the offset exceeds a preset threshold, ranking the parameters of each layer in the network according to the gradient information of the action output by the policy network; S34, selecting parameters with importance lower than a predetermined standard as fine-tuning targets, calculating the gradient of the deviation between the real-time action, the value function, and the historical benchmark on the selected parameters, and updating the parameters accordingly.

8. A distribution network flexibility support system based on improved security reinforcement learning, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Distributed new energy, energy storage and power distribution network planning method considering flexible investment

    CN112217202A

  • Method for determining flexibility of alternating-current and direct-current hybrid power distribution network

    CN114336793A