Power distribution network dispatching control method and device, computer equipment, readable storage medium and program product

By constructing a multi-objective optimization model and using deep learning algorithms to optimize the active power output of energy storage devices and generator sets, the problems of high scheduling calculation costs and safety and stability caused by distributed power sources such as photovoltaics have been solved, and efficient, energy-saving and emission-reducing distribution network scheduling has been achieved.

CN121840647APending Publication Date: 2026-04-10SHENZHEN POWER SUPPLY BUREAU
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies, when faced with the uncertainty of output from distributed power sources such as photovoltaics and the uncertainty of node loads, result in high computational costs for distribution network scheduling, making it difficult to ensure safe and stable operation and achieve the goals of energy conservation and emission reduction.

Method used

A multi-objective optimization model is constructed, which combines a dual-delay deep deterministic policy gradient model and a mixed-integer programming model to optimize the active power output of energy storage devices and generator sets. The scheduling strategy is optimized through a trained deep learning algorithm to meet the objectives of minimum operating cost, minimum voltage deviation, and minimum carbon emissions.

Benefits of technology

It reduces the computational cost of distribution network dispatching, improves dispatching efficiency, and ensures the safe and stable operation of the distribution network, while also achieving the goals of energy conservation and emission reduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121840647A_ABST
    Figure CN121840647A_ABST
Patent Text Reader

Abstract

The invention relates to a power distribution network dispatching control method and device, computer equipment, a readable storage medium and a program product. The method comprises the steps that the current net power of a power distribution network node in the current time period and the current charge state of an energy storage device in the current time period are acquired; inputting the current net power, the current charge state and the current time period into a double-delay depth deterministic strategy gradient model, and outputting first candidate active power of an energy storage device and second candidate active power of a generator set; constructing a mixed integer programming model embedded with constraint conditions of a multi-target optimization model of the power distribution network, inputting the first candidate active power and the second candidate active power into the mixed integer programming model, and outputting first target active power and second target active power; and controlling the energy storage device and the generator set to output the first target active power and the second target active power to the power distribution network node respectively. The method provided by the invention can reduce the calculation cost of the dispatching process of the power distribution network.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power distribution network optimization scheduling, in particular to a power distribution network scheduling control method and device, computer equipment, readable storage medium and program product. BACKGROUND

[0002] With the continuous access of distributed power sources such as photovoltaic, the uncertainty and volatility of its output bring serious challenges to the scheduling of power distribution network, and due to the uncertainty of node load in the power distribution network and the possible fluctuation of voltage in the power distribution network, further affecting the safe and stable operation of the power distribution network. At the same time, the current goal is constantly advancing, energy saving and emission reduction has become the current development trend, which needs to reduce carbon emissions as much as possible under the premise of ensuring power supply. The above factors bring serious challenges to the scheduling of power distribution network.

[0003] To solve the above problems, the prior art models the power distribution network in simulation software to generate a simulation scenario of the power distribution network, thereby optimizing the scheduling scheme of the power distribution network based on the simulation results; however, with the increasing complexity of the power distribution network, the calculation cost of this method will significantly increase. SUMMARY

[0004] Therefore, it is necessary to provide a power distribution network scheduling control method, device, computer equipment, readable storage medium and program product capable of effectively reducing the calculation cost in view of the above technical problems.

[0005] In a first aspect, the present application provides a power distribution network scheduling control method, which comprises:

[0006] For a power distribution network comprising a photovoltaic device, an energy storage device and a generator set, a multi-objective optimization model of the power distribution network and a constraint condition of the multi-objective optimization model are constructed; wherein the multi-objective optimization model comprises a minimum operating cost function, a minimum voltage deviation function and a minimum carbon emission target function;

[0007] The current net power of any one power distribution network node in the current period is obtained, and the current state of charge of the energy storage device in the current period is obtained;

[0008] The current net power, the current state of charge and the current period are input into the trained double-delay deep deterministic policy gradient model, and the first candidate active power of the energy storage device and the second candidate active power of the generator set are output;

[0009] A mixed integer programming model embedded with the constraint condition is constructed, and the first candidate active power and the second candidate active power are input into the mixed integer programming model, and the first target active power and the second target active power are output;

[0010] controlling the energy storage device to output the first target active power to the power grid node and controlling the generator set to output the second target active power to the power grid node.

[0011] In one of the embodiments, the training process of the double-delay deep deterministic policy gradient model comprises:

[0012] For a current training round in the training process, sample data of the current training round is obtained, wherein the sample data comprises a first sample period, a first sample net power of any one power grid node in the first sample period, and a first sample state of charge of the energy storage device in the first sample period;

[0013] The sample data is input into the double-delay deep deterministic policy gradient model before training, and a first predicted active power of the energy storage device and a second predicted active power of the generator set are output;

[0014] In a second sample period after the first sample period, the energy storage device is controlled to output the first predicted active power to the power grid node, and the generator set is controlled to output the second predicted active power to the power grid node, and a second sample net power of the power grid node in the second sample period and a second sample state of charge of the energy storage device in the second sample period are obtained;

[0015] In the case where the current training round is less than a preset training round, the second sample period is determined as a new first sample period, the second sample net power is determined as a new first sample net power, the second sample state of charge is determined as a new first sample state of charge, and a next training round of the current training round is determined as a new current training round, to obtain new sample data;

[0016] The step of inputting the sample data into the double-delay deep deterministic policy gradient model before training is returned and continued to be executed until the current training round is not less than the preset training round.

[0017] In one of the embodiments, after the step of controlling the energy storage device to output the first predicted active power to the power grid node and controlling the generator set to output the second predicted active power to the power grid node, the method further comprises:

[0018] The function value of the minimum operation cost function, the function value of the minimum voltage deviation function, and the function value of the minimum carbon emission target function obtained in the current training round are obtained;

[0019] Based on the function values and the constraint conditions, a reward value of the multi-objective optimization model is obtained.

[0020] optimizing model parameters of the double-delay deep deterministic policy gradient model based on the reward value.

[0021] In one of the embodiments, the obtaining of the reward value of the multi-objective optimization model based on the function value and the constraint condition comprises:

[0022] obtaining a preset optimal function value and a preset worst function value of any one function in the multi-objective optimization model, and obtaining a first Euclidean distance between the function value of the function and the preset optimal function value and a second Euclidean distance between the function value and the preset worst function value;

[0023] converting the constraint condition into a penalty function, and obtaining the reward value based on the first Euclidean distance, the second Euclidean distance and the penalty function.

[0024] In one of the embodiments, the obtaining of the current net power of any one power distribution network node in a current time period comprises:

[0025] obtaining a first active power output by the photovoltaic device to the power distribution network node in the current time period, a second active power output by the energy storage device to the power distribution network node in the current time period, a third active power output by the generator set to the power distribution network node in the current time period, and an active load power of the power distribution network node in the current time period;

[0026] determining the current net power of the power distribution network node in the current time period based on the first active power, the second active power, the third active power and the active load power.

[0027] In one of the embodiments, the obtaining process of the minimum operation cost function comprises:

[0028] obtaining a first cost function of the generator set and a second cost function of the energy storage device;

[0029] obtaining the minimum operation cost function based on the first cost function and the second cost function.

[0030] In a second aspect, the application further provides a power distribution network dispatching control device, which comprises:

[0031] a first construction module configured to construct, for a power distribution network comprising a photovoltaic device, an energy storage device and a generator set, a multi-objective optimization model of the power distribution network and a constraint condition of the multi-objective optimization model; wherein the multi-objective optimization model comprises a minimum operation cost function, a minimum voltage deviation function and a minimum carbon emission target function;

[0032] an acquisition module configured to acquire a current net power of any one power distribution network node in a current time period and acquire a current state of charge of the energy storage device in the current time period;

[0033] an input module configured to input the current net power, the current state of charge and the current time period into the trained double-delay deep deterministic policy gradient model, and output a first candidate active power of the energy storage device and a second candidate active power of the generator set;

[0034] a second construction module configured to construct a mixed integer programming model embedded with the constraint condition, and input the first candidate active power and the second candidate active power into the mixed integer programming model, and output a first target active power and a second target active power;

[0035] a control module configured to control the energy storage device to output the first target active power to the power distribution network node, and control the generator set to output the second target active power to the power distribution network node.

[0036] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps of the method in any one of the above embodiments are implemented.

[0037] In a fourth aspect, the present application also provides a computer readable storage medium. The computer readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the method in any one of the above embodiments are implemented.

[0038] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program. When the computer program is executed by a processor, the steps of the method in any one of the above embodiments are implemented.

[0039] The power distribution network dispatching control method, device, computer device, readable storage medium and program product construct a multi-objective optimization model of the power distribution network and constraint conditions of the multi-objective optimization model, wherein the multi-objective optimization model includes a minimum operating cost function, a minimum voltage deviation function and a minimum carbon emission target function; the current net power of any power distribution network node in the current period is obtained, and the current state of charge of the energy storage device in the current period is obtained; the current net power, the current state of charge and the current period are input into the trained double-delay deep deterministic policy gradient model, and the first candidate active power of the energy storage device and the second candidate active power of the generator set are output; a mixed integer programming model embedded with the constraint conditions is constructed, and the first candidate active power and the second candidate active power are input into the mixed integer programming model, and the first target active power and the second target active power are output; the energy storage device outputs the first target active power to the power distribution network node, and the generator set outputs the second target active power to the power distribution network node. The method provided in the application can reduce the calculation cost in the power distribution network dispatching process, improve the dispatching efficiency, and will not cause constraint violation problems. BRIEF DESCRIPTION OF DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the description of the embodiments of the present application or the related art will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other related drawings can be obtained by those skilled in the art without creative labor.

[0041] Figure 1 A flowchart of a power distribution network dispatching control method in an embodiment;

[0042] Figure 2 A structure diagram of a power distribution network in an embodiment;

[0043] Figure 3 A flowchart of a training process of a double-delay deep deterministic policy gradient model in an embodiment;

[0044] Figure 4 A structure block diagram of a power distribution network dispatching control device in an embodiment;

[0045] Figure 5 An internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION

[0046] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not intended to limit the present application.

[0047] It should be noted that the terms "first", "second", etc. used in the present application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "include" and "have" used in the present application and any variations thereof are intended to cover non-exclusive inclusion. The term "a plurality of" used in the present application means two or more. The term "and / or" used in the present application means one of the options or any combination of multiple options.

[0048] In one embodiment, as shown in Figure 1 A power distribution network scheduling control method is provided. The embodiment is exemplified by the method applied to a terminal. It should be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and can be realized through the interaction of the terminal and the server. In the embodiment, the method includes the following steps:

[0049] S102, for the power distribution network including photovoltaic devices, energy storage devices and generator sets, a multi-objective optimization model of the power distribution network and constraint conditions of the multi-objective optimization model are constructed; wherein the multi-objective optimization model includes a minimum operating cost function, a minimum voltage deviation function and a minimum carbon emission target function.

[0050] The multi-objective optimization model of the power distribution network is a multi-dimensional optimization system for safe, economic and low-carbon operation of the power distribution network, with the core optimization objectives of minimizing operating cost, minimizing node voltage deviation and minimizing total carbon emission, while embedding power distribution network physical constraints, device operation constraints and network security constraints. The minimum operating cost function is an economic objective function that quantifies the generation cost of the unit, the charging and discharging cost of the energy storage and the network loss cost during the operation of the power distribution network. The minimum voltage deviation function is a safety objective function that measures the deviation degree of the node voltage of the power distribution network from the rated voltage. The minimum carbon emission target function is an environmental protection objective function that quantifies the total carbon emission of fossil energy units during the operation of the power distribution network.

[0051] Optionally, the minimum operating cost function, the minimum voltage deviation function and the minimum carbon emission target function are as follows:

[0052]

[0053]

[0054]

[0055] wherein, , and are a minimum operating cost function, a minimum voltage deviation function and a minimum carbon emission objective function, respectively; is a set of optimization time periods; is a set of nodes of the power distribution network; is a set of nodes with energy storage devices; is a set of nodes with generator units; is the voltage of node m in time period t; is the carbon emission per unit power output of the generator unit in node m; is a first cost function of the generator unit, is a second cost function of the energy storage device.

[0056] Optionally, the constraint conditions of the multi-objective optimization model include active power balance conditions, reactive power balance conditions, voltage drop calculation conditions, branch current calculation conditions, energy storage device state of charge calculation conditions, energy storage device state of charge constraint conditions, energy storage device output constraint conditions, node voltage constraint conditions and branch current constraint conditions, which are sequentially shown in the following formulae:

[0057]

[0058]

[0059]

[0060]

[0061]

[0062]

[0063]

[0064]

[0065]

[0066] wherein, n and m are two adjacent nodes in the power distribution network, respectively; is a set of lines in the power distribution network; , , represent the active power, the reactive power and the current of line mn in time period t, respectively; , represent the resistance and the reactance of line mn, respectively; , , respectively represent the active power of the node with photovoltaic device, energy storage device and generator set corresponding to the device in period t, wherein the energy storage device power is positive indicating charging and negative indicating discharging; and respectively represent the active power and reactive power load demand of node m in period t; is the state of charge of the energy storage device, and correspond to the charging efficiency and discharging efficiency respectively, is the rated capacity of the energy storage device; and , and correspond to the upper and lower limits of the state of charge of the energy storage device and the upper and lower limits of the node voltage respectively; is the current upper limit corresponding to the line mn.

[0067] Optionally, the structure diagram of the power distribution network containing photovoltaic devices, energy storage devices and generator sets is as shown in Figure 2 , Figure 2 , wherein PV is a photovoltaic device, DG is a generator set, BESS is an energy storage device, and a solid point is a power distribution network node. Among them, the output of the photovoltaic device connected to the power distribution network has uncertainty, and the power in each period fluctuates within a certain range. The energy storage device can be charged and discharged in different periods to assist in reducing operating costs and controlling voltage fluctuations.

[0068] S104, acquiring the current net power of any one power distribution network node in the current period, and acquiring the current state of charge of the energy storage device in the current period.

[0069] Optionally, the method of deep reinforcement learning is used to solve the above multi-objective optimization problem, and the generator set and the energy storage device in the power distribution network are regarded as an agent, and the optimization model is converted into a Markov decision process (MDP) containing five tuples , wherein S is the state space, A is the action space, P is the state transition probability, R is the reward function, is the discount factor of future rewards. The interaction process between the agent and the power distribution network environment is as follows: after observing the state , the agent selects the action according to the strategy , interacts with the environment, and obtains the reward , and the state is transferred to the next state ; the agent obtains enough experience by continuously interacting to optimize the strategy to maximize the reward:

[0070]

[0071] Optionally, the state space reflects the perception range of the agent to the current environment, and the state space in the present scheme is defined in consideration of the operation state of each device in the power distribution network as follows:

[0072]

[0073] In the formula, t is the current period, is the net power of node m in period t, is the current state of charge of the energy storage device in the current period.

[0074] S106, input the current net power, the current state of charge and the current period into the trained double-delay deep deterministic policy gradient model, and output the first candidate active power of the energy storage device and the second candidate active power of the generator set.

[0075] Optionally, the scheduling action space is a set of actions available to the agent, and the action space is defined as follows:

[0076]

[0077] In the formula, is the first candidate active power of the energy storage device, is the second candidate active power of the generator set.

[0078] Optionally, the Twin Delayed Deep Deterministic policy gradient (TD3) algorithm relies on the Actor-Critic framework, which combines the strengths of value-based and policy-based methods and can efficiently solve continuous action space problems and long-term correlation problems. The Actor-Critic method includes two types of neural networks: the Actor network (policy network) takes the current state as input and outputs the corresponding action; the Critic network (action value network) takes the current state and action as input and outputs the value of the action in the current state. In addition, target networks corresponding to the Actor network and the Critic network are set up, and network stability is maintained through regular soft updates. During parameter updates, part of the experience is extracted from the experience replay buffer, and the Critic network is updated using an error function based on the time difference algorithm. Then, the Actor network is updated based on maximizing the action value output by the Critic network. After training, the Actor network and the Critic network with fixed parameters can be obtained. Compared with the classic deep deterministic policy gradient algorithm (DDPG) based on the Actor-Critic framework, the TD3 algorithm adds a Critic network, which outputs the smaller value of the two Critic networks during parameter updates, to some extent, alleviating the problem of value overestimation. By reducing the update frequency and delaying the update of the Actor network, the training stability is ensured, and truncated Gaussian noise is added to the action output by the target Actor network to prevent falling into a local optimal solution.

[0079] S108, a mixed integer programming model embedded with constraint conditions is constructed, and the first candidate active power and the second candidate active power are input into the mixed integer programming model to output the first target active power and the second target active power.

[0080] The mixed integer programming (MIP) is a mathematical optimization model that combines continuous variables and integer variables. The core of the model is to convert the value evaluation logic of the Critic network trained by TD3 into mathematical constraints, and embed the full-quantity hard constraints of the power distribution network. Under the premise of meeting the constraint conditions, the optimal scheduling action that balances the multiple objectives of "minimum operating cost", "minimum voltage deviation", and "minimum carbon emission" of the power distribution network is solved.

[0081] Optionally, this embodiment combines the idea of ​​model-based solution and establishes the Critic network as a mixed integer programming model. The original process of executing the highest-scoring action obtained based on the policy network to obtain the highest cumulative reward is transformed into using a solver to solve the problem of maximizing the cumulative action value including constraints. The input variables are the actions in the action space, thereby realizing the online real-time execution of unconstrained violations.

[0082] Optionally, the Critic network consists of K+1 layers, where layer 0 is the data input to the Critic network and layer K+1 is the output. For the Critic network, any... All layers There are n neurons, and the j-th neuron in the k-th layer is defined as... Define the j-th input to the Critic network as After being mapped by the neural network, each neuron in the k-th layer The output is Together they form the output vector of the kth layer. For any Layer, neuron The output vector can be represented as:

[0083]

[0084] In the formula, The activation function selected in this scheme is the ReLU function, defined as shown in the following equation. and Here are the weight matrix and bias coefficients for the k-th layer of the DNN:

[0085]

[0086] Optionally, based on the above definition of the Critic network and the characteristics of the ReLU function, for each neuron Use 0-1 variables This indicates the active state, thus transforming the Critic network into a mixed-integer programming model, as shown in the following formula:

[0087]

[0088] Optionally, the first, second, and third constraints of the mixed-integer programming model are shown in the following equations:

[0089]

[0090]

[0091]

[0092] wherein the first constraint condition is an output constraint of the activation function, and the second and third constraint conditions are upper and lower limits of the variable; and is a cost parameter, , , , are constant parameters, is an auxiliary variable, and when the output vector is greater than 0, is equal to 0, and vice versa represents a negative part of the neuron. When the input stage k is equal to 0, the upper and lower limits of the variable are the same as the input limit of the DNN, and when , the upper and lower limits are defined based on the parameter .

[0093] Optionally, considering the constraint conditions, the first constraint condition, the second constraint condition and the third constraint condition of the multi-objective optimization model, the final mixed integer programming model can be obtained as shown in the following formula:

[0094]

[0095] S110, control the energy storage device to output the first target active power to the power distribution network node, and control the generator set to output the second target active power to the power distribution network node.

[0096] Optionally, the first target active power and the second target active power are output to the power distribution network node, and the power distribution network node is controlled to operate; through accurate active power control of the energy storage device and the generator set, power supply and demand balance of the power distribution network node is realized, and multi-objective optimization targets of “minimum operating cost, minimum voltage deviation and minimum carbon emission” are achieved, thereby guaranteeing safe, stable, economic and low-carbon operation of the power distribution network.

[0097] In the power distribution network dispatching control method, for a power distribution network including a photovoltaic device, an energy storage device and a generator set, a multi-objective optimization model of the power distribution network and constraint conditions of the multi-objective optimization model are constructed; the multi-objective optimization model includes a minimum operation cost function, a minimum voltage deviation function and a minimum carbon emission target function; a current net power of any power distribution network node in a current time period is obtained, and a current state of charge of the energy storage device in the current time period is obtained; the current net power, the current state of charge and the current time period are input into the trained double-delay deep deterministic policy gradient model to output a first candidate active power of the energy storage device and a second candidate active power of the generator set; a mixed integer programming model embedded with the constraint conditions is constructed, and the first candidate active power and the second candidate active power are input into the mixed integer programming model to output a first target active power and a second target active power; the energy storage device is controlled to output the first target active power to the power distribution network node, and the generator set is controlled to output the second target active power to the power distribution network node. The method provided in the application can reduce the calculation cost in the power distribution network dispatching process, improve the dispatching efficiency, and will not cause constraint violation problems.

[0098] In some embodiments, as shown in Figure 3 The training process of the double-delay deep deterministic policy gradient model includes:

[0099] In S302, for a current training round in the training process, sample data of the current training round is obtained; the sample data includes a first sample time period, a first sample net power of any power distribution network node in the first sample time period, and a first sample state of charge of the energy storage device in the first sample time period.

[0100] In S304, the sample data is input into the double-delay deep deterministic policy gradient model before training to output a first predicted active power of the energy storage device and a second predicted active power of the generator set.

[0101] In S306, in a second sample time period after the first sample time period, the energy storage device is controlled to output the first predicted active power to the power distribution network node, and the generator set is controlled to output the second predicted active power to the power distribution network node, and a second sample net power of the power distribution network node in the second sample time period and a second sample state of charge of the energy storage device in the second sample time period are obtained.

[0102] In S308, when the current training round is less than a preset training round, the second sample time period is determined as a new first sample time period, the second sample net power is determined as a new first sample net power, the second sample state of charge is determined as a new first sample state of charge, and a next training round of the current training round is determined as a new current training round, to obtain new sample data.

[0103] S310, return to the step of inputting the sample data into the double-delay deep deterministic policy gradient model before training and continue to execute until the current training round is not less than the preset training round.

[0104] Optionally, the Actor network and the target Actor network are randomly initialized ; the Critic network and are initialized , the target Critic network and are initialized , and the parameters of the Actor network and the Critic network are randomly initialized ; the network parameters are copied to the corresponding target network, , , ; the number of training rounds and the number of steps of each round of training are set, the hyperparameters including the update frequency , the learning rate , the experience replay buffer size , the number of experience sampling samples , the soft update coefficient are set, the action corresponding to the current state is calculated , wherein is the noise ; the obtained action is used to interact with the environment to obtain the corresponding reward and the next state ; the experience is stored in the experience replay buffer; a plurality of experiences are randomly extracted from the experience replay buffer, and the action with added noise is output by the target Actor network , ; the target value is calculated using the target value network ; if the remainder of the previous loop variable i divided by is equal to 0, the Actor network is updated, and the soft update target network , , .

[0105] Optionally, after the state selects an action according to the policy and interacts with the environment, it is transferred to the next state The probability is defined as shown in the following formula:

[0106]

[0107] In this embodiment, the sample data closely follows the core state of the power distribution network (time period, node net power, and energy storage state of charge), accurately depicts the key input features of scheduling decision, and provides high-value data support for TD3 model training; the closed-loop training logic (sample input → power prediction → action execution → state feedback) is adopted to realize end-to-end iterative optimization of the TD3 model, so that the model continuously learns and adapts to the optimal scheduling strategy of the power distribution network.

[0108] In some embodiments, after controlling the energy storage device to output the first predicted active power to the power distribution network node and controlling the generator set to output the second predicted active power to the power distribution network node, the method further includes: obtaining a function value of the minimum operating cost function, a function value of the minimum voltage deviation function, and a function value of the minimum carbon emission target function obtained in the current training round; obtaining a reward value of the multi-objective optimization model based on the function values and the constraint conditions; and optimizing the model parameters of the double-delay deep deterministic policy gradient model based on the reward value.

[0109] Optionally, the reward functions of the minimum operating cost function, the minimum voltage deviation function, and the minimum carbon emission target function are shown in the following formulas in sequence:

[0110]

[0111]

[0112]

[0113] In this embodiment, the effectiveness of multi-objective optimization is quantified to provide accurate reward feedback basis for model optimization, and strong binding of multi-objective scheduling effect and model training is realized; the reward value is generated by combining the cost, voltage, and carbon emission target values with the constraint conditions, which takes into account the scheduling optimality and compliance, and ensures the accuracy of the model learning direction.

[0114] In some embodiments, based on the function values and the constraint conditions, the reward value of the multi-objective optimization model is obtained by: obtaining a preset optimal function value and a preset worst function value of any one function in the multi-objective optimization model, and obtaining a first Euclidean distance between the function value of the function and the preset optimal function value and a second Euclidean distance between the function value and the preset worst function value; converting the constraint conditions into a penalty function, and obtaining the reward value based on the first Euclidean distance, the second Euclidean distance, and the penalty function.

[0115] Optionally, since the units of the minimum operating cost function, the minimum voltage deviation function, and the minimum carbon emission target function are inconsistent, the better-worse distance method is adopted to construct the multi-objective reward function.

[0116] Optionally, the power distribution network operation plan is standardized based on the following formula.

[0117]

[0118] In the formula, is the power distribution network operation plan, and are the actual operation plan and the standardized operation plan of the optimization target i in the time period t, and correspond to the worst solution and the optimal solution of the optimization target, respectively.

[0119] Optionally, based on the above definition, the theoretical optimal solution is and the worst solution is Therefore, for the operation plan , the distance from the optimal solution and the distance from the worst solution can be calculated:

[0120]

[0121] Optionally, if the operation plan is closer to the optimal solution and farther from the worst solution, it means that the operation plan is better. Based on the distance, the immediate reward of the energy storage scheduling is constructed as follows:

[0122]

[0123] Optionally, the constraint condition is converted into an additional penalty in the reward function, and the above optimization problem containing the constraint condition is converted into an unconstrained optimization problem. The power upper and lower limit constraints have been met when adding the action space boundary, and the penalty function is defined as follows:

[0124]

[0125] In the formula, and are penalty coefficients, is a 0-1 variable, which is 0 when the power flow does not converge and 1 otherwise. The reward function of the deep reinforcement learning considering the constraint condition conversion is defined as follows:

[0126]

[0127] In this embodiment, the Euclidean distance quantization function value is used to quantify the difference between the optimal / worst value, to achieve accurate normalization evaluation of multi-objective optimization effect, and to eliminate the dimensional difference of different objective functions; the constraint condition is converted into a penalty function, which takes into account the multi-objective optimization effect and the power distribution network operation compliance, and guides the model to avoid illegal scheduling strategies.

[0128] In some embodiments, the current net power of any one power distribution network node in the current period is obtained by: obtaining a first active power output by the photovoltaic device to the power distribution network node in the current period, a second active power output by the energy storage device to the power distribution network node in the current period, a third active power output by the generator set to the power distribution network node in the current period, and an active load power of the power distribution network node in the current period; and determining the current net power of the power distribution network node in the current period based on the first active power, the second active power, the third active power, and the active load power.

[0129] Optionally, the current net power can be calculated by the following formula:

[0130]

[0131] In the formula, is the first active power, is the second active power, is the third active power, is the active load power.

[0132] In the embodiment, the node net power is directly calculated by splitting the output of the photovoltaic device, the energy storage device, and the generator set, and the node load power, clearly reflecting the power surplus or gap of the node, and providing a core state basis for power distribution network dispatching decision.

[0133] In some embodiments, the process of obtaining the minimum operation cost function includes: obtaining a first cost function of the generator set and a second cost function of the energy storage device; and obtaining the minimum operation cost function based on the first cost function and the second cost function.

[0134] Optionally, the first cost function and the second cost function are as follows:

[0135]

[0136]

[0137] In the formula, is the output power of the generator set i, , , and are cost coefficients, is the charge / discharge power of the energy storage device.

[0138] In the embodiment, the comprehensive operation cost is split into two independent modules of the generator set and the energy storage device, avoiding the modeling difficulty caused by the coupling of multiple device costs, and improving the interpretability and realizability of the cost function.

[0139] In one exemplary embodiment, another power distribution network scheduling control method is provided, which includes the following:

[0140] Step 1: Establish a typical power distribution network model containing photovoltaic and energy storage devices, construct the minimum operating cost, minimum voltage deviation, and minimum carbon emission objective functions, and corresponding constraint conditions.

[0141] Step 2: Based on Markov decision process, the above model-based optimization problem with constraints is converted into an unconstrained agent exploring optimal strategy problem based on reward function using deep reinforcement learning algorithm.

[0142] Step 3: The TD3 algorithm is used to solve the above problem, and a neural network with fixed parameters is obtained after training.

[0143] Step 4: Based on the Critic network obtained by training, a corresponding mixed integer programming model is established, and the optimal strategy under the constraint condition is obtained by considering the constraint condition established in Step 1.

[0144] Step 5: According to the actual power distribution network operation state, set the initial value of the above model and the algorithm parameter, train the agent using the TD3 algorithm, obtain the neural network with fixed parameters after training, and perform online execution using the method of Step 4 to obtain the corresponding optimal strategy in real time.

[0145] Based on the agent training process of Step 3, combined with the objective function and constraint condition established in Step 1 and the optimization problem conversion method of Step 2, the agent is trained using the deep reinforcement learning method, i.e. the TD3 algorithm, and the neural network with fixed parameters is obtained after optimization of the strategy. Then, combined with the neural network conversion process of Step 4, the Critic network is established as a mixed integer programming model, the problem is optimized and solved considering the safe and stable operation constraints of the power distribution network, and the optimal multi-objective control strategy satisfying the constraint condition is obtained in real time.

[0146] Step 6: The actual power distribution network model and parameters are brought into the above method for solving.

[0147] It should be understood that although each step in the flowchart involved in the above-described embodiments is shown in sequence according to the direction of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, there is no strict order limitation for the execution of these steps, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart involved in the above-described embodiments can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be alternately executed with at least part of other steps or steps or stages in other steps. It can be understood that the steps in different embodiments can be freely combined as needed, and various non-contradictory schemes formed by the combination are within the scope of protection of the present application.

[0148] Based on the same inventive concept, the embodiments of the present application also provide a power distribution network scheduling control device for implementing the above-mentioned power distribution network scheduling control method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more power distribution network scheduling control device embodiments provided below can refer to the limitations of the power distribution network scheduling control method described above, which will not be repeated here.

[0149] In one exemplary embodiment, as shown in Figure 4 A power distribution network scheduling control device 400 is provided, comprising a first construction module 401, an acquisition module 402, an input module 403, a second construction module 404, and a control module 405, wherein:

[0150] The first construction module 401 is configured to construct, for a power distribution network comprising a photovoltaic device, an energy storage device, and a generator set, a multi-objective optimization model of the power distribution network and constraint conditions of the multi-objective optimization model; wherein the multi-objective optimization model comprises a minimum operating cost function, a minimum voltage deviation function, and a minimum carbon emission target function.

[0151] The acquisition module 402 is configured to acquire a current net power of any one power distribution network node in a current period and acquire a current state of charge of the energy storage device in the current period.

[0152] The input module 403 is configured to input the current net power, the current state of charge, and the current period into the trained double-delay deep deterministic policy gradient model to output a first candidate active power of the energy storage device and a second candidate active power of the generator set.

[0153] The second construction module 404 is configured to construct a mixed integer programming model embedded with the constraint condition, and input the first candidate active power and the second candidate active power into the mixed integer programming model, and output a first target active power and a second target active power.

[0154] The control module 405 is configured to control the energy storage device to output the first target active power to the power grid node, and control the generator set to output the second target active power to the power grid node.

[0155] In some embodiments, the power grid scheduling control device 400 is specifically configured to, for a current training round in a training process, obtain sample data of the current training round, wherein the sample data comprises a first sample period, a first sample net power of any one power grid node in the first sample period, and a first sample state of charge of the energy storage device in the first sample period; input the sample data into a pre-training double-delay deep deterministic policy gradient model, and output a first predicted active power of the energy storage device and a second predicted active power of the generator set; in a second sample period after the first sample period, control the energy storage device to output the first predicted active power to the power grid node, and control the generator set to output the second predicted active power to the power grid node, and obtain a second sample net power of the power grid node in the second sample period and a second sample state of charge of the energy storage device in the second sample period; in a case where the current training round is less than a preset training round, determine the second sample period as a new first sample period, the second sample net power as a new first sample net power, the second sample state of charge as a new first sample state of charge, and a next training round of the current training round as a new current training round, to obtain new sample data; return to the step of inputting the sample data into the pre-training double-delay deep deterministic policy gradient model and continue to execute until the current training round is not less than the preset training round.

[0156] In some embodiments, the power grid scheduling control device 400 is further configured to obtain a function value of the minimum operation cost function, a function value of the minimum voltage deviation function, and a function value of the minimum carbon emission target function obtained in the current training round; based on the function values and the constraint condition, obtain a reward value of the multi-objective optimization model; and based on the reward value, optimize model parameters of the double-delay deep deterministic policy gradient model.

[0157] In some embodiments, the power distribution network dispatching control device 400 is further configured to obtain a preset optimal function value and a preset worst function value of any one function in the multi-objective optimization model, and obtain a first Euclidean distance between a function value of the function and the preset optimal function value and a second Euclidean distance between the function value and the preset worst function value; convert the constraint condition into a penalty function, and obtain the reward value based on the first Euclidean distance, the second Euclidean distance and the penalty function.

[0158] In some embodiments, the obtaining module 402 is further configured to obtain a first active power output by the photovoltaic device to the power distribution network node in the current time period, a second active power output by the energy storage device to the power distribution network node in the current time period, a third active power output by the generator set to the power distribution network node in the current time period, and an active load power of the power distribution network node in the current time period; and determine a current net power of the power distribution network node in the current time period based on the first active power, the second active power, the third active power and the active load power.

[0159] In some embodiments, the first constructing module 401 is further configured to obtain a first cost function of the generator set and a second cost function of the energy storage device; and obtain the minimum operation cost function based on the first cost function and the second cost function.

[0160] The above various modules in the power distribution network dispatching control device can be all or partially realized by software, hardware and combinations thereof. The above various modules can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in the computer device in software form, so as to be called and executed by a processor to perform the operations corresponding to the above various modules.

[0161] In one exemplary embodiment, a computer device is provided, which can be a terminal, and an internal structure diagram of the computer device can be as shown in Figure 5As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. Among them, the processor, the memory and the input / output interface are connected through the system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capability. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be realized through WIFI, mobile cellular network, Near Field Communication (NFC) or other technologies. The computer program is executed by the processor to realize a power distribution network scheduling control method.

[0162] Those skilled in the art can understand that, Figure 5 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0163] In one embodiment, a computer device is also provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to realize the steps in each of the above method embodiments.

[0164] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by the processor to realize the steps in each of the above method embodiments.

[0165] In one embodiment, a computer program product is provided, including a computer program, and the computer program is executed by the processor to realize the steps in each of the above method embodiments.

[0166] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.

[0167] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., without being limited thereto.

[0168] The technical features of the above embodiments can be combined in any manner. To make the description concise, all possible combinations of the technical features in the above embodiments are not described, but as long as the combinations of the technical features do not exist, they should be considered as the scope of the present application.

[0169] The above-described embodiments are merely illustrative of several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.

Claims

1. A distribution network dispatching and control method, characterized in that, The method includes: For a distribution network that includes photovoltaic devices, energy storage devices, and generator sets, a multi-objective optimization model for the distribution network and the constraints of the multi-objective optimization model are constructed; wherein, the multi-objective optimization model includes a minimum operating cost function, a minimum voltage deviation function, and a minimum carbon emission objective function; Obtain the current net power of any distribution network node in the current time period, and obtain the current state of charge of the energy storage device in the current time period; The current net power, the current state of charge, and the current time period are input into the trained dual-delay deep deterministic policy gradient model, and the first candidate active power of the energy storage device and the second candidate active power of the generator set are output. Construct a mixed integer programming model with the aforementioned constraints, and input the first candidate active power and the second candidate active power into the mixed integer programming model to output the first target active power and the second target active power; The energy storage device is controlled to output the first target active power to the distribution network node, and the generator set is controlled to output the second target active power to the distribution network node.

2. The method according to claim 1, characterized in that, The training process of the dual-delay deep deterministic policy gradient model includes: For the current training round during the training process, sample data for the current training round is obtained; wherein, the sample data includes a first sample time period, the first sample net power of any distribution network node in the first sample time period, and the first sample state of charge of the energy storage device in the first sample time period; The sample data is input into the double-delay deep deterministic policy gradient model before training, and the first predicted active power of the energy storage device and the second predicted active power of the generator set are output. In the second sample period following the first sample period, the energy storage device is controlled to output the first predicted active power to the distribution network node, and the generator set is controlled to output the second predicted active power to the distribution network node. The second sample net power of the distribution network node in the second sample period and the second sample state of charge of the energy storage device in the second sample period are obtained. If the current training round is less than the preset training round, the second sample time period is determined as the new first sample time period, the second sample net power is determined as the new first sample net power, the second sample state of charge is determined as the new first sample state of charge, and the next training round of the current training round is determined as the new current training round, so as to obtain new sample data. Return to the step of inputting the sample data into the dual-delay deep deterministic policy gradient model before training and continue execution until the current training round is not less than the preset training round.

3. The method according to claim 2, characterized in that, After controlling the energy storage device to output the first predicted active power to the distribution network node and controlling the generator set to output the second predicted active power to the distribution network node, the method further includes: Obtain the function values ​​of the minimum operating cost function, the minimum voltage deviation function, and the minimum carbon emission target function obtained in the current training round; Based on the function value and the constraints, the reward value of the multi-objective optimization model is obtained; The model parameters of the dual-delay deep deterministic policy gradient model are optimized based on the reward value.

4. The method according to claim 3, characterized in that, The step of obtaining the reward value of the multi-objective optimization model based on the function value and the constraints includes: Obtain the preset optimal function value and preset worst function value of any function in the multi-objective optimization model, and obtain the first Euclidean distance between the function value and the preset optimal function value, and the second Euclidean distance between the function value and the preset worst function value; The constraints are transformed into a penalty function, and the reward value is obtained based on the first Euclidean distance, the second Euclidean distance, and the penalty function.

5. The method according to claim 1, characterized in that, The step of obtaining the current net power of any distribution network node in the current time period includes: The system obtains the first active power output by the photovoltaic device to the distribution network node during the current time period, the second active power output by the energy storage device to the distribution network node during the current time period, the third active power output by the generator set to the distribution network node during the current time period, and the active load power of the distribution network node during the current time period. Based on the first active power, the second active power, the third active power, and the active load power, the current net power of the distribution network node in the current time period is determined.

6. The method according to claim 1, characterized in that, The process of obtaining the minimum operating cost function includes: Obtain the first cost function of the generator set and the second cost function of the energy storage device; Based on the first cost function and the second cost function, the minimum operating cost function is obtained.

7. A power distribution network dispatching and control device, characterized in that, The device includes: The first construction module is used to construct a multi-objective optimization model of a distribution network that includes photovoltaic devices, energy storage devices, and generator sets, as well as the constraints of the multi-objective optimization model; wherein, the multi-objective optimization model includes a minimum operating cost function, a minimum voltage deviation function, and a minimum carbon emission objective function; The acquisition module is used to acquire the current net power of any distribution network node in the current time period, and to acquire the current state of charge of the energy storage device in the current time period. The input module is used to input the current net power, the current state of charge, and the current time period into the trained dual-delay deep deterministic policy gradient model, and output the first candidate active power of the energy storage device and the second candidate active power of the generator set. The second construction module is used to construct a mixed integer programming model embedded with the constraints, and input the first candidate active power and the second candidate active power into the mixed integer programming model, and output the first target active power and the second target active power; The control module is used to control the energy storage device to output the first target active power to the distribution network node, and to control the generator set to output the second target active power to the distribution network node.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.