Power distribution network operation optimization method and device, equipment and storage medium

By constructing a second-order cone optimal power flow model and Markov decision process for the distribution network in multiple time periods and combining it with reinforcement learning methods, the distributed photovoltaic access capacity is optimized, solving the problems of insufficient distribution network carrying capacity and poor adaptability to dynamic scenarios, and achieving efficient dynamic operation optimization.

CN120728558APending Publication Date: 2025-09-30STATE GRID JIANGSU ECONOMIC RES INST
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510680105.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing technologies are difficult to effectively improve the distribution network's carrying capacity for distributed photovoltaics, and have poor adaptability and low optimization efficiency in dynamic scenarios.

Method used

A second-order cone optimal power flow model for the distribution network in multiple time periods is constructed. By combining the Markov decision process and reinforcement learning method, the distributed photovoltaic access capacity is optimized and the equipment operating status is dynamically adjusted through pre-training decision agents.

Benefits of technology

It improves the operating performance and carrying capacity of the distribution network in dynamic scenarios, meets the time requirements of real-time scheduling, and improves the overall optimization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120728558A_ABST
    Figure CN120728558A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a power distribution network operation optimization method and device, equipment and a storage medium, and the method comprises the steps: constructing a multi-period second-order cone optimal power flow model of a power distribution network, and solving to obtain a multi-period operation optimal solution of the power distribution network; converting the optimization problem of the second-order cone optimal power flow model into a Markov decision process to obtain a Markov decision model; determining a target state corresponding to the optimal solution of multi-period operation of the power distribution network, and constructing a pre-training data set based on the optimal solution and the corresponding target state; and pre-training the decision-making agent by using the pre-training data set, training the pre-trained decision-making agent by using an algorithm in reinforcement learning based on the Markov decision-making model, and outputting an optimal power distribution network multi-period operation strategy. The distributed photovoltaic bearing capacity can be effectively improved, the overall optimization efficiency is improved, and the problem of poor dynamic scene adaptation is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of power grid resource optimization, and in particular to a distribution network operation optimization method, apparatus, device, and storage medium. Background Art

[0002] With the advancement of the "dual carbon" goals, the penetration of distributed photovoltaics in distribution networks has rapidly increased, becoming a key force in energy transition. However, the intermittent and random nature of distributed photovoltaic output leads to increased voltage fluctuations at nodes within the distribution network. Especially during periods of abundant sunlight during the midday hours, a surge in photovoltaic power generation can cause local voltage limits to exceed certain limits. At night or on rainy days, power must be purchased from the main grid, exacerbating the risk of voltage drops. Therefore, in terms of optimizing distribution network operations, it is particularly important to conduct real-time research on maximizing the capacity of distributed photovoltaics within the distribution network to enhance the distribution network's carrying capacity for distributed photovoltaics.

[0003] Among related technologies, in terms of distribution network optimization, there is no effective solution to how to improve the distribution network's carrying capacity for distributed photovoltaics; in addition, most current research methods are based on the optimization of distribution network operation under static time sections, ignoring the continuous time coupling characteristics of some distribution network parameters, resulting in poor adaptability to dynamic scenarios; in terms of distribution network optimization, existing research methods also have the problem of low optimization efficiency. Summary of the Invention

[0004] The embodiments of the present application provide a distribution network operation optimization method, device, equipment and storage medium, which can effectively improve the distribution network's carrying capacity for distributed photovoltaics, improve the overall optimization efficiency, improve the distribution network's operating performance in dynamic scenarios, and effectively solve the problem of poor adaptability to dynamic scenarios.

[0005] In a first aspect, an embodiment of the present application provides a method for optimizing distribution network operation, comprising:

[0006] An objective function is constructed with the goal of maximizing the total amount of distributed photovoltaic power connected to the distribution network and minimizing the total operating cost of the distribution network. A second-order cone optimal power flow model for the distribution network in multiple time periods is constructed based on the objective function, with power flow constraints, voltage and current safety constraints, photovoltaic output constraints, energy storage dynamic constraints, reactive compensation equipment constraints, and on-load tap-changing transformer constraints as constraints.

[0007] Solving the second-order cone optimal power flow model through a preset solver to obtain the optimal solution for the multi-period operation of the distribution network;

[0008] Converting the optimization problem of the second-order cone optimal power flow model into a Markov decision process to obtain a Markov decision model;

[0009] Determining a target state corresponding to the optimal solution for the multi-period operation of the distribution network based on the state space of the Markov decision model, and constructing a pre-training data set based on the optimal solution for the multi-period operation of the distribution network and the corresponding target state;

[0010] The pre-training data set is used to pre-train the decision-making agent, and based on the Markov decision model, the algorithm in reinforcement learning is used to train the pre-trained decision-making agent, and the optimal distribution network multi-period operation strategy is output based on the trained decision-making agent, wherein the optimal distribution network multi-period operation strategy is the optimal distribution network multi-period operation plan.

[0011] In a second aspect, an embodiment of the present application provides a distribution network operation optimization device, comprising:

[0012] A model construction module is used to construct an objective function with the goal of maximizing the total amount of distributed photovoltaic power connected to the distribution network and minimizing the total operating cost of the distribution network. The module uses power flow constraints, voltage and current safety constraints, photovoltaic output constraints, energy storage dynamic constraints, reactive compensation equipment constraints, and on-load tap-changing transformer constraints as constraints, and combines the objective function to construct a multi-period second-order cone optimal power flow model for the distribution network.

[0013] A model solving module, configured to solve the second-order cone optimal power flow model using a preset solver to obtain an optimal solution for the multi-period operation of the distribution network;

[0014] a conversion module, configured to convert the optimization problem of the second-order cone optimal power flow model into a Markov decision process to obtain a Markov decision model;

[0015] A pre-training data set construction module is used to determine the target state corresponding to the optimal solution for the multi-period operation of the distribution network based on the state space of the Markov decision model, and to construct a pre-training data set based on the optimal solution for the multi-period operation of the distribution network and the corresponding target state;

[0016] A training output module is used to pre-train the decision-making agent using the pre-training data set, and based on the Markov decision model, use the algorithm in reinforcement learning to train the pre-trained decision-making agent, and output the optimal distribution network multi-period operation strategy based on the trained decision-making agent, wherein the optimal distribution network multi-period operation strategy is the optimal distribution network multi-period operation plan.

[0017] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method provided in the embodiment of the present application is implemented.

[0018] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed in a computer, it enables the computer to execute the method provided in the embodiment of the present application.

[0019] The technical solution provided in the embodiment of the present application constructs a second-order cone optimal power flow model, and obtains the optimal solution for the multi-period operation of the distribution network through the model, that is, obtains the distribution network operation optimization strategy, and obtains the Markov decision model by converting the optimization problem of the second-order cone optimal power flow model into a Markov decision process; the decision-making agent is pre-trained with the optimal solution for the multi-period operation of the distribution network, and based on the Markov decision model, the algorithm in reinforcement learning is used to train the pre-trained decision-making agent, and the trained decision-making agent outputs the optimal multi-period operation strategy of the distribution network, that is, in the Markov decision process in reinforcement learning, the pre-trained decision-making agent is trained with time-series environmental state data, and the trained decision-making agent obtains the optimal multi-period operation strategy of the distribution network through real-time state data; in summary, in the embodiment of the present application, the multi-period operation of the distribution network is obtained by constructing a second-order cone optimal power flow model for the distribution network. The optimal solution is to train the decision-making agent in sequence through the optimal solution of the multi-period operation of the distribution network and the time-series environmental state obtained based on the Markov decision model, which can optimize the distributed photovoltaic access capacity in real time, improve the operating performance of the distribution network in dynamic scenarios, and effectively solve the problem of poor adaptability to dynamic scenarios; the embodiment of the present application obtains the optimal multi-period operation strategy of the distribution network through the collaborative optimization of the second-order cone optimal power flow model and the reinforcement learning method, which can improve the distribution network's carrying capacity for distributed photovoltaics; the embodiment of the present application pre-trains the decision-making agent through the optimal solution of the multi-period operation of the distribution network output by the second-order cone optimal power flow model, which can accelerate the convergence of the decision-making agent, and train the pre-trained decision-making agent based on the Markov decision model and the algorithm in reinforcement learning. The trained decision-making agent can quickly make decisions, improve the overall optimization efficiency, and meet the time requirements of real-time scheduling of the distribution network. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 This is a flow chart of a distribution network operation optimization method provided by the present application;

[0021] Figure 2 This is a flow chart of a distribution network operation optimization method provided by the present application;

[0022] Figure 3 This is a flowchart of training a pre-trained decision-making agent provided in an embodiment of the present application;

[0023] Figure 4 This is a structural block diagram of a distribution network operation optimization device provided in an embodiment of the present application;

[0024] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The present application is further described in detail below through the accompanying drawings and specific implementation methods.

[0026] Figure 1 This is a flow chart of a distribution network operation optimization method provided in an embodiment of the present application. The method can be executed by a distribution network operation optimization device, the device can be implemented by software and / or hardware, and the device can be configured in electronic devices such as computers.

[0027] like Figure 1 The technical solution provided in the embodiment of the present application includes the following steps:

[0028] S110: Construct an objective function with the goal of maximizing the total amount of distributed photovoltaic power connected to the distribution network and minimizing the total operating cost of the distribution network, and use power flow constraints, voltage and current safety constraints, photovoltaic output constraints, energy storage dynamic constraints, reactive compensation equipment constraints and on-load tap-changing transformer constraints as constraints. Combined with the objective function, a second-order cone optimal power flow model for the distribution network in multiple time periods is constructed.

[0029] In this embodiment, for a distribution network with a high proportion of distributed photovoltaics in a typical application scenario, an objective function is established to minimize the total investment and operation cost for 24 hours and maximize the total capacity of connected distributed photovoltaics.

[0030] Wherein, the objective function is:

[0031]

[0032] Wherein, f1 is the total capacity of distributed photovoltaic connected to the distribution network; f2 is the total investment and operation cost (which can be the total investment and operation cost for 24 hours); C DG,i is the distributed photovoltaic capacity accessible to node i in the distribution network; C grid The unit price of electricity purchased from the main grid; P g (t) is the power purchased by the main grid in period t; I l (t) Current value of branch 1; R l is the resistance of branch 1, C loss is the unit cost of network loss, and N is the number of branches.

[0033] In this embodiment, key constraints also need to be set. The key constraints include power flow constraints, voltage and current safety constraints, photovoltaic output constraints, energy storage dynamic constraints, reactive compensation equipment constraints, and on-load tap-changing transformer constraints.

[0034] Among them, power flow constraints (second-order cone relaxation) can be used to relax the non-convex power equation into a convex constraint through second-order cone relaxation, ensuring that the problem is solvable and globally optimal. Specifically, these constraints include node power constraints, voltage equations, and second-order cone constraints. Node power constraints account for the injection and outflow of active and reactive power at each node; the voltage equation describes the relationship between voltage and power transfer between nodes; and second-order cone constraints are a key step in achieving problem transformation.

[0035] Specifically, in one embodiment, the node power constraint formula may be:

[0036]

[0037] Among them, P in,i,t is the active power injected by node i in the distribution network at the tth period, P load,i,t is the load of the node i in the tth period, P g,i,t is the active power input from the main grid to the node i in the tth period; P DG,i,t is the distributed photovoltaic power generated at the node i in the t period; P dch,i,t is the energy storage system discharge power at the node i in the t period, P ch,,i,t is the charging power of the energy storage system at node i in period t, Q in,i,t is the reactive power injected into the node i in the tth period, Q load,i,t is the load reactive power of the node i in the tth period; Q g,i,t Q is the reactive power input from the main grid to the node i in the t period; SVC,i,t is the reactive power generated by the static VAR compensator at the node i in the tth period, Q CB,i,t is the reactive power generated by the group-switched capacitors at the node i in the tth period;

[0038] Wherein, the voltage equation is:

[0039]

[0040] The formula for the second-order cone constraint is:

[0041]

[0042] Among them, U j,t is the voltage of node j in the tth period; U i,t is the voltage of the node i in the tth period; R ij is the resistance between the node i and the node j; X ij is the reactance between the node i and the node j; P ij,tQ is the active power transmitted from node i to node j in time period t; ij,t I is the reactive power transmitted from the node i to the node j in the tth period; ij,t is the current between the node i and the node j in the tth period.

[0043] In this embodiment, voltage and current safety constraints can be set upper and lower limits for voltage and current amplitudes to prevent voltage and current exceeding the limits from causing equipment damage, thereby ensuring the safe and stable operation of the distribution network. Photovoltaic output constraints can be used to limit the active and reactive power output of distributed photovoltaics to ensure that photovoltaic equipment operates within a safe and efficient range. Energy storage dynamic constraints can include charge and discharge power constraints, capacity dynamic constraints, and capacity upper and lower limits and cycle constraints. Among them, charge and discharge power constraints specify the charge and discharge status and power range of the energy storage system at different time periods; capacity dynamic constraints describe the changes in energy storage capacity as the charge and discharge process progresses; capacity upper and lower limits and cycle constraints ensure that the initial and final capacity of the energy storage system are the same each day and limit its capacity range. Reactive compensation equipment constraints can be set for capacitors and static var compensators (SVCs) respectively. Among them, capacitor constraints include real-time reactive power calculation and gear switching number limits; static var compensator constraints limit its real-time reactive power range. On-load tap changer (OLTC) constraints specify how to calculate the OLTC's transformation ratio and the number of gear switching times to achieve effective voltage regulation.

[0044] Specifically, the formula for the voltage and current safety constraint is:

[0045]

[0046] Among them, U i,min with U i,max are the upper limit and lower limit of the voltage amplitude of the node i, respectively. ij,max is the upper limit of the current amplitude between the node i and the node j;

[0047] The formula for the photovoltaic output constraint is:

[0048]

[0049] Among them, P DG,i,t With Q DG,i,t are the active and reactive power output of the distributed photovoltaic connected to the node i in the t period, i∈S N , S N is a set of all nodes in the distribution network;

[0050] Among them, P DG,i,max is the upper limit of the output active power of the distributed photovoltaic on the node i; Q DG,i,max is the upper limit of the output reactive power of the distributed photovoltaic at the node i;

[0051] The energy storage dynamic constraints include charge and discharge power constraints, capacity dynamic constraints, and capacity upper and lower limits and cycle constraints;

[0052] The formula for the charge and discharge power constraint is:

[0053]

[0054] Among them, u ch,i,t represents the charging state of energy storage system i at time period t; u dch,i,t represents the discharge state of the energy storage system i at the tth time period, where 1 represents execution and 0 represents non-execution; p ch,i,t represents the charging power of the energy storage system i in the t period; P ch,i,max represents the maximum charging power of the energy storage system i; p dch,i,t represents the discharge power of the energy storage system i in the tth period; P dch,i,max represents the maximum discharge power of the energy storage system i;

[0055] The formula for the dynamic capacity constraint is:

[0056]

[0057] Among them, E i,t and E i,t+1 are the energy storage capacity of the node i in the tth period and the t+1th period respectively; η ch,i is the charging efficiency of the node i, where η ch,i =0.9;η dch,i is the discharge efficiency of the node i, where η dch,i =0.9; discharge coefficient

[0058] The formulas for capacity upper and lower limits and cycle constraints are:

[0059]

[0060] Among them, E i,1 =E i,T+1 Indicates that the energy storage capacity of the node i in the first period and the T+1 period is the same, that is, it indicates that the energy storage capacity of the node i at the beginning and end of each day is the same; E i,min is the lower limit of the energy storage capacity of the node i; E i,maxis the upper limit of the energy storage capacity of the node i;

[0061] The reactive compensation equipment constraints include capacitor-related constraints and static VAR compensator constraints;

[0062] The formula for the capacitor-related constraints is:

[0063]

[0064] Among them, Q CB,i,t , Q step,i are respectively the real-time total reactive power of the capacitor on the node i and the reactive power that can be provided by one gear of the capacitor in the t period; θ CB,i,t,k represents the gear position of the capacitor on the node i in the tth period; θ IN,i,t and θ DE,i,t are the number of times the capacitor on node i is put into operation and the number of times it is removed in time period t;

[0065] The formula for the static VAR compensator constraint is:

[0066] -Q SVC,max ≤Q SVC,t ≤Q SVC,max

[0067] Among them, Q SVC,t , Q SVC,max are the real-time reactive power and maximum reactive power of the static VAR compensator respectively;

[0068] The formula for the on-load tap-changing transformer constraint is:

[0069]

[0070] Among them, r 1,t is the transformation ratio of the on-load tap-changing transformer at node 1 in the t period; r base is the basic transformation ratio, where r base =0.94; Δr k is the square increment of the gear ratio of gear k; θ IN,t and θ DE,t are the times the on-load tap-changing transformer is put into operation and removed in the t period, θ OLTC,t,k is the switching state of the on-load tap-changing transformer at the k-th gear in the t-th time period.

[0071] S120: Solve the second-order cone optimal power flow model through a preset solver to obtain an optimal solution for the multi-period operation of the distribution network.

[0072] In this embodiment, the second-order cone programming-based optimal power flow model (SOCP-OPF) can be solved by the CPLEX solver or the GUROBI solver to obtain the optimal solution for the multi-period operation of the distribution network. Specifically, the distribution network parameter data is obtained, brought into the second-order cone optimal power flow model, and the optimal solution for the multi-period operation of the distribution network is obtained by solving. The distribution network parameter data includes the objective function and the parameter data in the above-mentioned constraints. The distribution network parameter data can be historical data, so the obtained optimal solution for the multi-period operation of the distribution network can be a historical optimal solution. The optimal solution for the multi-period operation of the distribution network includes the photovoltaic output of each period, the energy storage charging and discharging power of each period, the reactive compensation equipment adjustment amount of each period, and the gear position of the on-load tap-changing transformer in each period.

[0073] S130: Converting the optimization problem of the second-order cone optimal power flow model into a Markov decision process to obtain a Markov decision model.

[0074] In this embodiment, the optimization problem is converted into a Markov decision process by obtaining the voltage amplitude, load, photovoltaic output, and energy storage capacity of each node, and constructing a state space based on the voltage amplitude, load, photovoltaic output, and energy storage capacity of each node. A dynamic space is constructed based on the distribution network parameter data to be solved in the second-order cone optimal power flow model, specifically, the state space is constructed based on the photovoltaic output adjustment, energy storage charging and discharging power, reactive compensation equipment adjustment, and the gear position of the on-load tap-changing transformer; a reward function is constructed based on the objective function in the second-order cone optimal power flow model; a state transition function is obtained based on the constraints in the second-order cone optimal power flow model; and a Markov decision process is constructed using the state space, action space, reward function, and state transition function to obtain a Markov decision model.

[0075] Specifically, the state space of the Markov decision model is:

[0076] s t =[|U t |,P t ,Q ct ,E t ]

[0077] Among them, s t is the state at time period t; |U t | is the set of voltage amplitudes of each node in time period t;

[0078] P t is the set of node loads in the t period; Q ctis the photovoltaic output of each node in the t period; E t is the energy storage capacity set of each node in the t period; where Q ct Satisfy the photovoltaic output constraint; E t satisfying the capacity dynamic constraint;

[0079] The dynamic space of the Markov decision model includes:

[0080]

[0081] Among them, a t Represents the choice of each action at time t;

[0082] in, The photovoltaic output adjustment, energy storage charging and discharging power, reactive compensation device adjustment, and on-load tap-changing transformer gear position are selected in time period t, respectively; the energy storage charging and discharging power satisfies the charging and discharging power constraint; the reactive compensation device adjustment satisfies the reactive compensation device constraint; and the on-load tap-changing transformer gear position satisfies the on-load tap-changing transformer constraint;

[0083] The reward function of the Markov decision model is:

[0084]

[0085] Among them, α, β, γ, and δ are weight coefficients respectively; V i,over (t) is the voltage over-limit value of the node i in the tth period, V i,over (t)=max(0,U i,t -U i,max )+max(0,U i,min -U i,t ); where -αP g (t) is the electricity purchase cost, r t is the immediate reward in period t; is the network loss cost, -γ∑V i,over (t) is the voltage over-limit penalty; δ·f1 is the carrying capacity reward. This reward function can guide the decision-making agent to comprehensively consider multiple factors during the decision-making process and make better decisions.

[0086] The state transition function of the Markov decision model is:

[0087]

[0088] Among them, P DG,i,t+1 is the photovoltaic output of the node i in the t+1 period; P pv,t+1 is the light intensity in the t+1 period; Uj,t+1 is the voltage of node j in the t+1th period; U i,t+1 is the voltage of the node i in the t+1th period; P ij,t+1 Q is the active power transmitted from the node i to the node j in the t+1th period; ij,t+1 I is the reactive power transmitted from the node i to the node j in the t+1 period; ij,t+1 is the current between the node i and the node j in the t+1th period.

[0089] S140: Determine a target state corresponding to the optimal solution for the multi-period operation of the distribution network based on the state space of the Markov decision model, and construct a pre-training data set based on the optimal solution for the multi-period operation of the distribution network and the corresponding target state.

[0090] In one implementation of this embodiment, determining the target state corresponding to the optimal solution for multi-period operation of the distribution network based on the state space of the Markov decision model includes: selecting target parameter data corresponding to state elements in the state space from the distribution network parameter data corresponding to the optimal solution for multi-period operation of the distribution network; and forming a state defined by the state space based on the target parameter data, and using the state as the target state corresponding to the optimal solution for multi-period operation of the distribution network. The distribution network parameter data corresponding to the optimal solution for multi-period operation of the distribution network may be the distribution network parameter data input to the second-order cone optimal power flow model when the optimal solution for multi-period operation of the distribution network is obtained using the second-order cone optimal power flow model. Specifically, the voltage amplitude, load, photovoltaic output, and energy storage capacity of each node may be selected from the distribution network parameter data to form a voltage amplitude set, load set, photovoltaic output set, and energy storage capacity set for each node. The voltage amplitude set, load set, photovoltaic output set, and energy storage capacity set of each node at this time are used to form a state defined by the state space, and used as the target state corresponding to the optimal solution for multi-period operation of the distribution network.

[0091] In this embodiment, the optimal solution of the distribution network operation in multiple time periods and the corresponding target states form a pre-training data set; wherein, in each time period, there is an optimal solution, and the optimal solution corresponds to a target state, and the pre-training data set is constructed based on the optimal solution of each time period and the corresponding target state.

[0092] S150: Pre-training the decision-making agent using the pre-training data set, and based on the Markov decision model, training the pre-trained decision-making agent using the algorithm in reinforcement learning, and outputting the optimal distribution network multi-period operation strategy based on the trained decision-making agent, wherein the optimal distribution network multi-period operation strategy is the optimal distribution network multi-period operation plan.

[0093] In this embodiment, a transfer learning initialization strategy is adopted to construct a pre-training data set using the optimal solution for multi-period operation of the distribution network generated by the second-order cone optimal power flow model and the corresponding target state, and the decision-making agent is pre-trained using the pre-training data set.

[0094] In this embodiment, real-time parameter data in the distribution network and the real-time state defined by the state space are collected and input into the trained decision-making agent to obtain the optimal multi-period operation plan of the distribution network.

[0095] In one embodiment, the pre-training data set is used to pre-train the decision agent, including: using the pre-training data set to pre-train the policy network in the decision agent. The decision agent can be in the distribution network dispatching center, and the environment can be a high-proportion distributed photovoltaic distribution network system. The decision agent can perceive the state of the environment, and the policy network in the decision agent is used to output actions. The decision agent can be a deep deterministic policy gradient algorithm agent or a proximal policy optimization algorithm agent. Specifically, determine the state s of the current period (for example, period t) t 、Action a t and instant rewards t , and determine the state s of the next period t+1 , the state s of the current period t 、Action a t and instant rewards t , and the state s of the next period t+1 As training samples, they are stored in the experience cache area. Through multiple iterations, a preset number of quantity samples are collected and stored in the experience cache area. The training samples in the experience cache area are used to train the pre-trained decision-making agent. Specifically, during the interaction between the decision-making agent and the environment, the agent takes actions on the environment according to a certain strategy. The environment feeds back the state and reward of the next time period after the action to the decision-making agent based on the action, and then starts the interaction process of the next time period and repeats it. It can also be understood that during the interaction between the decision-making agent and the environment, the agent takes actions on the environment according to a certain strategy. The environment feeds back the state and reward of the next time period after the action to the decision-making agent based on the action, and then starts the interaction process of the next time period and repeats it, so that the agent continues to learn and optimize until the optimal conditions are reached. The process of the technical solution provided in the embodiment of the present application can also be referred to. Figure 2 .

[0096] Specifically, Figure 3 This is a flowchart of training a pre-trained decision-making agent provided by an embodiment of the present application, such as Figure 2 As shown, the method includes:

[0097] S310: Acquire the state of the current period defined by the state space, and determine the action of the current period.

[0098] In this embodiment, the state of the current period includes the voltage amplitude set, load set, photovoltaic output set and energy storage capacity set of each node in the current period. For example, if the current period is t period, the state of the current period is s t =[|U t |,P t ,Q ct ,E t ].

[0099] In this embodiment, the action of the current period can be selected through the policy network.

[0100] S320: Determine an immediate reward for the current period according to the reward function, and determine a state for the next period based on the state transition function.

[0101] In this embodiment, the instantaneous reward r at the current moment can be determined by the reward function of the Markov decision model t The photovoltaic output, voltage amplitude and energy storage capacity of each node in the next period can be determined through the state transfer function. In the next period, the load of each node can be directly monitored and obtained, thus obtaining the state of the next period.

[0102] S330: The state of the current period, the action of the current period, the immediate reward of the current period, and the state of the next period are stored in an experience cache area as training samples.

[0103] In this embodiment, the state s of the current period is t 、Action a in the current period t 、The instant reward r of the current period t and the state s for the next period t+1 The formed set of data is used as a training sample and stored in the experience buffer area.

[0104] S340: Repeat the step of obtaining training samples, use the training samples in the experience cache area, and train the pre-trained decision-making agent based on the deep deterministic policy gradient algorithm or the proximal policy optimization algorithm.

[0105] In this embodiment, the state of the next period is used as the current state, and the process returns to the step of determining the action until a cutoff condition is met. The cutoff condition may be that the number of training samples meets a preset number, or the number of returns meets a preset number. The deep deterministic policy gradient algorithm or the proximal policy optimization algorithm is an algorithm used in reinforcement learning.

[0106] In this embodiment, the decision-making agent is exemplarily an agent based on a deep deterministic gradient algorithm, including the following:

[0107] Policy network (Actor network): Generates deterministic actions (device adjustment instructions).

[0108] Value Network (Critic Network): Evaluates the value of state-action pairs.

[0109] Target policy network (TargetActor network): used for stable training.

[0110] Target Critic Network: used to calculate the target Q value.

[0111] This example demonstrates the process of training a pre-trained decision-making agent using a deep deterministic gradient algorithm based on a Markov decision model. This algorithm addresses the instability and sample efficiency issues in reinforcement learning through experience replay and target network update mechanisms. The specific steps are as follows:

[0112] Step 1: Initialize network parameters.

[0113] (1) Initialize the value network parameters θ Q , that is, the initial parameters of the value network in the pre-trained decision agent are random parameters; the parameters of the pre-trained strategy network are used as the initial strategy network parameters θ of this training μ .

[0114] (2) Initialize the target strategy network parameters θ μ' and target value network parameter θ Q' , and set θ μ =θ μ' ,θ Q =θ Q' .

[0115] (3) Initialize the experience replay buffer R to store state transition samples (s t ,a t ,r t ,s t+1 ).

[0116] Step 2: Define the state space and action space of the Markov decision model.

[0117] (1) State space s t :s t =[|U t |,P t ,Q ct ,Et ].

[0118] (2) Action space a t :

[0119] Step 3: The policy network generates actions.

[0120] In the current period status s t Under this condition, the policy network μ(s t |θ μ ) Generate a deterministic action a t :

[0121]

[0122] in, Exploratory noise (such as Ornstein-Uhlenbeck noise) is used to increase exploratory power.

[0123] For example, the constraints of the action space can be embedded through the policy network. For example, physical constraints are introduced in the output layer of the policy network, such as the photovoltaic output limit: 0≤P DG,i, ≤P DG,i,max ; Energy storage charging and discharging power limit: 0≤p ch,i ≤u ch,i ·P ch,i,max OLTC gear discretization: mapping continuous output to actual gears (such as 12-gear transformer)

[0124] Step three: perform actions and observe environmental feedback.

[0125] (1) Execute action a in the distribution network environment t , observe the next state s t+1 and instant rewards t .

[0126] (2) The state transfer sample (s t ,a t ,r t , s t +1 ) is stored in the experience buffer area.

[0127] Step 5: Sampling from the experience buffer area.

[0128] (1) Randomly sample a batch of state transition samples (s i ,a i ,r i ,s i+1 ).

[0129] (2) Calculate the target Q value y i :

[0130] y i =r i +γQ'(s i+1 ,μ'(s i+1 |θ μ' )|θ Q' )

[0131] Among them, γ is the discount factor, Q' is the target value network, and μ' is the target policy network.

[0132] Step 6: Update the value network.

[0133] (1) Calculate the loss function of the value network:

[0134]

[0135] Among them, Q(s i ,a i |θ Q ) is the action function value approximated by the neural network, which can be understood as the true value of Q; i is the target Q value, that is, the value of Q calculated by the target value network.

[0136] (2) Update the value network parameters θ using gradient descent Q .

[0137] Step 7: Update the policy network.

[0138] Among them, the policy network can be a fully connected neural network, and the input is the state s i , the output is action a i .

[0139] The loss function of the policy network can be the mean square error (MSE), which minimizes the gap between the network output and the optimal solution of the distribution network multi-period operation output by the SOCP-OPF model:

[0140]

[0141] Optimization goal: Update the policy network parameters θ by gradient descent μ , so that it approaches the optimal solution of the multi-period operation of the distribution network output by the SOCP-OPF model. pre is the loss value of the loss function, M is the number of samples, a SOCP,i It is the optimal solution of multi-period operation of the i-th distribution network output by the SOCP-OPF model.

[0142] Specifically:

[0143] (1) Calculate the gradient of the policy network:

[0144]

[0145] (2) Use the gradient ascent method to update the policy network parameters θ μ .

[0146] Step 8: Update the target value network and target strategy network.

[0147] Use the soft update method to update the target network parameters:

[0148] θ μ' ←τθ μ +(1-τ)θ μ'

[0149] θ Q' ←τθ Q +(1-τ)θ Q'

[0150] Where τ is the soft update coefficient (usually 0.001).

[0151] Step nine: Repeat the training until convergence.

[0152] (1) Repeat steps 3 to 8 until the policy network and value network converge.

[0153] (2) During the training process, gradually reduce the exploration noise To increase the certainty of the strategy.

[0154] It should be noted that the distribution network can contain a high proportion of distributed photovoltaics. In addition, an action correction layer can be added inside or outside the decision-making agent. Specifically, when the strategy network outputs the original action a t Finally, add an action correction layer to ensure the legitimacy of the action through the following steps:

[0155] (1) Flow constraint check: Calculate the predicted state (voltage, line power) after the action is executed.

[0156] (2) Constraint violation detection: If the predicted state violates the safety threshold (Ui>1.05pu), the corrective action is taken.

[0157] (3) Projection correction: Project the illegal action to the nearest possible action space.

[0158] The technical solution provided in the embodiment of the present application constructs a second-order cone optimal power flow model, and obtains the optimal solution for the multi-period operation of the distribution network through the model, that is, obtains the distribution network operation optimization strategy, and obtains the Markov decision model by converting the optimization problem of the second-order cone optimal power flow model into a Markov decision process; the decision-making agent is pre-trained with the optimal solution for the multi-period operation of the distribution network, and based on the Markov decision model, the algorithm in reinforcement learning is used to train the pre-trained decision-making agent, and the trained decision-making agent outputs the optimal distribution network multi-period operation strategy, that is, in the Markov decision process in reinforcement learning, the pre-trained decision-making agent is trained with time-series environmental state data, and the trained decision-making agent obtains the optimal distribution network multi-period operation strategy through real-time state data; in summary, in the embodiment of the present application, the distribution network is obtained by constructing a second-order cone optimal power flow model for multi-periods. The optimal solution for multi-period operation of the power grid is obtained by sequentially training the decision-making agent through the optimal solution for multi-period operation of the distribution network and the time-series environmental state obtained based on the Markov decision model. This can optimize the access capacity of distributed photovoltaics in real time, improve the operating performance of the distribution network in dynamic scenarios, and effectively solve the problem of poor adaptability to dynamic scenarios. The optimal multi-period operation strategy of the distribution network is obtained by collaborative optimization of the second-order cone optimal power flow model and the reinforcement learning method, which can improve the distribution network's carrying capacity for distributed photovoltaics. The decision-making agent is pre-trained with the optimal solution for multi-period operation of the distribution network output by the second-order cone optimal power flow model, which can accelerate the convergence of the decision-making agent. The pre-trained decision-making agent is trained based on the Markov decision model and the algorithm in reinforcement learning. The trained decision-making agent can quickly make decisions, improve the overall optimization efficiency, and meet the time requirements of real-time scheduling of the distribution network.

[0159] In the related technologies, most of the existing research focuses on building a comprehensive evaluation model to assess the carrying capacity of distributed photovoltaics, but there are few in-depth studies on how to improve the carrying capacity. This application constructs a multi-period SOCP-OPF model for a distribution network with a high proportion of distributed photovoltaics to quantify the distributed photovoltaic carrying capacity of the distribution network; this application designs a dynamic adjustment strategy based on reinforcement learning (Markov decision process, state space, action space, reward function, state transfer function, collect real-time data to build training samples, use the algorithm in reinforcement learning to train the decision-making agent through the training samples, and make decisions through the decision-making agent) to optimize the operating status of distributed photovoltaic output, energy storage charging and discharging power, reactive compensation equipment and on-load tap-changing transformers in real time; through the collaborative optimization of the reinforcement learning method based on the Markov decision process and the SOCP-OPF model, the accuracy of the traditional optimization model and the dynamic adaptability of the reinforcement learning model are fully utilized, and the carrying capacity of the distribution network for distributed photovoltaics is effectively improved. It can effectively deal with the intermittent and random nature of distributed photovoltaic output and improve the operating stability and reliability of the distribution network under distributed photovoltaic access.

[0160] In the related art, most current studies are based on the optimization and scheduling of distribution networks under static time sections, ignoring the continuous time-series coupling characteristics of photovoltaic output and load, resulting in poor adaptability to dynamic scenarios. The embodiment of the present application fully considers the dynamic changes of photovoltaic output and load, and realizes real-time optimization of distributed photovoltaic access capacity by constructing a multi-period SOCP-OPF model and a dynamic adjustment strategy based on real-time data. In different time periods, the trained decision-making agent can dynamically adjust the operating status of the equipment according to real-time data such as photovoltaic output and load, thereby improving the operating performance of the distribution network in dynamic scenarios and effectively solving the problem of poor adaptability of the existing technology to dynamic scenarios.

[0161] In related technologies, traditional optimization methods such as second-order cone programming take tens of minutes to calculate, which cannot meet the minute-level scheduling requirements of the distribution network. This application uses a reinforcement learning method based on the Markov decision model in conjunction with the traditional SOCP-OPF model. The reinforcement learning method has a fast response speed and can make decisions quickly. At the same time, it uses a transfer learning initialization strategy and pre-trains the strategy network with the optimal solution output by SOCP-OPF to accelerate convergence, improve overall optimization efficiency, and meet the time requirements of real-time scheduling of the distribution network.

[0162] It should be noted that the embodiment of the present application utilizes digital twin simulation technology to build a high-fidelity simulation environment based on the actual topological structure of the regional distribution network, injecting random disturbances such as light fluctuations and load randomness into it, supporting large-scale parallel training, and enhancing the robustness of the model; building a virtual distribution network based on the IEEE 33-node system, including photovoltaic, energy storage, load, reactive compensation equipment, etc. Under the condition of high-proportion distributed photovoltaic access, the IEEE 33-node system successfully improved the distributed photovoltaic carrying capacity, reduced voltage fluctuations, and reduced operating costs, verifying the effectiveness and feasibility of the technical solution of the present invention. Among them, in terms of physical equation integration, the forward substitution method or the Newton-Raphson method can be used to solve the nonlinear power flow equation; based on the light intensity curve and the inverter efficiency curve, a photovoltaic output model is constructed, and the photovoltaic output data is obtained through the model; the charging and discharging efficiency and the state of charge (SOC) limit can be considered to construct a storage dynamic model, and the data corresponding to the model can be obtained; the embodiment of the present application injects random disturbances to simulate light fluctuations and load randomness, thereby enhancing the robustness of the model.

[0163] The solution provided by the embodiment of the present application effectively solves the problem of the dynamic adjustment strategy of reinforcement learning relying on a high-precision simulation environment and having poor noise interference resistance: although the dynamic optimization strategy of reinforcement learning responds quickly, it relies on a high-precision simulation environment, and the actual deployment is susceptible to noise interference. This application constructs a high-fidelity simulation environment for the distribution network to perform digital twin simulation, supports large-scale parallel training, and enhances the robustness of the model. At the same time, it embeds flow constraints in action selection, adopts a safe exploration mechanism, avoids invalid exploration, reduces the impact of noise on decision-making, and improves the applicability and reliability of the dynamic adjustment strategy of reinforcement learning in the actual distribution network environment.

[0164] Figure 4 This is a structural block diagram of a distribution network operation optimization device provided by an embodiment of the present application. Figure 4 As shown, the device comprises:

[0165] Model construction module 410 is used to construct an objective function with the goal of maximizing the total amount of distributed photovoltaic power connected to the distribution network and minimizing the total operating cost of the distribution network, and to construct a second-order cone optimal power flow model for the distribution network in multiple time periods based on the objective function, with power flow constraints, voltage and current safety constraints, photovoltaic output constraints, energy storage dynamic constraints, reactive compensation equipment constraints, and on-load tap-changing transformer constraints as constraints.

[0166] A model solving module 420 is configured to solve the second-order cone optimal power flow model using a preset solver to obtain an optimal solution for the multi-period operation of the distribution network;

[0167] a conversion module 430 for converting the optimization problem of the second-order cone optimal power flow model into a Markov decision process to obtain a Markov decision model;

[0168] A pre-training data set construction module 440 is used to determine the target state corresponding to the optimal solution for the multi-period operation of the distribution network based on the state space of the Markov decision model, and to construct a pre-training data set based on the optimal solution for the multi-period operation of the distribution network and the corresponding target state;

[0169] The training output module 450 is used to pre-train the decision-making agent using the pre-training data set, and based on the Markov decision model, use the reinforcement learning algorithm to train the pre-trained decision-making agent, and output the optimal distribution network multi-period operation strategy based on the trained decision-making agent, wherein the optimal distribution network multi-period operation strategy is the optimal distribution network multi-period operation plan.

[0170] In an optional embodiment, constructing an objective function with the goal of maximizing the total amount of distributed photovoltaic power connected to the distribution network and minimizing the total operating cost of the distribution network includes:

[0171] The objective function is determined based on the following formula:

[0172]

[0173] Where f1 is the total capacity of distributed photovoltaic connected to the distribution network; f2 is the total investment and operation cost; C DG,i is the distributed photovoltaic capacity accessible to node i in the distribution network; C grid The unit price of electricity purchased from the main grid; P g (t) is the power purchased by the main grid in period t; I l (t) is the current value of branch 1 in the distribution network; R l is the resistance of the branch 1, C loss is the unit cost of network loss, and N is the number of branches.

[0174] In an optional embodiment, the power flow constraint includes a node power constraint, a voltage equation, and a second-order cone constraint; wherein the node power constraint is formulated as follows:

[0175]

[0176] Among them, P in,i,t is the active power injected by node i in the distribution network during the tth period; P load,i,t is the load of the node i in the tth period; P g,i,t is the active power input from the main grid to the node i in the tth period; P DG,i,t is the distributed photovoltaic power generated at the node i in the t period; P dch,i,t is the energy storage system discharge power at the node i in the t period; P ch,,i,tis the charging power of the energy storage system at the node i in the t period; Q in,i,t is the reactive power injected into the node i in the tth period; Q load,i,t is the load reactive power of the node i in the tth period; Q g,i,t Q is the reactive power input from the main grid to the node i in the t period; SVC,i,t is the reactive power generated by the static VAR compensator at the node i in the tth period, Q CB,i,t is the reactive power generated by the group-switched capacitors at the node i in the tth period;

[0177] Wherein, the voltage equation is:

[0178]

[0179] The formula for the second-order cone constraint is:

[0180]

[0181] Among them, U j,t is the voltage of node j in the tth period; U i,t is the voltage of the node i in the tth period; R ij is the resistance between the node i and the node j; X ij is the reactance between the node i and the node j; P ij,t Q is the active power transmitted from node i to node j in time period t; ij,t I is the reactive power transmitted from the node i to the node j in the tth period; ij,t is the current between the node i and the node j in the tth period.

[0182] In an optional embodiment, the voltage and current safety constraint formula is:

[0183]

[0184] Among them, U i,min with U i,max are the upper and lower limits of the voltage amplitude of the node i, respectively. ij,max is the upper limit of the current amplitude between the node i and the node j;

[0185] The formula for the photovoltaic output constraint is:

[0186]

[0187] Among them, P DG,i,t With Q DG,i,tare the active and reactive power output of the distributed photovoltaic connected to the node i in the t period, i∈S N , S N is a set of all nodes in the distribution network;

[0188] Among them, P DG,i,max is the upper limit of the output active power of the distributed photovoltaic on the node i; Q DG,i,max is the upper limit of the output reactive power of the distributed photovoltaic at the node i;

[0189] The energy storage dynamic constraints include charge and discharge power constraints, capacity dynamic constraints, and capacity upper and lower limits and cycle constraints;

[0190] The formula for the charge and discharge power constraint is:

[0191]

[0192] Among them, u ch,i,t represents the charging state of energy storage system i at time period t; u dch,i,t represents the discharge state of the energy storage system i at the tth time period, where 1 represents execution and 0 represents non-execution; p ch,i,t represents the charging power of the energy storage system i in the t period; P ch,i,max represents the maximum charging power of the energy storage system i; P dch,i,t represents the discharge power of the energy storage system i in the tth period; P dch,i,max represents the maximum discharge power of the energy storage system i;

[0193] The formula for the dynamic capacity constraint is:

[0194]

[0195] Among them, E i,t and E i,t+1 are the energy storage capacity of the node i in the tth period and the t+1th period respectively; η ch,i is the charging efficiency of the node i, where η ch,i =0.9;η dch,i is the discharge efficiency of the node i, where η dcdh,i =0.9;

[0196] The formulas for capacity upper and lower limits and cycle constraints are:

[0197]

[0198] Among them, E i,1 =E i,T+1Indicates that the energy storage capacity of the node i in the first period and the T+1 period is the same, that is, it indicates that the energy storage capacity of the node i at the beginning and end of each day is the same; E i,min is the lower limit of the energy storage capacity of the node i; E i,max is the upper limit of the energy storage capacity of the node i;

[0199] The reactive compensation equipment constraints include capacitor-related constraints and static VAR compensator constraints;

[0200] The formula for the capacitor-related constraints is:

[0201]

[0202] Among them, Q CB,i,t , Q step,i are respectively the real-time total reactive power of the capacitor on the node i and the reactive power that can be provided by one gear of the capacitor in the t period; θ CB,i,t,k represents the gear position of the capacitor on the node i in the tth period; θ IN,i,t and θ DE,i,t are the number of times the capacitor on node i is put into operation and the number of times it is removed in time period t;

[0203] Among them, the formula for the static VAR compensator constraint is:

[0204] -Q SVC,max ≤Q SVC,t ≤Q SVC,max

[0205] Among them, Q SVC,t , Q SVC,max are the real-time reactive power and maximum reactive power of the static VAR compensator respectively;

[0206] The formula for the on-load tap-changing transformer constraint is:

[0207]

[0208] Among them, r 1,t is the transformation ratio of the on-load tap-changing transformer at node 1 in the t period; r base is the basic transformation ratio, where r base =0.94; Δr k is the square increment of the ratio of the kth gear; θ IN,t and θ DE,t are the times the on-load tap-changing transformer is put into operation and removed in the t period, θ OLTC,t,k is the switching state of the on-load tap-changing transformer at the k-th gear in the t-th time period.

[0209] In an optional embodiment, the state space of the Markov decision model is:

[0210] s t =[|U t |,P t ,Q ct ,E t ]

[0211] Among them, s t is the state at time period t; |U t | is the voltage amplitude set of each node in the tth period;

[0212] P t is the load set of each node in the t period; Q ct is the photovoltaic output of each node in the t period; E t is the energy storage capacity set of each node in the t period; where Q ct Satisfy the photovoltaic output constraint; E t satisfying the capacity dynamic constraint;

[0213] The dynamic space of the Markov decision model includes:

[0214]

[0215] Among them, a t Represents the choice of each action at time t;

[0216] in, The photovoltaic output adjustment, energy storage charging and discharging power, reactive power compensation device adjustment, and on-load tap-changing transformer gear position are selected in time period t, respectively; the energy storage charging power satisfies the charging and discharging power constraint; the reactive power compensation device adjustment satisfies the reactive power compensation device constraint; and the on-load tap-changing transformer gear position satisfies the on-load tap-changing transformer constraint;

[0217] The reward function of the Markov decision model is:

[0218]

[0219] Among them, α, β, γ, and δ are weight coefficients respectively; r t is the immediate reward in period t; V i,over (t) is the voltage over-limit value of the node i in the tth period, V i,over (t)=max(0,U i,t -U i,max )+max(0,U i,min -U i,t );

[0220] The state transition function of the Markov decision model is:

[0221]

[0222] Among them, P DG,i,t+1 is the photovoltaic output of the node i in the t+1 period; P pv,t+1 is the light intensity in the t+1 period; U j,t+1 is the voltage of node j in the t+1th period; U i,t+1 is the voltage of the node i in the t+1th period; P ij,t+1 Q is the active power transmitted from the node i to the node j in the t+1th period; ij,t+1 I is the reactive power transmitted from the node i to the node j in the t+1 period; ij,t+1 is the current between the node i and the node j in the t+1th period.

[0223] In an optional embodiment, determining the target state corresponding to the optimal solution for the multi-period operation of the distribution network based on the state space of the Markov decision model includes:

[0224] Selecting target parameter data corresponding to the state element in the state space from the distribution network parameter data corresponding to the optimal solution of the distribution network multi-period operation;

[0225] The state defined in the state space is formed based on the target parameter data and used as the target state corresponding to the optimal solution for the multi-period operation of the distribution network.

[0226] In an optional embodiment, the decision-making agent is an agent based on a deep deterministic policy gradient algorithm or a proximal policy optimization algorithm;

[0227] The pre-training of the decision agent using the pre-training data set includes:

[0228] Pre-training the policy network in the decision agent using the pre-training data set;

[0229] The method of training the pre-trained decision agent based on the Markov decision model using a reinforcement learning algorithm includes:

[0230] Obtaining the state of the current period defined by the state space and determining the action of the current period;

[0231] Determine the immediate reward for the current period based on the reward function, and determine the state for the next period based on the state transition function;

[0232] The state of the current period, the action of the current period, the immediate reward of the current period, and the state of the next period are stored as training samples in an experience buffer area;

[0233] Repeating the step of obtaining training samples, using the training samples in the experience buffer area to train the pre-trained decision agent based on the deep deterministic policy gradient algorithm or the proximal policy optimization algorithm;

[0234] In which, during the training process of the trained decision-making agent based on the deep deterministic policy gradient algorithm or the proximal policy optimization algorithm, the initial parameters of the value network in the pre-trained decision-making agent are random parameters.

[0235] An embodiment of the present application provides a distribution network operation optimization device that executes the method provided in an embodiment of the present application.

[0236] like Figure 5 As shown, an embodiment of the present application provides an electronic device, including a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.

[0237] Memory 113, for storing computer programs;

[0238] In one embodiment of the present application, the processor 111 is configured to execute a program stored in the memory 113 to implement the method provided by any of the aforementioned method embodiments, including:

[0239] An objective function is constructed with the goal of maximizing the total amount of distributed photovoltaic power connected to the distribution network and minimizing the total operating cost of the distribution network. A second-order cone optimal power flow model for the distribution network in multiple time periods is constructed based on the objective function, with power flow constraints, voltage and current safety constraints, photovoltaic output constraints, energy storage dynamic constraints, reactive compensation equipment constraints, and on-load tap-changing transformer constraints as constraints.

[0240] Solving the second-order cone optimal power flow model through a preset solver to obtain the optimal solution for the multi-period operation of the distribution network;

[0241] Converting the optimization problem of the second-order cone optimal power flow model into a Markov decision process to obtain a Markov decision model;

[0242] Determining a target state corresponding to the optimal solution for the multi-period operation of the distribution network based on the state space of the Markov decision model, and constructing a pre-training data set based on the optimal solution for the multi-period operation of the distribution network and the corresponding target state;

[0243] The pre-training data set is used to pre-train the decision-making agent, and based on the Markov decision model, the algorithm in reinforcement learning is used to train the pre-trained decision-making agent, and the optimal distribution network multi-period operation strategy is output based on the trained decision-making agent, wherein the optimal distribution network multi-period operation strategy is the optimal distribution network multi-period operation plan.

[0244] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method provided in any of the aforementioned method embodiments are implemented.

[0245] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0246] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, or of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiment.

[0247] The above embodiments are provided for illustrative purposes only and are not intended to limit the scope of implementation. Those skilled in the art will appreciate that other variations or modifications based on the above descriptions are possible. It is not necessary and impossible to provide an exhaustive list of all implementations. Obvious variations or modifications arising therefrom remain within the scope of protection of this application.

Claims

1. A distribution network operation optimization method, characterized in that: include: An objective function is constructed with the goal of maximizing the total amount of distributed photovoltaic power connected to the distribution network and minimizing the total operating cost of the distribution network. A second-order cone optimal power flow model for the distribution network in multiple time periods is constructed based on the objective function, with power flow constraints, voltage and current safety constraints, photovoltaic output constraints, energy storage dynamic constraints, reactive compensation equipment constraints, and on-load tap-changing transformer constraints as constraints. Solving the second-order cone optimal power flow model through a preset solver to obtain the optimal solution for the multi-period operation of the distribution network; Converting the optimization problem of the second-order cone optimal power flow model into a Markov decision process to obtain a Markov decision model; Determining a target state corresponding to the optimal solution for the multi-period operation of the distribution network based on the state space of the Markov decision model, and constructing a pre-training data set based on the optimal solution for the multi-period operation of the distribution network and the corresponding target state; The pre-training data set is used to pre-train the decision-making agent, and based on the Markov decision model, the algorithm in reinforcement learning is used to train the pre-trained decision-making agent, and the optimal distribution network multi-period operation strategy is output based on the trained decision-making agent.

2. The method according to claim 1, characterized in that The objective function is constructed with the goal of maximizing the total amount of distributed photovoltaic power connected to the distribution network and minimizing the total operating cost of the distribution network, including: The objective function is determined based on the following formula: Where f1 is the total capacity of distributed photovoltaic connected to the distribution network; f2 is the total investment and operation cost; C DG,i is the distributed photovoltaic capacity accessible to node i in the distribution network; C grid The unit price of electricity purchased from the main grid; P g (t) is the power purchased by the main grid in period t; I l (t) is the current value of branch 1 in the distribution network; R l is the resistance of the branch 1, C loss is the unit cost of network loss, and N is the number of branches.

3. The method according to claim 2, characterized in that The power flow constraint includes a node power constraint, a voltage equation, and a second-order cone constraint; wherein the node power constraint is formulated as follows: Among them, P in,i,t is the active power injected by node i in the distribution network during the tth period; P load,i,t is the load of the node i in the tth period; P g,i,t is the active power input from the main grid to the node i in the tth period; P DG,i,t is the distributed photovoltaic power generated at the node i in the t period; P dch,i,t is the energy storage system discharge power at the node i in the t period; P ch,,i,t is the charging power of the energy storage system at the node i in the t period; Q in,i,t is the reactive power injected into the node i in the tth period; Q load,i,t is the load reactive power of the node i in the tth period; Q g,i,t Q is the reactive power input from the main grid to the node i in the t period; SVC,i,t is the reactive power generated by the static VAR compensator at the node i in the tth period, Q CB,i,t is the reactive power generated by the group-switched capacitors at the node i in the tth period; Wherein, the voltage equation is: The formula for the second-order cone constraint is: Among them, U j,t is the voltage of node j in the tth period; U i,t is the voltage of the node i in the tth period; R ij is the resistance between the node i and the node j; X ij is the reactance between the node i and the node j; P ij,t Q is the active power transmitted from node i to node j in time period t; ij,t I is the reactive power transmitted from the node i to the node j in the tth period; ij,t is the current between the node i and the node j in the tth period.

4. The method according to claim 3, characterized in that The formula for the voltage and current safety constraint is: Among them, U i,min with U i,max are the upper and lower limits of the voltage amplitude of the node i, respectively. ij,max is the upper limit of the current amplitude between the node i and the node j; The formula for the photovoltaic output constraint is: Among them, P DG,i,t With Q DG,i,t are the active and reactive power output of the distributed photovoltaic connected to the node i in the t period, i∈S N , S N is a set of all nodes in the distribution network; Among them, P DG,i,max is the upper limit of the output active power of the distributed photovoltaic on the node i; Q DG,i,max is the upper limit of the output reactive power of the distributed photovoltaic at the node i; The energy storage dynamic constraints include charge and discharge power constraints, capacity dynamic constraints, and capacity upper and lower limits and cycle constraints; The formula for the charge and discharge power constraint is: Among them, u ch,i,t represents the charging state of energy storage system i at time period t; u dch,i,t represents the discharge state of the energy storage system i at the tth time period, where 1 represents execution and 0 represents non-execution; p ch,i,t represents the charging power of the energy storage system i in the t period; P ch,i,max represents the maximum charging power of the energy storage system i; p dch,i,t represents the discharge power of the energy storage system i in the tth period; P dch,i,max represents the maximum discharge power of the energy storage system i; The formula for the dynamic capacity constraint is: Among them, E i,t and E i,t+1 are the energy storage capacity of the node i in the tth period and the t+1th period respectively; η ch,i is the charging efficiency of the node i, where η ch,i =0.9;η dch,i is the discharge efficiency of the node i, where η dch,i =0.9; The formulas for capacity upper and lower limits and cycle constraints are: Among them, E i,1 =E i,T+1 Indicates that the energy storage capacity of the node i in the first period and the T+1 period is the same, that is, it indicates that the energy storage capacity of the node i at the beginning and end of each day is the same; E i,min is the lower limit of the energy storage capacity of the node i; E i,max is the upper limit of the energy storage capacity of the node i; The reactive compensation equipment constraints include capacitor-related constraints and static VAR compensator constraints; The formula for the capacitor-related constraints is: Among them, Q CB,i,y , Q step,i are respectively the real-time total reactive power of the capacitor on the node i and the reactive power that can be provided by one gear of the capacitor in the t period; θ CB,i,t,k represents the gear position of the capacitor on the node i in the tth period; θ IN,i,t and θ DE,i,t are the number of times the capacitor on node i is put into operation and the number of times it is removed in time period t; Among them, the formula for the static VAR compensator constraint is: -Q SVC,max ≤Q SVC,t ≤Q SVC,max Among them, Q SVC,t , Q SVC,max are the real-time reactive power and maximum reactive power of the static VAR compensator respectively; The formula for the on-load tap-changing transformer constraint is: Among them, r 1,t is the transformation ratio of the on-load tap-changing transformer at node 1 in the t period; r base is the basic transformation ratio, where r base =0.94; Δr k is the square increment of the ratio of the kth gear; θ IN,t and θ DE,t are the times the on-load tap-changing transformer is put into operation and removed in the t period, θ OLTC , t,k is the switching state of the on-load tap-changing transformer at the k-th gear in the t-th time period.

5. The method according to claim 4, characterized in that The state space of the Markov decision model is: s t =[|U t |,P t ,Q ct ,E t ] Among them, s t is the state at time period t; |U t | is the voltage amplitude set of each node in the tth period; P t is the load set of each node in the t period; Q ct is the photovoltaic output of each node in the t period; E t is the energy storage capacity set of each node in the t period; where Q ct Satisfy the photovoltaic output constraint; E t satisfying the capacity dynamic constraint; The dynamic space of the Markov decision model includes: Among them, a t Represents the choice of each action at time t; in, The photovoltaic output adjustment, energy storage charging and discharging power, reactive power compensation device adjustment, and on-load tap-changing transformer gear position are selected in time period t, respectively; the energy storage charging power satisfies the charging and discharging power constraint; the reactive power compensation device adjustment satisfies the reactive power compensation device constraint; and the on-load tap-changing transformer gear position satisfies the on-load tap-changing transformer constraint; The reward function of the Markov decision model is: Among them, α, β, γ, and δ are weight coefficients respectively; r t is the immediate reward in period t; V i,over (t) is the voltage over-limit value of the node i in the tth period, V i,over (t)=max(0,U i,t -U i,max )+max(0,U i,min -U i,t ); The state transition function of the Markov decision model is: Among them, P DG,i,t+1 is the photovoltaic output of the node i in the t+1 period; P pv,t+1 is the light intensity in the t+1 period; U j,t+1 is the voltage of node j in the t+1th period; U i,t+1 is the voltage of the node i in the t+1th period; P ij,t+1 Q is the active power transmitted from the node i to the node j in the t+1th period; ij,t+1 I is the reactive power transmitted from the node i to the node j in the t+1 period; ij,t+1 is the current between the node i and the node j in the t+1th period.

6. The method according to claim 5, characterized in that The determining of the target state corresponding to the optimal solution for the multi-period operation of the distribution network based on the state space of the Markov decision model includes: Selecting target parameter data corresponding to the state element in the state space from the distribution network parameter data corresponding to the optimal solution of the distribution network multi-period operation; The state defined in the state space is formed based on the target parameter data and used as the target state corresponding to the optimal solution for the multi-period operation of the distribution network.

7. The method according to claim 5, characterized in that The decision-making agent is an agent based on a deep deterministic policy gradient algorithm or a proximal policy optimization algorithm; The pre-training of the decision agent using the pre-training data set includes: Pre-training the policy network in the decision agent using the pre-training data set; The method of training the pre-trained decision agent based on the Markov decision model using a reinforcement learning algorithm includes: Obtaining the state of the current period defined by the state space and determining the action of the current period; Determine the immediate reward for the current period based on the reward function, and determine the state for the next period based on the state transition function; The state of the current period, the action of the current period, the immediate reward of the current period, and the state of the next period are stored as training samples in an experience buffer area; Repeating the step of obtaining training samples, using the training samples in the experience buffer area to train the pre-trained decision agent based on the deep deterministic policy gradient algorithm or the proximal policy optimization algorithm; In which, during the training process of the trained decision-making agent based on the deep deterministic policy gradient algorithm or the proximal policy optimization algorithm, the initial parameters of the value network in the pre-trained decision-making agent are random parameters.

8. A distribution network operation optimization device, characterized in that: include: A model construction module is used to construct an objective function with the goal of maximizing the total amount of distributed photovoltaic power connected to the distribution network and minimizing the total operating cost of the distribution network. The module uses power flow constraints, voltage and current safety constraints, photovoltaic output constraints, energy storage dynamic constraints, reactive compensation equipment constraints, and on-load tap-changing transformer constraints as constraints, and combines the objective function to construct a multi-period second-order cone optimal power flow model for the distribution network. A model solving module, configured to solve the second-order cone optimal power flow model using a preset solver to obtain an optimal solution for the multi-period operation of the distribution network; a conversion module, configured to convert the optimization problem of the second-order cone optimal power flow model into a Markov decision process to obtain a Markov decision model; A pre-training data set construction module is used to determine the target state corresponding to the optimal solution for the multi-period operation of the distribution network based on the state space of the Markov decision model, and to construct a pre-training data set based on the optimal solution for the multi-period operation of the distribution network and the corresponding target state; A training output module is used to pre-train the decision-making agent using the pre-training data set, and based on the Markov decision model, use the algorithm in reinforcement learning to train the pre-trained decision-making agent, and output the optimal distribution network multi-period operation strategy based on the trained decision-making agent.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed in a computer, the computer is caused to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Power distribution network optimization scheduling method and system based on multi-target deep reinforcement learning

    CN121906489A

  • User-side-oriented multi-target decision execution method and system

    CN122022400A

  • User-side multi-target decision execution method and system

    CN122022400B