Reinforcement learning-based reactive power and voltage optimization method, device, equipment and medium
Patent Information
- Application Number
- CN202211593877.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-12
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2042-12-12
AI Technical Summary
[0004]本申请提供了一种基于强化学习的无功电压优化方法、装置、设备及介质,以解决基于模型驱动的传统优化方法难以获得高精度模型使得决策精度低,导致难以适应业务需求的技术问题
[0043] By utilizing deep learning algorithms, a pre-defined deep learning optimizer is trained using historical operating data of the actual distribution network as input and reactive power and voltage optimization data obtained based on the nominal model as training labels. This yields a first strategy model, where the historical operating data includes generator active power, load active power, and load reactive power. The optimization results of the nominal model provide a reference for deep reinforcement learning. A reactive power optimization agent is generated based on the first strategy model using a Markov decision process, transforming the distribution network reactive power optimization problem into a Markov decision process and generating a user-reinforcement-learned reactive power optimization agent. Finally, the reactive power optimization agent interacts with the actual distribution network in real time to optimize reactive power and voltage. The reactive power optimization agent is updated using reinforcement learning algorithms, thereby gradually reducing the optimization intensity of the nominal model by leveraging the improved reactive power optimization capability of the reinforcement learning agent, thus eliminating dependence on the distribution network model and improving reactive power optimization accuracy.
Smart Images

Figure CN115833147B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voltage control technology, and in particular to a method, apparatus, device and medium for reactive voltage optimization based on reinforcement learning. Background Technology
[0002] As more and more distributed generation sources are connected to digital distribution networks, the high penetration rate of distributed generation sources can cause voltage fluctuations or excessively high voltages leading to their disconnection from the grid. This severely restricts the ability of active distribution networks to absorb renewable energy generation, wasting grid resources and renewable energy. To address this, active distribution networks utilize reactive power and voltage control algorithms to reduce grid losses and improve voltage by controlling adjustable reactive power resources.
[0003] Currently, traditional distribution network voltage control algorithms are mainly model-driven, including centralized and distributed algorithms. Centralized algorithms require real-time acquisition of grid state information and are easily affected by communication quality. Distributed algorithms typically do not require communication with neighbors, but their optimization performance heavily depends on the accuracy of the distribution network model and parameters. Active distribution networks possess complex characteristics such as high nonlinearity, heterogeneity, and time-varying nature, as well as factors like large network scale and sparse measurements, leading to significant uncertainties in model parameters. This makes it difficult to obtain high-precision distribution network models in practical applications. Therefore, traditional model-driven optimization control methods are ill-suited to business requirements. Summary of the Invention
[0004] This application provides a reactive voltage optimization method, apparatus, device, and medium based on reinforcement learning to solve the technical problem that traditional model-driven optimization methods are difficult to obtain high-precision models, resulting in low decision accuracy and difficulty in adapting to business needs.
[0005] To address the aforementioned technical problems, firstly, this application provides a reactive power voltage optimization method based on reinforcement learning, comprising:
[0006] Using a deep learning algorithm, a preset deep learning optimizer is trained with historical operating data of the actual distribution network as input and reactive voltage optimization data obtained based on the nominal model as training labels to obtain the first strategy model. The historical operating data includes generator active power, load active power and load reactive power.
[0007] Using Markov decision process, a reactive power optimization agent is generated based on the first policy model;
[0008] The reactive power optimization agent interacts with the actual power distribution network in real time to optimize the reactive power and voltage of the actual power distribution network, and uses reinforcement learning algorithms to update the reactive power optimization agent.
[0009] In some implementations, the use of deep learning algorithms, taking historical operating data of the actual distribution network as input and reactive power and voltage optimization data obtained based on the nominal model as training labels, trains a preset deep learning optimizer to obtain a first strategy model, including:
[0010] Using the nominal model, the reactive power voltage optimization data is output based on the historical operating condition data;
[0011] Using the preset deep learning optimizer, reactive power voltage control data is output based on the historical operating condition data;
[0012] Based on the reactive voltage optimization data and the reactive voltage control data, calculate the first loss function of the preset deep learning optimizer;
[0013] The preset deep learning optimizer is updated based on the first loss function, and it is determined whether the preset deep learning optimizer has reached the convergence condition.
[0014] If the first loss function reaches its minimum value, then the preset deep learning optimizer is determined to have reached the convergence condition, and the first policy model is obtained.
[0015] In some implementations, the step of using the nominal model to output the reactive power voltage optimization data based on the historical operating condition data includes:
[0016] Using the nominal model and based on the historical operating data, power flow analysis is performed on the actual distribution network to output the reactive power and voltage optimization data. The nominal model is:
[0017]
[0018] Where, r p (x t ,u t x represents grid loss or power generation cost. t u is the dependent variable. t To control the variable, D t Let b be the disturbance variable containing the historical operating data, A be the model parameter of the active distribution network model, g be the topology of the active distribution network model, and h be the power flow equation. v The inequality constraint equations represent the relationship between voltage and control variables.
[0019] In some implementations, the step of using a Markov decision process to generate a reactive power optimization agent based on the first policy model includes:
[0020] In the Markov decision-making process, a preset reinforcement learning agent observes the first state information of the actual distribution network at the current moment. The first state information includes node injected active power, node injected reactive power, node voltage, and reactive output power.
[0021] Using the first policy model, select the first action information corresponding to the first state information, and calculate the first reward information and the state information at the next moment of observation of the preset reinforcement learning agent to generate the reactive power optimization agent.
[0022] In some implementations, the step of real-time interaction between the reactive power optimization agent and the actual distribution network to optimize reactive power and voltage in the actual distribution network, and updating the reactive power optimization agent using a reinforcement learning algorithm, includes:
[0023] Based on the model parameters of the first strategy model, a second strategy model is generated, and two preset critic networks and a data buffer are initialized.
[0024] A target strategy model is generated based on the second strategy model, and two target critic networks are generated based on the two critic networks.
[0025] If the amount of data in the data buffer is less than the preset amount of data, then based on the observation of the second state information of the actual distribution network at the current moment by the reactive power optimization agent, and according to the target strategy model, the second action information corresponding to the second state information is selected to optimize the reactive power voltage of the actual distribution network, and the data buffer is updated based on the second state information and the second action information.
[0026] If the amount of data in the data buffer is not less than the preset amount of data, then based on the reactive power optimization agent's observation of the third state information of the actual distribution network at the current moment, and selecting the deterministic action information corresponding to the third state information, reactive power voltage optimization is performed on the actual distribution network, and the target policy model is updated based on the data buffer using the target critic network.
[0027] In some implementations, selecting the deterministic action information corresponding to the third state information includes:
[0028] Based on a preset strategy function, the deterministic action information corresponding to the third state information is selected. The preset strategy function is:
[0029]
[0030] Where 'a' represents deterministic action information. Let be the training neural network policy, ∈ be the exploration noise, typically a small Gaussian noise, and a LOW a is the minimum adjustable capability of a reactive power adjustable device. High This represents the maximum adjustable capacity of the reactive power adjustable device.
[0031] In some implementations, updating the target policy model using the target critic network based on the data buffer includes:
[0032] Randomly select several sets of sample data from the data buffer;
[0033] Based on the sample data, calculate the objective value of the target critic network function;
[0034] Based on the objective value of the function, calculate the second loss function of the target critic network and the third loss function of the target policy model;
[0035] Based on the second and third loss functions, the regularization coefficients, the target critic network, and the target policy model are updated.
[0036] Secondly, this application also provides a reactive power voltage optimization device based on reinforcement learning, comprising:
[0037] The training module is used to train a preset deep learning optimizer using a deep learning algorithm, taking historical operating data of the actual distribution network as input and reactive voltage optimization data obtained based on the nominal model as training labels, to obtain the first strategy model. The historical operating data includes generator active power, load active power and load reactive power.
[0038] The generation module is used to generate a reactive power optimization agent based on the first policy model using a Markov decision process.
[0039] An optimization module is used to perform real-time interaction between the reactive power optimization agent and the actual distribution network to optimize the reactive power and voltage of the actual distribution network, and to update the reactive power optimization agent using a reinforcement learning algorithm.
[0040] Thirdly, this application also provides a computer device, including a processor and a memory, the memory being used to store a computer program, which, when executed by the processor, implements the reactive voltage optimization method based on reinforcement learning as described in the first aspect.
[0041] Fourthly, this application also provides a computer-readable storage medium, characterized in that it stores a computer program, which, when executed by a processor, implements the reactive voltage optimization method based on reinforcement learning as described in the second aspect.
[0042] Compared with the prior art, this application has at least the following beneficial effects:
[0043] By utilizing deep learning algorithms, a pre-defined deep learning optimizer is trained using historical operating data of the actual distribution network as input and reactive power and voltage optimization data obtained based on the nominal model as training labels. This yields a first strategy model, where the historical operating data includes generator active power, load active power, and load reactive power. The optimization results of the nominal model provide a reference for deep reinforcement learning. A reactive power optimization agent is generated based on the first strategy model using a Markov decision process, transforming the distribution network reactive power optimization problem into a Markov decision process and generating a user-reinforcement-learned reactive power optimization agent. Finally, the reactive power optimization agent interacts with the actual distribution network in real time to optimize reactive power and voltage. The reactive power optimization agent is updated using reinforcement learning algorithms, thereby gradually reducing the optimization intensity of the nominal model by leveraging the improved reactive power optimization capability of the reinforcement learning agent, thus eliminating dependence on the distribution network model and improving reactive power optimization accuracy.
[0044] Furthermore, this application utilizes the optimization results of a reactive power optimization algorithm based on an imprecise model to reduce the learning cost of reinforcement learning, assist reinforcement learning decision-making, improve the convergence speed and accuracy of the reinforcement learning algorithm, and ultimately achieve reactive power optimization results superior to those based on nominal model optimization and reinforcement learning methods. In the initial stage of reinforcement learning training, the strategy obtained from the training results based on the nominal model optimization is used. Initialize the policy model to ensure that, in the initial M steps, the reinforcement learning agent has and The same reactive power optimization capability ensures that the reinforcement learning agent achieves good results in reducing network losses and improving voltage quality in actual power grids. This is achieved through two strategies. and regular terms The updates are restricted to avoid the negative impact of an inaccurate critic network in the early stages of training on the policy model, thus ensuring the stability of learning. At the same time, this regularization term can prevent policy... The resulting overfitting problem. Attached Figure Description
[0045] Figure 1 This is a flowchart illustrating a reinforcement learning-based reactive voltage optimization method according to an embodiment of this application.
[0046] Figure 2 This is a schematic diagram of the structure of a reactive voltage optimization device based on reinforcement learning, as shown in an embodiment of this application.
[0047] Figure 3This is a schematic diagram of the structure of a computer device shown in an embodiment of this application. Detailed Implementation
[0048] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0049] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a reinforcement learning-based reactive power optimization method provided in an embodiment of this application. The reinforcement learning-based reactive power optimization method of this application can be applied to computer devices, including but not limited to smartphones, laptops, tablets, desktop computers, physical servers, and cloud servers. Figure 1 As shown, the reactive power voltage optimization method based on reinforcement learning in this embodiment includes steps S101 to S103, which are detailed below:
[0050] Step S101: Using a deep learning algorithm, a preset deep learning optimizer is trained with historical operating data of the actual distribution network as input and reactive voltage optimization data obtained based on the nominal model as training labels to obtain the first strategy model. The historical operating data includes generator active power, load active power and load reactive power.
[0051] In this step, the nominal model is an imprecise active distribution network model. Using the optimization results of the nominal model to train the policy model can reduce the learning cost of deep reinforcement learning.
[0052] In some embodiments, step S101 includes:
[0053] Using the nominal model, the reactive power voltage optimization data is output based on the historical operating condition data;
[0054] Using the preset deep learning optimizer, reactive power voltage control data is output based on the historical operating condition data;
[0055] Based on the reactive voltage optimization data and the reactive voltage control data, calculate the first loss function of the preset deep learning optimizer;
[0056] The preset deep learning optimizer is updated based on the first loss function, and it is determined whether the preset deep learning optimizer has reached the convergence condition.
[0057] If the first loss function reaches its minimum value, then the preset deep learning optimizer is determined to have reached the convergence condition, and the first policy model is obtained.
[0058] In this embodiment, the reactive power-voltage optimization based on the model is essentially an optimal power flow problem, which can be simplified into a constrained optimization problem, i.e., the nominal model is:
[0059]
[0060] Where, r p (x t ,u t x represents grid loss or power generation cost. t The dependent variable includes active power injection P. t Reactive power injection Q t and voltage amplitude V t ;u t For reactive power optimization problems, the control variable refers to the reactive power generated by static var generators and inverter-based distributed power sources; D t The disturbance variable that includes the historical operating data is also called the uncontrollable variable, namely the active power generation P of distributed energy resources. G,t and electrical load P D,t Q D,t b represents the model parameters of the active distribution network model, A represents the topology of the active distribution network model; g represents the power flow equations, such as resistance and reactance; h v The inequality constraint equations represent the relationship between voltage and control variables.
[0061] It should be noted that when the power flow model is inaccurate, the result u obtained from the model optimization solution... t In real-world power distribution network environments, achieving satisfactory results is often difficult and may even lead to voltage safety issues. However, this optimization result... t It has some reference value and can provide a reference for reinforcement learning.
[0062] Optionally, since optimization based on the nominal model requires a large number of optimization iterations and is time-consuming, it is not suitable for real-time reactive power and voltage control optimization. Therefore, this embodiment performs Monte Carlo sampling on historical load and distributed generation data to generate a large amount of operating condition data P. D Q D P G These data are then fed into a reactive power voltage optimization algorithm based on a nominal model to solve for the corresponding controllable reactive power output u. m , that is u t Then, a policy model π is trained using a deep learning algorithm. θ It is used to mimic optimization strategies based on nominal models. Let π be a deep neural network, where θ represents the learnable parameters of that network. θ Input [P] D Q D P G The output is the predicted controllable reactive power output. The training objective of a deep neural network is to minimize the loss function:
[0063]
[0064] Furthermore, by selecting appropriate neural network hyperparameters such as network structure, learning rate, optimizer, and regularization term, the neural network is trained to be suitable for the dataset {[P}. D Q D P G ], u m It has good predictive ability.
[0065] Step S102: Using a Markov decision process, generate a reactive power optimization agent based on the first strategy model.
[0066] In this step, the reactive power optimization agent belongs to the reinforcement learning agent. In this embodiment, the reactive power optimization problem of the distribution network is modeled as a Markov decision process.
[0067] In some embodiments, step S101 includes:
[0068] In the Markov decision-making process, a preset reinforcement learning agent observes the first state information of the actual distribution network at the current moment. The first state information includes node injected active power, node injected reactive power, node voltage, and reactive output power.
[0069] Using the first policy model, select the first action information corresponding to the first state information, and calculate the first reward information and the state information at the next moment of observation of the preset reinforcement learning agent to generate the reactive power optimization agent.
[0070] In this embodiment, during the Markov decision process, at each time step t, the agent observes the state s and selects the action a corresponding to the policy π(s) to obtain the reward r and a new state s′ of the environment. The Markov decision process can be stored as a tuple (s t a t s t+1 r t d t ). d t s t+1 Is it in a terminated state? The unlimited range of discounted gains R is defined as the sum of discounted gains from all rewards received by the agent. Where γ∈(0,1) is the discount factor that determines the long-term reward priority. For the reactive voltage control problem, the corresponding state space, action space, and reward function are defined as follows:
[0071] Action: For reactive power optimization problems, the action is the reactive power output 'a' of all inverter-based reactive power adjustable devices. t = G,t,i ,in i is the sequence number of the inverter-based distributed power source, S G,i For capacity, It is the active power capacity of the distributed power source of the inverter.
[0072] Status: Set in this invention Where P t Q t V t Inject active power, reactive power, and node voltage into the node; T represents transpose.
[0073] Rewards: The rewards include two items: active power loss reward r p,t Voltage over-limit reward r v,t r p,t The sum of active power injections is calculated as follows:
[0074] Where N is the number of buses.
[0075] Voltage over-limit reward r v,t for:
[0076]
[0077] Step S103: Based on the real-time interaction between the reactive power optimization agent and the actual distribution network, the reactive power voltage of the actual distribution network is optimized, and the reactive power optimization agent is updated using a reinforcement learning algorithm.
[0078] In this step, the algorithm interacts online with a real power grid model to generate data for training a deep reinforcement learning agent. This algorithm is compatible with existing reinforcement learning algorithms such as Trust Domain Policy Optimization, Proximal Policy Optimization, Deep Deterministic Policy Gradient, Two-Delay DDPG (TD3), and Actor-Critic.
[0079] In some embodiments, step S103 includes:
[0080] Based on the model parameters of the first strategy model, a second strategy model is generated, and two preset critic networks and a data buffer are initialized.
[0081] A target strategy model is generated based on the second strategy model, and two target critic networks are generated based on the two critic networks.
[0082] If the amount of data in the data buffer is less than the preset amount of data, then based on the observation of the second state information of the actual distribution network at the current moment by the reactive power optimization agent, and according to the target strategy model, the second action information corresponding to the second state information is selected to optimize the reactive power voltage of the actual distribution network, and the data buffer is updated based on the second state information and the second action information.
[0083] If the amount of data in the data buffer is not less than the preset amount of data, then based on the reactive power optimization agent's observation of the third state information of the actual distribution network at the current moment, and selecting the deterministic action information corresponding to the third state information, reactive power voltage optimization is performed on the actual distribution network, and the target policy model is updated based on the data buffer using the target critic network.
[0084] In this embodiment, optionally, selecting the deterministic action information corresponding to the third state information includes: selecting the deterministic action information corresponding to the third state information based on a preset strategy function, wherein the preset strategy function is:
[0085]
[0086] Where 'a' represents deterministic action information. Let be the training neural network policy, ∈ be the exploration noise, typically a small Gaussian noise, and a LOW a is the minimum adjustable capability of a reactive power adjustable device. High This represents the maximum adjustable capacity of the reactive power adjustable device.
[0087] Optionally, updating the target policy model using the target critic network based on the data buffer includes: randomly selecting several sets of sample data from the data buffer;
[0088] Based on the sample data, calculate the objective function value of the target critic network; based on the objective function value, calculate the second loss function of the target critic network and the third loss function of the target policy model; based on the second loss function and the third loss function, update the regularization coefficients, the target critic network, and the target policy model.
[0089] For example, consider the dual-delay depth-deterministic policy gradient (TD3):
[0090] 1. Define a strategy model and two critic networks It also initializes its parameters and defines the data buffer. And initialize it. The policy trained in step 2... Parameters copied In the equation θ2←θ1, the canonical coefficient λ is defined as the decay rate λ1 of the canonical coefficient. λ1 is usually a number close to 1, such as 0.9999.
[0091] 2. Define the target policy model and the target critic network, and then apply the policy model π θ and critics network The parameters are copied to the target policy model and the critic network θ targ ←θ2, φ targ,1 ←φ1, φ targ,2 ←φ2.
[0092] 3. Repeat the following steps until convergence:
[0093] a. If the buffer contains less than M data items: Select an action based on the state s observed by the deep reinforcement learning agent. Execute action 'a' in the power grid environment, observe the next state 's', reward 'r', and store {s, a, r, s'} into the data buffer. middle.
[0094] b. If the buffer contains more than or equal to M data items:
[0095] i. Select an action based on the state s observed by the deep reinforcement learning agent. Where ∈ represents exploration noise, often a relatively small Gaussian noise. Action a is performed in the power grid environment, the next state s′ is observed, reward r is given, and {s, a, r, s′} is stored in the data buffer. middle.
[0096] ii. From the data buffer Randomly select B sets of data {s, a, r, s′}.
[0097] iii. Calculate the objective value of the critic function:
[0098]
[0099] in, Let c be Gaussian noise with variance σ, and c be the upper and lower bounds of the exploration noise, which is generally a constant less than 1, such as 0.2.
[0100] iv. Update the critic network by minimizing the loss function: for i = 1, 2.
[0101] v. Update the policy model by maximizing the loss function:
[0102] vi. Update coefficient λ = λ1 * λ.
[0103] vii. Update the policy model, the target network of the critic network:
[0104] φ targ,i ←ρφ targ,i +(1-ρ)φ i for i = 1, 2
[0105] θ targ ←ρθ targ +(1-ρ)θ2.
[0106] It should be noted that this application optimizes the strategy obtained from the training results based on the nominal model through steps 1 and 3.b).v. Integrate into deep reinforcement learning algorithms. Step 1 uses a strategy obtained from optimizing the training results based on the nominal model. Initializing the policy model ensures that, in the initial M steps, the reinforcement learning agent possesses and The same reactive power optimization capability ensures that the reinforcement learning agent will not cause significant network losses or voltage overruns in the real power grid environment.
[0107] When deep reinforcement learning algorithms begin learning, the critic network has a large error. At this point, using an inaccurate critic network to model the policy is problematic. To mitigate this problem, step 3.b).v. involves two strategies to perform an error update. and regular terms Restrict updates to ensure Will not deviate Too far. During training, the accuracy of the critic network gradually improves, step 3.b).vi. gradually decays as the decay rate λ1 decreases the value of λ.
[0108] To implement the reinforcement learning-based reactive voltage optimization method corresponding to the above method embodiments, and to achieve the corresponding functions and technical effects, see [link to relevant documentation]. Figure 2 , Figure 2 This diagram illustrates a structural block diagram of a reactive power voltage optimization device based on reinforcement learning, according to an embodiment of this application. For ease of explanation, only the parts relevant to this embodiment are shown. The reactive power voltage optimization device based on reinforcement learning provided in this embodiment includes:
[0109] Training module 201 is used to train a preset deep learning optimizer using a deep learning algorithm, taking historical operating data of the actual distribution network as input and reactive voltage optimization data obtained based on the nominal model as training labels, to obtain a first strategy model. The historical operating data includes generator active power, load active power and load reactive power.
[0110] The generation module 202 is used to generate a reactive power optimization agent based on the first strategy model using a Markov decision process.
[0111] The optimization module 203 is used to perform real-time interaction between the reactive power optimization agent and the actual distribution network to optimize the reactive power voltage of the actual distribution network, and to update the reactive power optimization agent using a reinforcement learning algorithm.
[0112] In some embodiments, the training module 201 is specifically used for:
[0113] Using the nominal model, the reactive power voltage optimization data is output based on the historical operating condition data;
[0114] Using the preset deep learning optimizer, reactive power voltage control data is output based on the historical operating condition data;
[0115] Based on the reactive voltage optimization data and the reactive voltage control data, calculate the first loss function of the preset deep learning optimizer;
[0116] The preset deep learning optimizer is updated based on the first loss function, and it is determined whether the preset deep learning optimizer has reached the convergence condition.
[0117] If the first loss function reaches its minimum value, then the preset deep learning optimizer is determined to have reached the convergence condition, and the first policy model is obtained.
[0118] In some embodiments, the nominal model is:
[0119]
[0120] Where, r p (x t ,u t x represents grid loss or power generation cost. t u is the dependent variable. t To control the variable, D t Let b be the disturbance variable containing the historical operating data, A be the model parameter of the active distribution network model, g be the topology of the active distribution network model, and h be the power flow equation. v The inequality constraint equations represent the relationship between voltage and control variables.
[0121] In some embodiments, the generation module 202 is specifically used for:
[0122] In the Markov decision-making process, a preset reinforcement learning agent observes the first state information of the actual distribution network at the current moment. The first state information includes node injected active power, node injected reactive power, node voltage, and reactive output power.
[0123] Using the first policy model, select the first action information corresponding to the first state information, and calculate the first reward information and the state information at the next moment of observation of the preset reinforcement learning agent to generate the reactive power optimization agent.
[0124] In some embodiments, the optimization module 203 is specifically used for:
[0125] Based on the model parameters of the first strategy model, a second strategy model is generated, and two preset critic networks and a data buffer are initialized.
[0126] A target strategy model is generated based on the second strategy model, and two target critic networks are generated based on the two critic networks.
[0127] If the amount of data in the data buffer is less than the preset amount of data, then based on the observation of the second state information of the actual distribution network at the current moment by the reactive power optimization agent, and according to the target strategy model, the second action information corresponding to the second state information is selected to optimize the reactive power voltage of the actual distribution network, and the data buffer is updated based on the second state information and the second action information.
[0128] If the amount of data in the data buffer is not less than the preset amount of data, then based on the reactive power optimization agent's observation of the third state information of the actual distribution network at the current moment, and selecting the deterministic action information corresponding to the third state information, reactive power voltage optimization is performed on the actual distribution network, and the target policy model is updated based on the data buffer using the target critic network.
[0129] In some embodiments, the optimization module 203 is further configured to:
[0130] Based on a preset strategy function, the deterministic action information corresponding to the third state information is selected. The preset strategy function is:
[0131]
[0132] Where 'a' represents deterministic action information. Let be the training neural network policy, ∈ be the exploration noise, typically a small Gaussian noise, and a LOWa is the minimum adjustable capability of a reactive power adjustable device. High This represents the maximum adjustable capacity of the reactive power adjustable device.
[0133] In some embodiments, the optimization module 203 is further configured to:
[0134] Randomly select several sets of sample data from the data buffer;
[0135] Based on the sample data, calculate the objective value of the target critic network function;
[0136] Based on the objective value of the function, calculate the second loss function of the target critic network and the third loss function of the target policy model;
[0137] Based on the second and third loss functions, the regularization coefficients, the target critic network, and the target policy model are updated.
[0138] The aforementioned reactive power voltage optimization device based on reinforcement learning can implement the reactive power voltage optimization method based on reinforcement learning in the above method embodiments. The options in the above method embodiments are also applicable to this embodiment, and will not be detailed here. The remaining content of this application's embodiments can be referred to the content of the above method embodiments, and will not be repeated in this embodiment.
[0139] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Figure 3 As shown, the computer device 3 of this embodiment includes: at least one processor 30 ( Figure 3 (Only one is shown in the diagram), memory 31, and computer program 32 stored in said memory 31 and executable on said at least one processor 30, wherein said processor 30 executes said computer program 32 to implement the steps in any of the above method embodiments.
[0140] The computer device 3 can be a smartphone, tablet, desktop computer, cloud server, or other computing device. This computer device may include, but is not limited to, a processor 30 and a memory 31. Those skilled in the art will understand that... Figure 3 The computer device 3 is merely an example and does not constitute a limitation on the computer device 3. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.
[0141] The processor 30 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0142] In some embodiments, the memory 31 may be an internal storage unit of the computer device 3, such as a hard disk or memory of the computer device 3. In other embodiments, the memory 31 may be an external storage device of the computer device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 3. Furthermore, the memory 31 may include both internal and external storage units of the computer device 3. The memory 31 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 31 can also be used to temporarily store data that has been output or will be output.
[0143] In addition, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above method embodiments.
[0144] This application provides a computer program product that, when run on a computer device, enables the computer device to execute the steps described in the various method embodiments above.
[0145] In the several embodiments provided in this application, it will be understood that each block in the flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved.
[0146] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0147] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application for those skilled in the art.
Claims
1. A reactive power voltage optimization method based on reinforcement learning, characterized in that, include: Using a deep learning algorithm, a preset deep learning optimizer is trained with historical operating data of the actual distribution network as input and reactive voltage optimization data obtained based on the nominal model as training labels to obtain the first strategy model. The historical operating data includes generator active power, load active power and load reactive power. Using Markov decision process, a reactive power optimization agent is generated based on the first policy model; The reactive power optimization agent interacts with the actual power distribution network in real time to optimize the reactive power and voltage of the actual power distribution network, and uses reinforcement learning algorithms to update the reactive power optimization agent. The real-time interaction between the reactive power optimization agent and the actual distribution network to optimize reactive power and voltage in the actual distribution network, and the updating of the reactive power optimization agent using a reinforcement learning algorithm, include: Based on the model parameters of the first strategy model, a second strategy model is generated, and two preset critic networks and a data buffer are initialized. A target strategy model is generated based on the second strategy model, and two target critic networks are generated based on the two critic networks. If the amount of data in the data buffer is less than the preset amount of data, then based on the observation of the second state information of the actual distribution network at the current moment by the reactive power optimization agent, and according to the target strategy model, the second action information corresponding to the second state information is selected to optimize the reactive power voltage of the actual distribution network, and the data buffer is updated based on the second state information and the second action information. If the amount of data in the data buffer is not less than the preset amount of data, then based on the reactive power optimization agent's observation of the third state information of the actual distribution network at the current moment, and selecting the deterministic action information corresponding to the third state information, reactive power voltage optimization is performed on the actual distribution network, and the target policy model is updated based on the data buffer using the target critic network.
2. The reactive power voltage optimization method based on reinforcement learning as described in claim 1, characterized in that, The method utilizes a deep learning algorithm, taking historical operating data of the actual power distribution network as input and reactive power and voltage optimization data obtained based on the nominal model as training labels, to train a preset deep learning optimizer and obtain a first strategy model, including: Using the nominal model, the reactive power voltage optimization data is output based on the historical operating condition data; Using the preset deep learning optimizer, reactive power voltage control data is output based on the historical operating condition data; Based on the reactive voltage optimization data and the reactive voltage control data, calculate the first loss function of the preset deep learning optimizer; The preset deep learning optimizer is updated based on the first loss function, and it is determined whether the preset deep learning optimizer has reached the convergence condition. If the first loss function reaches its minimum value, then the preset deep learning optimizer is determined to have reached the convergence condition, and the first policy model is obtained.
3. The reactive power voltage optimization method based on reinforcement learning as described in claim 2, characterized in that, The step of using the nominal model to output the reactive power voltage optimization data based on the historical operating condition data includes: Using the nominal model and based on the historical operating data, power flow analysis is performed on the actual distribution network to output the reactive power and voltage optimization data. The nominal model is: ; ; in, For grid losses or power generation costs, As the dependent variable, To control variables, The disturbance variable is the one that incorporates the historical operating data. These are the model parameters for the active distribution network model. This is the topology of the active distribution network model. Represents the power flow equation. The inequality constraint equations represent the relationship between voltage and control variables.
4. The reactive power voltage optimization method based on reinforcement learning as described in claim 1, characterized in that, The process of generating a reactive power optimization agent based on the first policy model using a Markov decision process includes: In the Markov decision-making process, a preset reinforcement learning agent observes the first state information of the actual distribution network at the current moment. The first state information includes node injected active power, node injected reactive power, node voltage, and reactive output power. Using the first policy model, select the first action information corresponding to the first state information, and calculate the first reward information and the state information at the next moment of observation of the preset reinforcement learning agent to generate the reactive power optimization agent.
5. The reactive power voltage optimization method based on reinforcement learning as described in claim 1, characterized in that, The selection of the deterministic action information corresponding to the third state information includes: Based on a preset strategy function, the deterministic action information corresponding to the third state information is selected. The preset strategy function is: ; in, For deterministic action information, For the training neural network strategy, To explore noise, This represents the minimum adjustable capability of a reactive power adjustable device. This represents the maximum adjustable capacity of the reactive power adjustable device.
6. The reactive power voltage optimization method based on reinforcement learning as described in claim 1, characterized in that, The step of updating the target policy model using the target critic network and based on the data buffer includes: Randomly select several sets of sample data from the data buffer; Based on the sample data, calculate the objective value of the target critic network function; Based on the objective value of the function, calculate the second loss function of the target critic network and the third loss function of the target policy model; Based on the second and third loss functions, the regularization coefficients, the target critic network, and the target policy model are updated.
7. A reactive power voltage optimization device based on reinforcement learning, characterized in that, include: The training module is used to train a preset deep learning optimizer using a deep learning algorithm, taking historical operating data of the actual distribution network as input and reactive voltage optimization data obtained based on the nominal model as training labels, to obtain the first strategy model. The historical operating data includes generator active power, load active power and load reactive power. The generation module is used to generate a reactive power optimization agent based on the first policy model using a Markov decision process. An optimization module is used to perform real-time interaction between the reactive power optimization agent and the actual distribution network to optimize the reactive power voltage of the actual distribution network, and to update the reactive power optimization agent using a reinforcement learning algorithm. The real-time interaction between the reactive power optimization agent and the actual distribution network to optimize reactive power and voltage in the actual distribution network, and the updating of the reactive power optimization agent using a reinforcement learning algorithm, include: Based on the model parameters of the first strategy model, a second strategy model is generated, and two preset critic networks and a data buffer are initialized. A target strategy model is generated based on the second strategy model, and two target critic networks are generated based on the two critic networks. If the amount of data in the data buffer is less than the preset amount of data, then based on the observation of the second state information of the actual distribution network at the current moment by the reactive power optimization agent, and according to the target strategy model, the second action information corresponding to the second state information is selected to optimize the reactive power voltage of the actual distribution network, and the data buffer is updated based on the second state information and the second action information. If the amount of data in the data buffer is not less than the preset amount of data, then based on the reactive power optimization agent's observation of the third state information of the actual distribution network at the current moment, and selecting the deterministic action information corresponding to the third state information, reactive power voltage optimization is performed on the actual distribution network, and the target policy model is updated based on the data buffer using the target critic network.
8. A computer device, characterized in that, It includes a processor and a memory, the memory being used to store a computer program, which, when executed by the processor, implements the reactive voltage optimization method based on reinforcement learning as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the reactive voltage optimization method based on reinforcement learning as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Power distribution network topology voltage regulation method based on deep reinforcement learning
CN113036772A
Power distribution network optimization method based on multi-agent deep reinforcement learning
CN114725936A