Power grid power flow control method and device, storage medium and electronic equipment

Through real-time monitoring and neural network model power flow control method, the power flow control is automated, which solves the problem of difficult manual control of the power grid and improves the efficiency of power grid management.

CN120728602APending Publication Date: 2025-09-30EAST CHINA BRANCH OF STATE GRID CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410362223.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-28
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Manual methods are difficult to implement during power grid flow control, which leads to increased complexity in grid management and a sharp increase in the number of controlled objects, making manual management difficult.

Method used

Adopting the grid element state parameter vector based on real-time monitoring, the preset power flow control model and neural network model are used for strategy matching and control, including constructing the neural network model, training and optimization, and generating the grid power flow control strategy.

Benefits of technology

It realizes automatic control without manual adjustment of the power grid status, saves manpower and material resources, and improves the efficiency of power grid flow control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120728602A_ABST
    Figure CN120728602A_ABST
Patent Text Reader

Abstract

The invention discloses a power grid power flow control method and device, a storage medium and electronic equipment, and is applied to the technical field of power grid operation control. The method comprises the steps that first state parameter vectors of all power grid elements of a current power grid framework are obtained through real-time monitoring, and the power grid elements comprise one or more of a transformer substation, a line, a load and a generator; based on the first state parameter vectors, a preset power flow control model is adopted for strategy matching, and a power grid power flow control strategy corresponding to the current power grid architecture is obtained; and controlling the current power grid architecture by adopting the power grid power flow control strategy to obtain a power grid power flow control result. The power flow control method does not need to manually adjust the power grid state, manpower and material resources are saved, and the power grid power flow control efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power grid operation control, and in particular to a power grid current control method, device, storage medium and electronic equipment. Background Art

[0002] With the increasing complexity of power systems, the continued penetration of renewable energy, and the drive towards carbon reduction, the operation and control of power systems are becoming increasingly challenging. These changes have led to a significant increase in system complexity, a sharp increase in the number of control objects in the power grid, and increasingly difficult manual management. Summary of the Invention

[0003] In view of this, the present invention provides a power grid flow control method, device, storage medium and electronic equipment, the main purpose of which is to solve the problem that it is difficult to achieve power grid flow control using manual methods during the current power grid flow control process.

[0004] To solve the above problems, the present application provides a power grid power flow control method, including:

[0005] Real-time monitoring to obtain a first state parameter vector of each grid element of the current grid architecture, wherein the grid elements include: one or more of a substation, a line, a load, and a generator;

[0006] Based on each of the first state parameter vectors, a preset power flow control model is used to perform strategy matching to obtain a power grid power flow control strategy corresponding to the current power grid architecture;

[0007] The power grid flow control strategy is used to control the current power grid architecture to obtain a power grid flow control result.

[0008] Optionally, before performing calculation processing using a preset power flow control model based on each of the state parameter vectors, the method further includes: constructing the preset power flow control model, specifically including:

[0009] Obtaining a preprocessed state data set and an action data set based on a target power grid architecture and predetermined constraints for ensuring safe operation of the target power grid architecture;

[0010] Determining structural parameters of a neural network model for training the preset power flow control model based on the state data set and the action data set;

[0011] The neural network model is configured using the structural parameters to obtain a first neural network for action selection and a second neural network for expected value calculation;

[0012] Using the state data set and the action data set as training samples to train the first neural network and the second neural network, updating model parameters of the first neural network and the second neural network, and generating an initial power flow control model;

[0013] The initial power flow control model is optimized by using predetermined hyperparameters of the preset power flow control model to construct the preset power flow control model.

[0014] Optionally, the using the state data set and the action data set as training samples to train the first neural network and the second neural network, updating the model parameters of the first neural network and the second neural network, and generating the preset power flow control model specifically includes:

[0015] In the same training cycle, starting from the initial state of the state data set, the first neural network is used to perform action selection for the initial state to obtain initial pre-action sets corresponding to the initial state and first historical expected values ​​corresponding to the initial pre-action sets;

[0016] Screening each of the initial pre-action sets and each of the first historical expected values ​​to obtain a target initial pre-action set and a target first historical expected value;

[0017] Based on a number of previously acquired historical experience data corresponding to the initial state, the second neural network is used to perform calculation processing to obtain a second historical expected value of the target corresponding to the initial state;

[0018] Based on the first historical expected value of the target, the second historical expected value of the target, and a preset loss function, the back propagation algorithm and the gradient descent method are used to calculate the parameters of the first neural network to obtain the historical target parameter values;

[0019] updating the model parameters of the first neural network based on each of the historical target parameter values;

[0020] The target power grid architecture is controlled based on the target initial pre-action set to obtain an intermediate state, and historical target parameter values ​​corresponding to the intermediate state are determined using an updated first neural network and an updated second neural network. The first neural network is updated based on the historical target parameter values ​​corresponding to the intermediate state until a current training cycle ends and the current training cycle is greater than or equal to a preset cycle threshold, thereby obtaining the preset power flow control model.

[0021] Optionally, performing strategy matching based on each of the first state parameter vectors using a preset power flow control model to obtain a power grid power flow control strategy corresponding to the current power grid architecture specifically includes:

[0022] Performing action selection using a first neural network based on each of the first state parameter vectors to obtain each of the first action sets corresponding to the first state vectors and a first expected value corresponding to each of the first action sets;

[0023] Screening each of the first expected values ​​to obtain a first target expected value;

[0024] The first action sets are screened based on the first target expected value to obtain the power grid power flow control strategy.

[0025] Optionally, after performing strategy matching using a preset power flow control model based on each of the first state parameter vectors to obtain a power flow control strategy corresponding to the current power grid architecture, the method further includes:

[0026] querying a preset experience data set based on the first state parameter vector to obtain a plurality of first experience data corresponding to the first state parameter vector;

[0027] Performing calculations using a second neural network based on each of the first empirical data sets to obtain a second target expected value;

[0028] Based on the first target expected value, the second target expected value and a preset loss function, a back propagation algorithm is used to perform calculation processing to obtain a gradient value;

[0029] Based on the gradient value, a gradient descent method is used to perform calculation processing to obtain the value of each target parameter;

[0030] The first neural network is updated using each of the target parameter values ​​to update the preset power flow control model.

[0031] Optionally, the method further includes:

[0032] According to the preset update frequency, the target parameter values ​​corresponding to the current first neural network are collected and obtained;

[0033] The model parameters of the second neural network are updated based on the target parameter values ​​to update the preset power flow control model.

[0034] Optionally, after the power grid flow control strategy is adopted to control the current power grid architecture, a second state parameter vector of each power grid element of the current power grid architecture is obtained and a target reward value for controlling the current power grid architecture using the power grid flow control strategy is randomly generated. The method further includes:

[0035] Based on the first state parameter vector, the first action set, the target reward value, and the second state parameter vector, construct an experience data set corresponding to the first state parameter vector;

[0036] The experience data set is stored in a predetermined storage area.

[0037] To solve the above problems, the present application provides a power grid flow control device, comprising:

[0038] A monitoring module, configured to monitor in real time and obtain a first state parameter vector of each grid element of the current grid architecture, wherein the grid elements include one or more of a substation, a line, a load, and a generator;

[0039] a strategy matching module, configured to perform strategy matching using a preset power flow control model based on each of the first state parameter vectors to obtain a power flow control strategy corresponding to the current power grid architecture;

[0040] The control module is used to control the current power grid architecture by adopting the power grid power flow control strategy to obtain a power grid power flow control result.

[0041] To solve the above problem, the present application provides a storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned power grid power flow control method are implemented.

[0042] To solve the above problems, the present application provides an electronic device, which includes at least a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps of the above-mentioned power grid flow control method when executing the computer program in the memory.

[0043] The beneficial effects of the present application are as follows: the present application adopts a preset flow control model to perform strategy matching based on the first state parameter vectors obtained in real time to obtain a grid flow control strategy for grid flow control, and controls the current grid structure. There is no need to manually adjust the grid state, saving manpower and material resources and improving the efficiency of grid flow control.

[0044] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:

[0046] Figure 1 A schematic diagram of a flow chart of a power grid power flow control method provided in an embodiment of the present application is shown;

[0047] Figure 2 A schematic diagram of a flow chart of a power grid power flow control method provided by another embodiment of the present application is shown;

[0048] Figure 3 A schematic diagram of a power grid architecture using an IEEE 14-bus test system is shown;

[0049] Figure 4 A schematic diagram of the action space topology of the grid architecture of the adopted IEEE 14-bus test system is shown;

[0050] Figure 5 A structural block diagram of a power grid flow control device provided in another embodiment of the present application is shown. DETAILED DESCRIPTION

[0051] Various aspects and features of the present application are described herein with reference to the accompanying drawings.

[0052] It should be understood that various modifications may be made to the embodiments of the present application. Therefore, the above description should not be considered as limiting, but merely as an example of an embodiment. Other modifications within the scope and spirit of the present application will occur to those skilled in the art.

[0053] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the present application and, together with the general description of the present application given above and the detailed description of the embodiments given below, serve to explain the principles of the present application.

[0054] These and other characteristics of the present application will become apparent from the following description of a preferred form of embodiment given as a non-limiting example with reference to the accompanying drawings.

[0055] It should also be understood that although the present application has been described with reference to certain specific examples, those skilled in the art will readily be able to implement many other equivalent forms of the present application.

[0056] The above and other aspects, features and advantages of the present application will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings.

[0057] Specific embodiments of the present application will be described hereinafter with reference to the accompanying drawings; however, it should be understood that the embodiments described are merely examples of the present application and may be implemented in a variety of ways. Familiar and / or repetitive functions and structures are not described in detail to avoid obscuring the present application with unnecessary or redundant details. Therefore, the specific structural and functional details described herein are not intended to be limiting, but rather serve merely as a basis and representative basis for the claims to teach those skilled in the art to variously utilize the present application with substantially any suitable detailed structure.

[0058] This specification may use the phrases "in one embodiment," "in another embodiment," "in yet another embodiment," or "in other embodiments," which may all refer to one or more of the same or different embodiments according to the present application.

[0059] The present application provides a method for controlling power flow in a power grid. Figure 1 As shown, including:

[0060] Step S101: real-time monitoring to obtain a first state parameter vector of each grid element of the current grid architecture, wherein the grid elements include one or more of a substation, a line, a load, and a generator;

[0061] During the specific implementation of this step, the grid architecture is composed of several grid elements, each of which includes: substations, lines, loads and generators, etc. The generators include but are not limited to hydropower generators, wind power generators, solar power generators, nuclear power generators and wind power generators, etc. The first state parameter vector includes: time parameters, active power, reactive power, voltage, current, line capacity, topology vector, line state, time step after line overload, substation shutdown time step and line shutdown time step.

[0062] Step S102: performing strategy matching using a preset power flow control model based on each of the first state parameter vectors to obtain a power flow control strategy corresponding to the current power grid architecture;

[0063] During the specific implementation of this step, a first neural network is used to select actions based on each of the first state parameter vectors to obtain first action sets corresponding to the first state vectors and first expected values ​​corresponding to each of the first action sets. Each of the first expected values ​​is screened to obtain a first target expected value. Specifically, the first expected value with the largest value among the first expected values ​​can be determined as the first target expected value. Each of the first action sets is screened based on the first target expected value to obtain the power grid power flow control strategy.

[0064] Step S103: adopting the power grid flow control strategy to control the current power grid architecture and obtain a power grid flow control result.

[0065] During the specific implementation of this step, the first action set corresponding to the power grid flow control strategy is used to control the status of each power grid element of the current power grid architecture to obtain a power grid flow control result.

[0066] This application uses a preset flow control model to perform strategy matching based on the first state parameter vectors obtained in real time to obtain a grid flow control strategy for grid flow control, and controls the current grid structure. There is no need to manually adjust the grid state, saving manpower and material resources and improving the efficiency of grid flow control.

[0067] Another embodiment of the present application provides another power grid power flow control method, such as Figure 2 As shown, including:

[0068] Step S201: real-time monitoring to obtain a first state parameter vector of each grid element in the current grid architecture;

[0069] During the specific implementation of this step, the grid architecture is composed of several grid elements, each of which includes: substations, lines, loads and generators, etc. The generators include but are not limited to hydropower generators, wind power generators, solar power generators, nuclear power generators and wind power generators, etc. The first state parameter vector includes: time parameters, active power, reactive power, voltage, current, line capacity, topology vector, line state, time step after line overload, substation shutdown time step and line shutdown time step.

[0070] Step S202: obtaining a pre-processed state data set and an action data set based on a target power grid architecture and predetermined constraints for ensuring safe operation of the target power grid architecture;

[0071] During the specific implementation of this step, multiple operational constraints affecting the safe operation of the target grid architecture are determined for the target grid architecture to ensure the normal operation of the grid architecture. These operational constraints include: system requirements must be fully met; no generators or loads can be disconnected; electrical islands cannot be formed due to topology changes; and AC power flow must converge. If hard constraints are violated, a power system blackout will occur. A target operation mode is determined for the target grid architecture to ensure the safe operation of the target grid architecture. This target operation mode may be an operational adjustment of the topology. Based on this target operation mode, multiple operational constraints are determined to ensure the safe operation of the target grid architecture. These operational constraints may include: reconnecting or disconnecting a line; changing the substation configuration; and other operational constraints. These operational constraints may include: the minimum number of components that must be connected to each substation is two; and the components connected to the substation must have at least two lines. A state space set and an action space set are constructed for the target grid architecture. The state space describes the power grid perceived by the intelligent agent. The power flow control method of this application can be executed on the intelligent agent to control the power flow. Data cleaning is performed on the state space set and the action space set to obtain the state data set and the action data set. Figure 3 The grid architecture of the IEEE 14-bus test system is shown in the figure. The power system model includes 14 substations, 20 lines, 11 loads and 6 generators. The current energy mix includes hydropower, nuclear power, thermal power, wind power and two solar power plants. The dataset of the IEEE 14-node test system contains 1004 monthly scenarios, each representing 28 days with a time interval of 5 minutes. By carefully analyzing the actions involving bus splitting, the final substation configuration is selected as the action space. The number of actions depends on the total number of components connected to the substation, which is 2 n-1 , where n is the total number of elements connected to the substation. The number of actions also depends on the number of false actions, such as disconnecting loads or disconnecting generators, which can be calculated as 2 B -1, where B is the number of generators and loads in the substation. The analysis of the substation configuration parameters is shown in Table 1. Excluding erroneous actions and actions that do not take any action, the total number of bus splitting actions in the test network is 179.

[0072]

[0073] Table 1 Configuration parameters of each substation

[0074] Substation 7 is connected to only one line and one generator, so it cannot be split. The number of actions given in this manner is too large for the selected power grid. Expert knowledge indicates that this power grid does not possess the necessary topologies to meet the grid control requirements, so further action selection is required. This action selection principle is based on the fundamental principle of network planning, n-1 safety. This principle states that if a component fails or shuts down while the system is operating at maximum load and power, network safety must still be maintained. This imposes the following constraints: the minimum number of components connected to each substation must be two, and each component must have at least two lines connected to the substation. This reduces the number of substations that can be split by half. Since only one topological action can be selected per time step, the final selected actions are those that still meet the constraints after the actions are performed. For this power grid, the number of substations that can be split simultaneously must be less than or equal to three. Therefore, only three substations, 1, 3, and 4, are considered as potential final topology choices. In addition to the selected actions, the action space also includes an action that sets the substation's initial topology (each busbar component is connected to the same busbar segment) and a "do nothing" action. Thus, the final action space contains only nine action quantities, such as Figure 4 Figure 2 shows the action space topology for the grid architecture of the IEEE 14-bus test system. The state space describes the grid as perceived by the agent. The initial state space set is cleaned to include the voltage and current at both ends of the line, the line capacity, and the topological vectors for each object (load, generator, and both ends of the line) connected to the substation. This yields the state data sets described above.

[0075] Step S203: determining structural parameters of a neural network model for training the preset power flow control model based on the state data set and the action data set;

[0076] In the specific implementation process of this step, the neural network model used in this application includes: an input layer, an output layer and a hidden layer; the size of the input layer is determined by the size of the state space; the size of the output layer is determined by the size of the action space; there are two hidden layers, and each hidden layer contains 128 neurons. For example: Figure 4 The action space topology diagram of the grid architecture used in the IEEE 14-bus test system is shown. The state space includes 157 features, and the action space size is 9. The structural parameters of the neural network model are an input layer size of 157, an output layer size of 9, and two hidden layers containing 128 neurons each. The structural parameters are determined by analyzing the state and action datasets of the target grid structure in actual application.

[0077] Step S204: configuring the neural network model using the structural parameters to obtain a first neural network for action selection and a second neural network for expected value calculation;

[0078] During the specific implementation of this step, the present application uses the same structural parameters to configure the first neural network and the second neural network, so that the input layer size, output layer size, number of hidden layers, and structural parameters of each hidden layer containing 128 neurons of the first neural network and the second neural network are the same. The present application uses an interleaving double Q network (dueling double DQN, referred to as DDDQN). The first neural network can be the main neural Q network of DDDQN, and the second neural network can be the target neural Q network of DDDQN. The first neural network is used to select the next action, and the second neural network is used to calculate the expected value, thereby more stably updating the parameters of the first neural network.

[0079] Step S205: using the state data set and the action data set as training samples to train the first neural network and the second neural network, updating the model parameters of the first neural network and the second neural network, and generating an initial power flow control model;

[0080] During the specific implementation of this step, in the same training cycle, starting from the initial state of the state data set, the first neural network is used to select actions for the initial state to obtain initial pre-action sets corresponding to the initial state and first historical expected values ​​corresponding to each initial pre-action set; each initial pre-action set and each first historical expected value are screened to obtain a target initial pre-action set and a target first historical expected value; based on a number of pre-acquired historical experience data corresponding to the initial state, the second neural network is used to perform calculations to obtain a target second historical expected value corresponding to the initial state; based on the target first historical expected value, the target second historical expected value and a preset loss function, the back propagation algorithm and the gradient descent method are used to calculate the parameters of the first neural network to obtain historical target parameter values; specifically, the preset loss function is used to calculate the target first historical expected value and the target second historical expected value to obtain a loss value; when the loss value is greater than or equal to the preset loss value threshold, the preset loss function is optimized using the back propagation algorithm to obtain a gradient value, and the parameters of the first neural network are calculated based on the gradient value to obtain historical target parameter values ​​of the first neural network. The model parameters of the first neural network are updated based on the historical target parameter values; the target power grid architecture is controlled based on the target initial pre-action set to obtain an intermediate state, and the historical target parameter values ​​corresponding to the intermediate state are determined using the updated first neural network and the updated second neural network. The first neural network is updated based on the historical target parameter values ​​corresponding to the intermediate state, until the current training cycle ends and the current training cycle is greater than or equal to a preset cycle threshold, thereby obtaining the preset power flow control model. Alternatively, the preset power flow control model is obtained when the current training cycle is less than the preset cycle threshold and the loss value is less than the preset loss value threshold.

[0081] Step S206: Optimizing the initial power flow control model using predetermined hyperparameters of the preset power flow control model to construct the preset power flow control model;

[0082] In the specific implementation process of this step, the predetermined hyperparameters include: neural network hidden layer structure, neural network hidden layer activation function, number of parallel training work processes, number of steps for training start, number of samples for each training batch, target network update frequency, learning rate related parameters, exploration parameters, etc.; wherein, the neural network hidden layer structure is [128,128], the neural network hidden layer activation function is relu, the number of parallel training work processes is 4, the number of steps for training start is 5000, the number of samples for each training batch is 64, the target network update frequency is 256, and the learning rate related parameters include: initial learning rate 10×10-5 , the final learning rate is 10×10 -6 , learning rate scheduling time 100,000, exploration parameters include: initial exploration rate 0.4, final exploration rate 1 / 2016, ε exploration time step 100,000; the hyperparameters of this application also include the default parameters of the DQN algorithm in the RLlib library.

[0083] Step S207: performing action selection using a first neural network based on each of the first state parameter vectors to obtain each of the first action sets corresponding to the first state vectors and a first expected value corresponding to each of the first action sets;

[0084] During the specific implementation of this step, the action value function of the current first neural network is used to perform calculations based on each of the first state parameter vectors to obtain each of the first action sets and the first expected value corresponding to each of the first action sets; the mathematical expression of the action value function can be expressed as follows:

[0085] q π (s,a)=E π [G t |S t =s,A t =a]=E π [R t+1 +γR t+2 +γ 2 R t+2 +…|S t =s,A t =a] (1);

[0086] Among them, s t Indicates the state, s t Belongs to the state set S, A t is an action belonging to the action set A; R is the reward value, and γ is a constant between 0 and 1, which is used to determine the relative value of delayed reward and immediate reward.

[0087] Step S208: screening each of the first expected values ​​to obtain a first target expected value;

[0088] During the specific implementation of this step, a predetermined optimal action function is used to screen each of the first expected values ​​to obtain a first target expected value; the mathematical expression of the predetermined optimal action function is shown in the following formula (2): (2); where s∈S, a∈A(s). The largest first expected value among the first expected values ​​is determined as the first target expected value.

[0089] Step S209: screening each of the first action sets based on the first target expected value to obtain the power grid power flow control strategy;

[0090] During the specific implementation of this step, each of the first action sets is screened based on the first target expected value. Specifically, each of the first action sets is screened based on the first target expected value and the mathematical expression of the Bellman equation to obtain the power grid power flow control strategy. Specifically, the mathematical expression of the Bellman equation can be expressed as the following formula (3):

[0091]

[0092] Step S210: querying a preset experience data set based on the first state parameter vector to obtain a plurality of first experience data corresponding to the first state parameter vector;

[0093] During the specific implementation of this step, the preset experience data set includes several historical state parameter vectors, the historical first action set corresponding to each of the historical state parameter vectors, the next state parameter vector corresponding to each of the historical state vectors obtained by screening, and the reward value obtained after the power grid is controlled by the historical first action set. This lays the foundation for subsequent calculation and processing based on each of the first experience data sets using the second neural network to obtain the second target expected value. This application uses experience data with experience replay to calculate the second target expected value. By repeatedly using past experience data for training, the main Q network can better learn the long-term laws and strategies of the environment. Experience replay can reduce volatility in the training process, improve the stability and convergence speed of the algorithm, and make the trained main Q network more robust.

[0094] Step S211: performing calculations using a second neural network based on each of the first experience data sets to obtain a second target expected value;

[0095] During the specific implementation of this step, a second neural network is used to calculate and process each of the first empirical data sets to obtain the second target expected value. Through the interaction between the first and second neural networks, the DDDQN algorithm can more effectively learn the action-value function and achieve better performance in complex reinforcement learning tasks. The first neural network is responsible for action selection and prediction of the first target expected value, while the second neural network is used to calculate the stable target Q value, namely the second target expected value. The cooperation between the two can improve the stability and convergence speed of the algorithm.

[0096] Step S212: performing calculations using a back propagation algorithm based on the first target expected value, the second target expected value, and a preset loss function to obtain a gradient value;

[0097] During the specific implementation of this step, the preset loss function can calculate the mean square error (MSE) loss to measure the difference between the predicted first target expected value and the target Q value, that is, the second target expected value, and use the back propagation algorithm for calculation and processing to obtain the gradient value.

[0098] Step S213: performing calculation based on the gradient value using a gradient descent method to obtain target parameter values;

[0099] In the specific implementation process of this step, the target parameter values ​​are calculated based on the gradient value and the set learning rate related parameters. The mathematical formula for calculating the target parameter value can be expressed as follows:

[0100]

[0101] where w t 、w t+1 is the vector of network weights at time t and time t+1, A t is the action chosen at time t, S t and S t+1 are the state parameter vectors of the network at time t and t+1 respectively.

[0102] Step S214: updating the first neural network using the target parameter values ​​to update the preset power flow control model;

[0103] During the specific implementation of this step, each of the target parameter values ​​is used to update the model parameters of the first neural network to update the preset power flow control model.

[0104] Step S215: acquiring target parameter values ​​corresponding to the current first neural network according to a preset update frequency;

[0105] In the specific implementation process of this step, based on a preset update frequency, for example, a hard update is performed every N steps. When the preset update frequency is reached, the target parameter values ​​corresponding to the current first neural network are collected and obtained.

[0106] Step S216: updating the model parameters of the second neural network based on the target parameter values ​​to update the preset power flow control model;

[0107] During the specific implementation of this step, a preset update frequency is defined, for example, a hard update is performed every N steps. When the predetermined update frequency is reached, the model parameters of the first neural network are copied to the second neural network for updating. The model parameters of the second neural network can also be updated by a soft update method. In the soft update, the model parameters of the second neural network will gradually approach the model parameters of the first neural network with a certain degree of smoothness. The parameters of the second neural network can be updated using the Exponential Moving Average (EMA) method. The model parameters of the second neural network are updated regularly. This can maintain the consistency of the second neural network with the first neural network and improve the stability and convergence speed of the algorithm.

[0108] Step S217: adopting the power grid flow control strategy to control the current power grid architecture and obtain a power grid flow control result.

[0109] During the specific implementation of this step, the power flow control strategy is used to control the current power grid architecture, changing the connection mode of the current power grid architecture to obtain a second state parameter vector for each power element in the current power grid architecture, and randomly generating a target reward value for controlling the current power grid architecture using the power flow control strategy. Based on the first state parameter vector, the first action set, the target reward value, and the second state parameter vector, an experience dataset corresponding to the first state parameter vector is constructed; the experience dataset is stored in a predetermined storage area. A prioritized experience replay method is used to reuse past experience, replaying important transformations more frequently, thereby achieving more effective learning.

[0110] This application uses a dual-interleaved deep Q network training with priority experience replay to obtain a power flow control model that can autonomously adjust the topology, successfully operate the power grid in as many scenarios as possible through topology adjustment operations alone, simplify the action space, and design a simple and effective reward mechanism, so that the intelligent agent can safely control the operation of the power grid for a long time without the help of experts. In scenarios with different renewable energy shares, the intelligent agent can effectively control the operation of the power grid for a longer time. Various constraints are taken into account in the training model process to effectively prevent line overload and power outages, while making the maximum line capacity utilization rate lower, demonstrating good adaptability.

[0111] Another embodiment of the present application provides a power grid flow control device, such as Figure 5 As shown, including:

[0112] Monitoring module 1, for real-time monitoring and obtaining a first state parameter vector of each grid element of the current grid architecture, wherein the grid elements include one or more of a substation, a line, a load, and a generator;

[0113] a strategy matching module 2, configured to perform strategy matching using a preset power flow control model based on each of the first state parameter vectors to obtain a power flow control strategy corresponding to the current power grid architecture;

[0114] The control module 3 is used to control the current power grid architecture by adopting the power grid power flow control strategy to obtain a power grid power flow control result.

[0115] During the specific implementation process, the power grid flow control device also includes: a construction module, which is specifically used to: construct the preset flow control model, specifically including: obtaining a preprocessed state data set and an action data set based on the target power grid architecture and predetermined constraints to ensure the safe operation of the target power grid architecture; determining the structural parameters of the neural network model used to train the preset flow control model based on the state data set and the action data set; configuring the neural network model using the structural parameters to obtain a first neural network for action selection and a second neural network for expected value calculation; using the state data set and the action data set as training samples to train the first neural network and the second neural network, updating the model parameters of the first neural network and the second neural network, and generating an initial flow control model; optimizing the initial flow control model using each predetermined hyperparameter of the preset flow control model to construct the preset flow control model.

[0116] In a specific implementation process, the construction module is also used to: in the same training cycle, starting from the initial state of the state data set, use the first neural network to select an action for the initial state, and obtain each initial pre-action set corresponding to the initial state and the first historical expected value corresponding to each initial pre-action set; screen each initial pre-action set and each first historical expected value to obtain a target initial pre-action set and a target first historical expected value; based on a number of historical experience data corresponding to the initial state obtained in advance, use the second neural network to perform calculation processing to obtain a target second historical expected value corresponding to the initial state; based on the target first historical expected value, the target The method comprises the following steps: a first neural network is used to determine the historical target parameter values ​​​​of the first neural network and the second historical expected value and the preset loss function, and a back propagation algorithm and a gradient descent method are used to calculate the parameters of the first neural network to obtain the historical target parameter values; the model parameters of the first neural network are updated based on the historical target parameter values; the target power grid architecture is controlled based on the target initial pre-action set to obtain an intermediate state, and the updated first neural network and the updated second neural network are used to determine the historical target parameter values ​​corresponding to the intermediate state, and the first neural network is updated based on the historical target parameter values ​​corresponding to the intermediate state, until the current training cycle ends and the current training cycle is greater than or equal to the preset cycle threshold, so as to obtain the preset power flow control model.

[0117] During the specific implementation process, the strategy matching module 2 is specifically used to: use the first neural network to select actions based on each first state parameter vector to obtain each first action set corresponding to the first state vector and the first expected value corresponding to each first action set; screen each first expected value to obtain a first target expected value; screen each first action set based on the first target expected value to obtain the power grid flow control strategy.

[0118] During the specific implementation process, the power grid flow control device also includes: a first neural network update module, which is specifically used to: query a preset experience data set based on the first state parameter vector to obtain a number of first experience data corresponding to the first state parameter vector; use a second neural network to perform calculation processing based on each of the first experience data sets to obtain a second target expected value; use a back propagation algorithm to perform calculation processing based on the first target expected value, the second target expected value and a preset loss function to obtain a gradient value; use a gradient descent method to perform calculation processing based on the gradient value to obtain each target parameter value; use each of the target parameter values ​​to update the first neural network to update the preset flow control model.

[0119] During the specific implementation process, the power grid flow control device also includes: a second neural network update module, which is specifically used to: collect and obtain the target parameter values ​​corresponding to the current first neural network according to the preset update frequency; update the model parameters of the second neural network based on the target parameter values ​​to update the preset flow control model.

[0120] During the specific implementation process, the power grid flow control device also includes: a storage module, which is specifically used to: construct an experience data set corresponding to the first state parameter vector based on the first state parameter vector, the first action set, the target reward value and the second state parameter vector; and store the experience data set in a predetermined storage area.

[0121] This application uses a preset flow control model to perform strategy matching based on the first state parameter vectors obtained in real time to obtain a grid flow control strategy for grid flow control, and controls the current grid structure. There is no need to manually adjust the grid state, saving manpower and material resources and improving the efficiency of grid flow control.

[0122] Another embodiment of the present application provides a storage medium storing a computer program. When the computer program is executed by a processor, the following method steps are implemented:

[0123] Step 1: Real-time monitoring to obtain a first state parameter vector of each grid element of the current grid architecture, wherein the grid elements include: one or more of a substation, a line, a load, and a generator;

[0124] Step 2: Based on each of the first state parameter vectors, a preset power flow control model is used to perform strategy matching to obtain a power flow control strategy corresponding to the current power grid architecture;

[0125] Step 3: Use the power grid flow control strategy to control the current power grid architecture and obtain a power grid flow control result.

[0126] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0127] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0128] The specific implementation process of the above method steps can be referred to the embodiments of any of the above power grid power flow control methods, and will not be repeated in this embodiment.

[0129] This application uses a preset flow control model to perform strategy matching based on the first state parameter vectors obtained in real time to obtain a grid flow control strategy for grid flow control, and controls the current grid structure. There is no need to manually adjust the grid state, saving manpower and material resources and improving the efficiency of grid flow control.

[0130] Another embodiment of the present application provides an electronic device, which may be a server, and the electronic device includes a processor, a memory, a network interface, and a database connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external client via a network connection. When the electronic device program is executed by the processor, it implements the functions or steps on the server side of a power grid flow control method.

[0131] In one embodiment, an electronic device is provided, which may be a client. The electronic device includes a processor, memory, a network interface, a display screen, and an input device connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external server via a network connection. When the electronic device program is executed by the processor, it implements the functions or steps on the client side of a power grid power flow control method.

[0132] Another embodiment of the present application provides an electronic device, comprising at least a memory and a processor, wherein the memory stores a computer program, and the processor implements the following method steps when executing the computer program in the memory:

[0133] Step 1: Real-time monitoring to obtain a first state parameter vector of each grid element of the current grid architecture, wherein the grid elements include: one or more of a substation, a line, a load, and a generator;

[0134] Step 2: Based on each of the first state parameter vectors, a preset power flow control model is used to perform strategy matching to obtain a power flow control strategy corresponding to the current power grid architecture;

[0135] Step 3: Use the power grid flow control strategy to control the current power grid architecture and obtain a power grid flow control result.

[0136] The specific implementation process of the above method steps can be referred to the embodiments of any of the above power grid power flow control methods, and will not be repeated in this embodiment.

[0137] This application uses a preset flow control model to perform strategy matching based on the first state parameter vectors obtained in real time to obtain a grid flow control strategy for grid flow control, and controls the current grid structure. There is no need to manually adjust the grid state, saving manpower and material resources and improving the efficiency of grid flow control.

[0138] The above embodiments are merely exemplary embodiments of the present application and are not intended to limit the scope of the present application. The scope of protection of the present application is defined by the claims. Those skilled in the art may make various modifications or equivalent substitutions to the present application within the essence and scope of protection of the present application, and such modifications or equivalent substitutions shall also be deemed to fall within the scope of protection of the present application.

Claims

1. A power grid power flow control method, characterized in that: include: Real-time monitoring to obtain a first state parameter vector of each grid element of the current grid architecture, wherein the grid elements include: one or more of a substation, a line, a load, and a generator; Based on each of the first state parameter vectors, a preset power flow control model is used to perform strategy matching to obtain a power grid power flow control strategy corresponding to the current power grid architecture; The power grid flow control strategy is used to control the current power grid architecture to obtain a power grid flow control result.

2. The method according to claim 1, wherein Before performing calculation processing using a preset power flow control model based on each of the state parameter vectors, the method further includes: constructing the preset power flow control model, specifically including: Obtaining a preprocessed state data set and an action data set based on a target power grid architecture and predetermined constraints for ensuring safe operation of the target power grid architecture; Determining structural parameters of a neural network model for training the preset power flow control model based on the state data set and the action data set; The neural network model is configured using the structural parameters to obtain a first neural network for action selection and a second neural network for expected value calculation; Using the state data set and the action data set as training samples to train the first neural network and the second neural network, updating model parameters of the first neural network and the second neural network, and generating an initial power flow control model; The initial power flow control model is optimized by using predetermined hyperparameters of the preset power flow control model to construct the preset power flow control model.

3. The method according to claim 2, wherein The step of using the state data set and the action data set as training samples to train the first neural network and the second neural network, updating the model parameters of the first neural network and the second neural network, and generating the preset power flow control model specifically includes: In the same training cycle, starting from the initial state of the state data set, the first neural network is used to perform action selection for the initial state to obtain initial pre-action sets corresponding to the initial state and first historical expected values ​​corresponding to the initial pre-action sets; Screening each of the initial pre-action sets and each of the first historical expected values ​​to obtain a target initial pre-action set and a target first historical expected value; Based on a number of previously acquired historical experience data corresponding to the initial state, the second neural network is used to perform calculation processing to obtain a second historical expected value of the target corresponding to the initial state; Based on the first target historical expected value, the second target historical expected value, and a preset loss function, the back propagation algorithm and the gradient descent method are used to calculate the parameters of the first neural network to obtain the historical target parameter values; updating the model parameters of the first neural network based on each of the historical target parameter values; The target power grid architecture is controlled based on the target initial pre-action set to obtain an intermediate state, and historical target parameter values ​​corresponding to the intermediate state are determined using an updated first neural network and an updated second neural network. The first neural network is updated based on the historical target parameter values ​​corresponding to the intermediate state until a current training cycle ends and the current training cycle is greater than or equal to a preset cycle threshold, thereby obtaining the preset power flow control model.

4. The method according to claim 1, wherein The step of performing strategy matching based on each of the first state parameter vectors using a preset power flow control model to obtain a power flow control strategy corresponding to the current power grid architecture specifically includes: Performing action selection using a first neural network based on each of the first state parameter vectors to obtain each of the first action sets corresponding to the first state vectors and a first expected value corresponding to each of the first action sets; Screening each of the first expected values ​​to obtain a first target expected value; The first action sets are screened based on the first target expected value to obtain the power grid power flow control strategy.

5. The method according to claim 4, wherein After performing strategy matching using a preset power flow control model based on each of the first state parameter vectors to obtain a power flow control strategy corresponding to the current power grid architecture, the method further includes: querying a preset experience data set based on the first state parameter vector to obtain a plurality of first experience data corresponding to the first state parameter vector; Performing calculations using a second neural network based on each of the first empirical data sets to obtain a second target expected value; Based on the first target expected value, the second target expected value and a preset loss function, a back propagation algorithm is used to perform calculation processing to obtain a gradient value; Based on the gradient value, a gradient descent method is used to perform calculation processing to obtain the value of each target parameter; The first neural network is updated using each of the target parameter values ​​to update the preset power flow control model.

6. The method according to claim 5, wherein The method further comprises: According to the preset update frequency, the target parameter values ​​corresponding to the current first neural network are collected and obtained; The model parameters of the second neural network are updated based on the target parameter values ​​to update the preset power flow control model.

7. The method according to claim 5, wherein After the power grid flow control strategy is used to control the current power grid architecture, a second state parameter vector of each power grid element of the current power grid architecture is obtained and a target reward value for controlling the current power grid architecture using the power grid flow control strategy is randomly generated. The method further includes: Based on the first state parameter vector, the first action set, the target reward value, and the second state parameter vector, construct an experience data set corresponding to the first state parameter vector; The experience data set is stored in a predetermined storage area.

8. A power grid flow control device, characterized in that: include: A monitoring module, configured to monitor in real time and obtain a first state parameter vector of each grid element of the current grid architecture, wherein the grid elements include one or more of a substation, a line, a load, and a generator; a strategy matching module, configured to perform strategy matching using a preset power flow control model based on each of the first state parameter vectors to obtain a power flow control strategy corresponding to the current power grid architecture; The control module is used to control the current power grid architecture by adopting the power grid power flow control strategy to obtain a power grid power flow control result.

9. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the power grid power flow control method according to any one of claims 1 to 7 are implemented.

10. An electronic device, characterized in that: The system comprises at least a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the power grid flow control method according to any one of claims 1 to 7 when executing the computer program on the memory.