A method and related apparatus for generating dynamic overload decisions for power transmission networks.
By applying Markov decision theory and deep reinforcement learning models to the power transmission network, and combining reward functions for load shortage penalties and fault penalties, the decision-making complexity of transmission line overload in new power systems is solved, enabling rapid and effective power grid regulation and improving the operational reliability and security of the power transmission network.
Patent Information
- Application Number
- CN202111641984.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-29
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2041-12-29
AI Technical Summary
Existing technologies fail to fully consider the characteristics of new power systems, lack specificity, and cannot guarantee the speed of optimization when there are many control measures, resulting in the inability to meet the complexity of decision-making and speed requirements for transmission line overload.
A dynamic decision generation method for power transmission network overload based on Markov decision theory and deep reinforcement learning model is adopted. By constructing a deep reinforcement learning model and combining a reward function of load shortage penalty, fault penalty and action cost, the power transmission network topology is controlled. The Actor-Critic algorithm is used to train the model to achieve fast decision-making.
It enables efficient and targeted overload control of the transmission network in the new power system, improves the speed and quality of decision-making, ensures that transmission lines are not overloaded, and enhances the reliability and safety of power grid operation.
Smart Images

Figure CN114417710B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power transmission network control technology, and in particular to a method and related apparatus for generating dynamic overload decisions for power transmission networks. Background Technology
[0002] In new power systems, the increasing proportion of renewable energy generation injected into the main grid is significantly impacting grid volatility. Simultaneously, distributed generation is being integrated into the distribution system in large numbers, leading to more flexible and complex operating characteristics. With the gradual transformation of traditional distribution networks, the means and capabilities for interaction between transmission and distribution networks are strengthening, enabling coordinated operation between them.
[0003] Currently, power grid control methods are beginning to shift from single self-control to coordinated control of transmission and distribution networks, especially on the 110kV side. The randomness and variability at the source and load ends of the power system increase the probability of transmission line overload and exceeding limits; while the diversification of implementable control measures also brings complexity to decision-making. Furthermore, the requirements for decision-making speed are increasing. Therefore, for the new power system model, it is crucial to propose targeted power grid control decisions after transmission line overload occurs.
[0004] Modifying the transmission network topology can alleviate line overload problems and improve the reliability and security of grid operation. Previously, grid regulation methods mostly involved altering generation and load, resulting in a simplistic approach and uneconomical cost. The power network, as the carrier of electrical energy, has not had its transmission capacity fully utilized, and research and application of system-level optimization regulation of its topology are limited. In general, without fully considering the new characteristics and problems of the new power system, or with diversified regulation methods, the speed of finding optimal decisions cannot be guaranteed. Summary of the Invention
[0005] This application provides a method and related apparatus for generating dynamic overload decisions for power transmission networks, which addresses the technical problems that existing technologies do not fully consider the characteristics of new power systems, lack specificity, and cannot guarantee the speed of optimization when there are many control measures.
[0006] In view of this, the first aspect of this application provides a method for generating dynamic overload decisions for power transmission networks, comprising:
[0007] The power grid operation status and power flow data under line overload conditions are obtained based on the pre-set power grid simulation model, which includes a simplified transmission network circuit and a simplified distribution network circuit.
[0008] Based on Markov decision theory, a deep reinforcement learning model constructed based on the power flow data is trained according to the power grid operating status to obtain a transmission network topology control model. The deep reinforcement learning model includes a reward function consisting of load shortage penalty, fault penalty and action cost.
[0009] The power grid topology control model is used to perform decision analysis on real-time line overload operation data to obtain the target power grid control decision.
[0010] Optionally, the step of obtaining power grid operating status and power flow data under line overload conditions based on a pre-set power grid simulation model, wherein the pre-set power grid simulation model includes a simplified transmission network circuit and a simplified distribution network circuit, including:
[0011] The transmission and distribution networks in the target area are converted into simplified transmission network circuits and simplified distribution network circuits, respectively.
[0012] Based on the substation node approach, the simplified circuits of the transmission network and the simplified circuits of the distribution network are graphically described to obtain a pre-set power grid simulation model.
[0013] The power grid operating status of the preset power grid simulation model within a preset continuous time step under line overload conditions is obtained, and power flow calculation is performed simultaneously to obtain power flow data.
[0014] Optionally, the deep reinforcement learning model based on the power flow data, constructed according to the power grid operating state, is trained based on Markov decision theory to obtain a transmission network topology control model. The deep reinforcement learning model includes a reward function composed of load shortage penalties, fault penalties, and action costs, comprising:
[0015] Based on Markov decision theory, the control process of the power grid is transformed into a Markov decision process according to the power flow data, resulting in a preset state space, a preset action space, and a reward function.
[0016] The preset state space includes the power grid topology, generator output, total power demand of the distribution network and line transmission power; the preset action space includes bus allocation and line switching; and the reward function includes load shortage penalty, fault penalty and action cost.
[0017] A deep reinforcement learning model is constructed based on the Actor-Critic algorithm, according to the preset state space, the preset action space, and the reward function.
[0018] The deep reinforcement learning model is trained using the power grid operating status to obtain the power grid topology control model.
[0019] Optionally, the step of performing decision analysis on real-time line overload operation data through the transmission network topology control model to obtain target power grid control decisions further includes:
[0020] The target power grid control decision is used to regulate the transmission network and avoid overload of transmission lines.
[0021] A second aspect of this application provides an overload dynamic decision generation device for a power transmission network, comprising:
[0022] The data acquisition module is used to acquire the power grid operation status and power flow data under line overload conditions according to the preset power grid simulation model. The preset power grid simulation model includes a simplified transmission network circuit and a simplified distribution network circuit.
[0023] The model training module is used to train a deep reinforcement learning model based on the power flow data according to the power grid operating status, based on Markov decision theory, to obtain a transmission network topology control model. The deep reinforcement learning model includes a reward function consisting of load shortage penalty, fault penalty and action cost.
[0024] The decision generation module is used to perform decision analysis on real-time line overload operation data through the transmission network topology control model to obtain target power grid control decisions.
[0025] Optionally, the data acquisition module is specifically used for:
[0026] The transmission and distribution networks in the target area are converted into simplified transmission network circuits and simplified distribution network circuits, respectively.
[0027] Based on the substation node approach, the simplified circuits of the transmission network and the simplified circuits of the distribution network are graphically described to obtain a pre-set power grid simulation model.
[0028] The power grid operating status of the preset power grid simulation model within a preset continuous time step under line overload conditions is obtained, and power flow calculation is performed simultaneously to obtain power flow data.
[0029] Optionally, the model training module is specifically used for:
[0030] Based on Markov decision theory, the control process of the power grid is transformed into a Markov decision process according to the power flow data, resulting in a preset state space, a preset action space, and a reward function.
[0031] The preset state space includes the power grid topology, generator output, total power demand of the distribution network and line transmission power; the preset action space includes bus allocation and line switching; and the reward function includes load shortage penalty, fault penalty and action cost.
[0032] A deep reinforcement learning model is constructed based on the Actor-Critic algorithm, according to the preset state space, the preset action space, and the reward function.
[0033] The deep reinforcement learning model is trained using the power grid operating status to obtain the power grid topology control model.
[0034] Optional, also includes:
[0035] The overload control module is used to regulate the transmission network by adopting the target power grid control decision to avoid overload of transmission lines.
[0036] A third aspect of this application provides an overload dynamic decision generation device for a power transmission network, the device including a processor and a memory;
[0037] The memory is used to store program code and transmit the program code to the processor;
[0038] The processor is used to execute the overload dynamic decision generation method for the power transmission network as described in the first aspect, according to the instructions in the program code.
[0039] A fourth aspect of this application provides a computer-readable storage medium for storing program code for executing the overload dynamic decision generation method for power transmission networks described in the first aspect.
[0040] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:
[0041] This application provides a method for generating dynamic overload decisions for a power transmission network, comprising: acquiring power grid operating status and power flow data under line overload conditions based on a pre-set power grid simulation model, the pre-set power grid simulation model including a simplified transmission network circuit and a simplified distribution network circuit; training a deep reinforcement learning model based on power flow data according to the power grid operating status based on Markov decision theory to obtain a power transmission network topology control model, the deep reinforcement learning model including a reward function consisting of load shortage penalty, fault penalty, and action cost; and performing decision analysis on real-time line overload operating data through the power transmission network topology control model to obtain a target power grid control decision.
[0042] The overload dynamic decision generation method for power transmission networks provided in this application employs a deep reinforcement learning model for decision analysis, which is more capable of extracting effective information from the power grid operating environment, thereby obtaining the optimal action decision and achieving high-quality overload control of the power transmission network. Because combining deep learning and reinforcement learning enables one-to-one correspondence learning from perception to action, the trained model can be directly used in the real-time control process of the power transmission network, offering high efficiency and greater targeting. Therefore, this application addresses the technical problems of existing technologies failing to fully consider the characteristics of new power systems, lacking specificity, and being unable to guarantee optimization speed when multiple control methods are available. Attached Figure Description
[0043] Figure 1 A flowchart illustrating a method for generating dynamic overload decisions for a power transmission network, provided in an embodiment of this application;
[0044] Figure 2 A schematic diagram of the structure of an overload dynamic decision generation device for a power transmission network provided in an embodiment of this application;
[0045] Figure 3 This is a schematic diagram illustrating the relationship between the simulation process of the distribution network and the transmission network provided in the embodiments of this application;
[0046] Figure 4 This is a theoretical schematic diagram of the deep reinforcement learning model provided in the embodiments of this application. Detailed Implementation
[0047] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0048] For easier understanding, please refer to Figure 1 An embodiment of a dynamic overload decision generation method for power transmission networks provided in this application includes:
[0049] Step 101: Obtain the power grid operation status and power flow data under line overload conditions based on the pre-set power grid simulation model. The pre-set power grid simulation model includes a simplified transmission network circuit and a simplified distribution network circuit.
[0050] Further, step 101 includes:
[0051] The transmission and distribution networks in the target area are converted into simplified transmission network circuits and simplified distribution network circuits, respectively.
[0052] Based on the substation node approach, the simplified circuits of the transmission network and the simplified circuits of the distribution network are graphically described to obtain a pre-set power grid simulation model;
[0053] The power grid operation status of the preset power grid simulation model within a preset continuous time step under line overload conditions is obtained, and power flow calculation is performed to obtain power flow data.
[0054] By using a power grid simulator based on directly available information about the transmission and distribution network, a pre-built power grid simulation model can be constructed, as shown in this embodiment. The parameters in the model can be calculated, entered, and stored. This pre-built power grid simulation model allows for the acquisition of information such as power grid power, current, and changes in power grid topology, establishing a user-friendly interactive environment.
[0055] For details, please refer to Figure 3 This model clarifies the relationship between the transmission and distribution networks. The power grid can be described as a graph composed of nodes, each corresponding to a substation connected to loads, generators, and power lines. Generators produce electricity, loads consume electricity, and transmission lines transmit electricity between substations. A substation can be viewed as a router in the network, determining the location of power transmission. Common busbar connection methods in transmission networks (double busbar connection, single busbar segmented with bypass, 3 / 2 connection) can all operate in split mode. Without loss of generality, we assume here that the substation uses a double busbar connection, where the transmission lines connected to the substation can be distributed to one of the two busbars, and power is transmitted only through components on the same busbar. Therefore, each substation under busbar splitting can be considered as two nodes. Transmission and distribution networks are connected by boundary equipment, and distribution networks are connected by tie switches, which are open during normal grid operation.
[0056] Once the pre-set power grid simulation model is built, it can generate continuous time-step power grid operating states. Therefore, it can obtain the power grid operating states within a pre-set continuous time step, and can also perform power flow calculations to obtain power flow data. The power grid operating states include: the power grid topology, the output of each generator, and the total power required by each distribution network; the power flow data mainly includes information such as the power transmitted by each line, and specific information can be added or removed according to the actual situation, and appropriate information can be obtained.
[0057] Step 102: Based on Markov decision theory, train the deep reinforcement learning model built on power flow data according to the power grid operation status to obtain the transmission network topology control model. The deep reinforcement learning model includes a reward function consisting of load power shortage penalty, fault penalty and action cost.
[0058] Further, step 102 includes:
[0059] Based on Markov decision theory, the control process of the power grid is transformed into a Markov decision process according to power flow data, resulting in a preset state space, a preset action space, and a reward function.
[0060] The pre-set state space includes the power grid topology, generator output, total power demand of the distribution network and line transmission power; the pre-set action space includes bus allocation and line switching; and the reward function includes load shortage penalty, fault penalty and action cost.
[0061] Based on the Actor-Critic algorithm, a deep reinforcement learning model is constructed according to the pre-set state space, pre-set action space and reward function;
[0062] The deep reinforcement learning model is trained using the power grid operating status to obtain the power grid topology control model.
[0063] A Markov Decision Process (MDP) is a mathematical model of sequential decision-making used to simulate stochastic policies and rewards achievable by an agent in environments where the system state exhibits Markov properties. A Markov Decision Process includes states, actions, policies, and rewards. For details, please refer to [link to relevant documentation]. Figure 4 In a Multi-Level Processing (MDP), the agent is a proxy for machine learning. It can perceive the state of the external environment, make decisions, take actions in response to the environment, and adjust its decisions based on feedback from the environment. The environment is the collection of all things outside the agent in the MDP model. Its state is affected by the agent's actions, and these changes can be fully or partially perceived by the agent. After each decision, the environment may provide the agent with a corresponding reward.
[0064] The control process of the transmission network is transformed into the determination of the preset state space, preset action space, and reward function. The preset state space includes the topology of the transmission and distribution networks (the connectivity of each transmission line and the bus allocation of each substation, as well as the connectivity of the transmission and distribution network boundary lines and the state of the tie switches between distribution networks), the output of each generator, the total power required by each distribution network, and the power transmitted by each line; the upper and lower limits of these state variables together determine the state space. The preset action space includes operations such as bus allocation and line switching. The model allows actions to be applied to the proportion of renewable energy on the generation side, the power consumption on the load side, substations, lines, and tie switches at the boundary of the transmission and distribution networks to regulate the grid; among them, the action on the substation is called bus allocation, and each transmission line of the substation can be connected to one of the two buses; the action on the transmission line is called line switching, which involves disconnecting a line or reconnecting a disconnected line, including tie lines at the boundary of the transmission and distribution networks. Therefore, each action includes the following action variables: the bus number n (0 or 1) selected for connection to each transmission line in the substation, and the on / off status in_service (0 or 1) of each transmission line. The reward function includes load shortage penalty, fault penalty, and action cost, and can be expressed as:
[0065]
[0066] Among them, t over This is the number of time steps completed in this round of reinforcement learning, prod t It is the total power generation of the power grid at time step t, load t It is the total load of the power grid at time step t, (prod t -load t The total active power loss of the power grid at time step t is represented by the failure penalty, which is the sum of constant penalties for the remaining time steps when the current cycle ends. t C(a) is the action taken by the decision-maker at time step t. t ) is action a t The goal of training this reinforcement learning model is to simultaneously minimize the active power loss of the power grid, avoid faults, and achieve the cost of grid control.
[0067] The Actor-Critic algorithm, simply put, is a process where the Actor function or model and the Critic function or model collide and optimize their training to achieve better results. The Actor takes a corresponding action based on its current state, and the Critic scores the Actor's performance based on the state and the corresponding action. The Actor, in turn, adjusts the parameters in its function or model (i.e., its training strategy) based on the Critic's scores to obtain better scores. The Critic adjusts its own scoring strategy based on the system's rewards and other scores to achieve more accurate and reliable ratings.
[0068] This embodiment employs two neural networks: a target network for evaluating actions and a decision network for selecting actions. Both networks have identical structures and initial parameters, but the target network's parameters are updated with a delay, thereby reducing data correlation. Furthermore, the control measures in this embodiment include: power grid generation regulation, power load regulation, power grid topology and parameter control, and transmission-distribution coordination. A reinforcement learning algorithm model is used to autonomously learn the dynamic control process of the power grid after line overload. The trained deep reinforcement learning agent can provide fast and autonomous topology control strategies, making real-time decisions on the overload state of the power grid lines to maintain normal power supply to the power system.
[0069] Based on the above, the training process of the power grid topology control model is as follows: First, a memory unit D with N data entries is set up, and the number of traversal rounds is episode = 0, 1, 2, ..., M, where M is the maximum number of rounds. The environment generates the initial state s0 = [{top}, {P}. generator},{P load},{P line The time steps for traversal are step = 0, 1, 2, ..., T, with 15-minute intervals, dividing the day into 96 time steps, T = 96; then, the decision network is initialized, weights ω are randomly generated, and a target network is generated with the same structure and parameters, whose weights are expressed as ω', and ω' = ω; then, the target network uses the ∈-greedy strategy to generate action a. t The system selects a random action with probability ε, or chooses the action in the current state as the output result; then executes the action in the power grid simulation environment. If the action is valid, the power flow value is solved, and the reward value r is calculated according to the rules. t Otherwise, the round ends, and a failure penalty r is applied. t =r fail The reward r t and the corresponding new state s t+1 The state transition process constitutes (s) t ,a t,r t ,s t+1 The data is stored in memory unit D; finally, a sample of data (s) is uniformly and randomly sampled from D. j ,a j ,r j ,s j+1 Construct the input as (s) j ,a j The output is The decision network is trained using training samples. This can be understood as copying the parameters of the decision network to the target network every τ steps, using the "input s" generated by the target network. j The output is The training samples are used to train the parameters in the network.
[0070] Step 103: Analyze the real-time line overload operation data using the transmission network topology control model to obtain the target power grid control decision.
[0071] Real-time line overload operation data is data from the online operation of the transmission network. The trained model can be directly used in the operating transmission network for decision analysis and to obtain target control and rejection data.
[0072] Furthermore, step 103, followed by:
[0073] The target power grid control decision is used to regulate the transmission network and avoid overload of transmission lines.
[0074] The overload dynamic decision generation method for power transmission networks provided in this application uses a deep reinforcement learning model for decision analysis, which is more capable of extracting effective information from the power grid operating environment, thereby obtaining the optimal action decision and achieving high-quality overload control of the power transmission network. Because combining deep learning and reinforcement learning enables one-to-one learning from perception to action, the trained model can be directly used in the real-time control process of the power transmission network, offering high efficiency and greater targeting. Therefore, this application's embodiments can solve the technical problems of existing technologies not fully considering the characteristics of new power systems, lacking targeting, and failing to guarantee optimization speed when there are many control methods.
[0075] For easier understanding, please refer to Figure 2 This application provides an embodiment of an overload dynamic decision generation device for a power transmission network, comprising:
[0076] The data acquisition module 201 is used to acquire the power grid operation status and power flow data under line overload conditions according to the preset power grid simulation model. The preset power grid simulation model includes a simplified transmission network circuit and a simplified distribution network circuit.
[0077] The model training module 202 is used to train a deep reinforcement learning model based on power flow data according to the power grid operating status based on Markov decision theory, so as to obtain a transmission network topology control model. The deep reinforcement learning model includes a reward function consisting of load power shortage penalty, fault penalty and action cost.
[0078] The decision generation module 203 is used to perform decision analysis on real-time line overload operation data through the transmission network topology control model to obtain the target power grid control decision.
[0079] Furthermore, the data acquisition module 201 is specifically used for:
[0080] The transmission and distribution networks in the target area are converted into simplified transmission network circuits and simplified distribution network circuits, respectively.
[0081] Based on the substation node approach, the simplified circuits of the transmission network and the simplified circuits of the distribution network are graphically described to obtain a pre-set power grid simulation model;
[0082] The power grid operation status of the preset power grid simulation model within a preset continuous time step under line overload conditions is obtained, and power flow calculation is performed to obtain power flow data.
[0083] Furthermore, the model training module 202 is specifically used for:
[0084] Based on Markov decision theory, the control process of the power grid is transformed into a Markov decision process according to power flow data, resulting in a preset state space, a preset action space, and a reward function.
[0085] The pre-set state space includes the power grid topology, generator output, total power demand of the distribution network and line transmission power; the pre-set action space includes bus allocation and line switching; and the reward function includes load shortage penalty, fault penalty and action cost.
[0086] Based on the Actor-Critic algorithm, a deep reinforcement learning model is constructed according to the pre-set state space, pre-set action space and reward function;
[0087] The deep reinforcement learning model is trained using the power grid operating status to obtain the power grid topology control model.
[0088] Furthermore, it also includes:
[0089] The overload control module 204 is used to control the transmission network by adopting the target power grid control decision to avoid overload of transmission lines.
[0090] This application also provides an overload dynamic decision generation device for a power transmission network, the device including a processor and a memory;
[0091] The memory is used to store program code and transfer the program code to the processor;
[0092] The processor is used to execute the overload dynamic decision generation method for the power transmission network in the above method embodiment according to the instructions in the program code.
[0093] This application also provides a computer-readable storage medium for storing program code for executing the overload dynamic decision generation method for power transmission networks in the above method embodiments.
[0094] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0095] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0096] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0097] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for executing all or part of the steps of the methods described in the various embodiments of this application through a computer device (which may be a personal computer, server, or network device, etc.). The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.
[0098] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for generating dynamic overload decisions for a power transmission network, characterized in that, include: The power grid operation status and power flow data under line overload conditions are obtained based on the pre-set power grid simulation model, which includes a simplified transmission network circuit and a simplified distribution network circuit. Based on Markov decision theory, a deep reinforcement learning model constructed from the power flow data is trained according to the power grid operating state to obtain a transmission network topology control model. The deep reinforcement learning model includes a reward function composed of load shortage penalty, fault penalty, and action cost. The specific process is as follows: Based on Markov decision theory, the control process of the power grid is transformed into a Markov decision process according to the power flow data, resulting in a preset state space, a preset action space, and a reward function. The preset state space includes the power grid topology, generator output, total power demand of the distribution network and line transmission power; the preset action space includes bus allocation and line switching; and the reward function includes load shortage penalty, fault penalty and action cost. A deep reinforcement learning model is constructed based on the Actor-Critic algorithm, according to the preset state space, the preset action space, and the reward function. The deep reinforcement learning model is trained using the power grid operating status to obtain the power grid topology control model; The power grid topology control model is used to perform decision analysis on real-time line overload operation data to obtain the target power grid control decision.
2. The overload dynamic decision generation method for power transmission networks according to claim 1, characterized in that, The process involves obtaining power grid operating status and power flow data under line overload conditions based on a pre-set power grid simulation model. This pre-set power grid simulation model includes simplified transmission network circuits and simplified distribution network circuits, comprising: The transmission and distribution networks in the target area are converted into simplified transmission network circuits and simplified distribution network circuits, respectively. Based on the substation node approach, the simplified circuits of the transmission network and the simplified circuits of the distribution network are graphically described to obtain a pre-set power grid simulation model. The power grid operating status of the preset power grid simulation model within a preset continuous time step under line overload conditions is obtained, and power flow calculation is performed simultaneously to obtain power flow data.
3. The overload dynamic decision generation method for power transmission networks according to claim 1, characterized in that, The step of performing decision analysis on real-time line overload operation data using the transmission network topology control model to obtain target power grid control decisions further includes: The target power grid control decision is used to regulate the transmission network and avoid overload of transmission lines.
4. An overload dynamic decision generation device for a power transmission network, characterized in that, include: The data acquisition module is used to acquire the power grid operation status and power flow data under line overload conditions according to the preset power grid simulation model. The preset power grid simulation model includes a simplified transmission network circuit and a simplified distribution network circuit. The model training module is used to train a deep reinforcement learning model constructed based on the power flow data according to the power grid operating state, based on Markov decision theory, to obtain a transmission network topology control model. The deep reinforcement learning model includes a reward function composed of load shortage penalty, fault penalty, and action cost. Specifically, the model training module is used for: Based on Markov decision theory, the control process of the power grid is transformed into a Markov decision process according to the power flow data, resulting in a preset state space, a preset action space, and a reward function. The preset state space includes the power grid topology, generator output, total power demand of the distribution network and line transmission power; the preset action space includes bus allocation and line switching; and the reward function includes load shortage penalty, fault penalty and action cost. A deep reinforcement learning model is constructed based on the Actor-Critic algorithm, according to the preset state space, the preset action space, and the reward function. The deep reinforcement learning model is trained using the power grid operating status to obtain the power grid topology control model; The decision generation module is used to perform decision analysis on real-time line overload operation data through the transmission network topology control model to obtain target power grid control decisions.
5. The overload dynamic decision generation device for power transmission networks according to claim 4, characterized in that, The data acquisition module is specifically used for: The transmission and distribution networks in the target area are converted into simplified transmission network circuits and simplified distribution network circuits, respectively. Based on the substation node approach, the simplified circuits of the transmission network and the simplified circuits of the distribution network are graphically described to obtain a pre-set power grid simulation model. The power grid operating status of the preset power grid simulation model within a preset continuous time step under line overload conditions is obtained, and power flow calculation is performed simultaneously to obtain power flow data.
6. The overload dynamic decision generation device for power transmission networks according to claim 4, characterized in that, Also includes: The overload control module is used to regulate the transmission network by adopting the target power grid control decision to avoid overload of transmission lines.
7. An overload dynamic decision generation device for a power transmission network, characterized in that, The device includes a processor and a memory; The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the overload dynamic decision generation method for the power transmission network according to any one of claims 1-3, based on the instructions in the program code.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program code for executing the overload dynamic decision generation method for the power transmission network according to any one of claims 1-3.
Citation Information
Patent Citations
Active safety verification method for power transmission network line based on extreme learning machine
CN107069708A
Power grid reactive voltage distributed control method and system
CN111799808A