A method and system for improving the transmission capacity of AC / DC hybrid power transmission corridors
By selecting key transmission corridors in AC/DC hybrid power grids, identifying controllable equipment, constructing Markov decision process models, and employing maximum entropy secure deep reinforcement learning methods to optimize the control of DC lines and generator sets, the problem of limited transmission capacity was solved, thereby improving transmission capacity and enhancing the level of renewable energy consumption.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2026-03-10
AI Technical Summary
The existing power grid planning and control decision-making system is unable to cope with the large-scale fluctuations of new energy sources and the complexity of AC/DC hybrid operation, resulting in limited power transmission capacity, especially in key power transmission corridors where there are problems of excessive and unbalanced power transmission pressure.
A method for enhancing the transmission capacity of AC/DC hybrid transmission corridors is adopted. By selecting key transmission corridors, identifying controllable equipment, conducting sensitivity analysis, constructing a Markov decision process model, and using the maximum entropy security deep reinforcement learning method to solve for the optimal control strategy, the control of DC lines and generator sets is optimized.
It enables flexible and controllable power flow distribution in key transmission corridors, improves transmission capacity and renewable energy consumption, and provides rapid adaptability and effective control strategies.
Smart Images

Figure CN119813402B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to power systems, and more particularly to a method and system for improving the transmission capacity of AC / DC hybrid transmission corridors. Background Technology
[0002] With the large-scale grid connection of massive renewable energy and power electronic equipment, the operation of ultra-high voltage AC / DC hybrid systems, the large-scale cross-regional power transmission, and the continuous increase in the proportion of new loads, the operating characteristics of new power systems have undergone profound and complex changes. The high randomness and strong uncertainty of its source and load bring huge challenges to power system planning decisions and online safety and stability. Under extreme faults, local disturbances evolve into cascading faults, and the risk of systemic instability increases significantly. To meet such challenges, the existing power grid planning, operation, and control decision-making system and methods have many limitations, specifically: (1) Existing operation and control decisions are mostly based on offline simulation analysis plus experience judgment. Although multiple scenarios are considered, it is difficult to cope with the many possibilities brought about by the large-scale fluctuations of high-proportion new energy sources and the variable operating characteristics such as AC / DC hybrid operation. It cannot cover all operating mode scenarios to effectively cope with extreme power grid conditions. (2) Power grid operation and management data come from different business systems and still have redundancy and quality problems. It is difficult to effectively extract typical and abnormal operating modes and deeply explore their potential mechanisms. Due to the limitations of the simulation analysis speed and modeling accuracy of large power grids, traditional methods are difficult to form optimal control decisions from massive operating data and simulation results. (3) Regional energy management systems typically lack in-depth exploration, effective accumulation, and application of historical experience; at the decision-making level, there is a lack of effective utilization of load elasticity, and traditional algorithms rely heavily on precise models, which are inflexible and difficult to support the information management and control of massive adjustable resources in the future. In addition, the key transmission corridors of AC / DC hybrid power grids are constrained by safety and stability, resulting in severely limited power transmission capacity.
[0003] Taking the Jiangsu power grid as an example, with the increase in installed capacity of new energy sources in northern Jiangsu, many DC transmission lines are operating at full capacity. It is estimated that by 2026, the maximum transmission power of Jiangsu's river-crossing channels (including AC and DC channels) will reach 32.19 million kilowatts, exceeding the transmission capacity by approximately 7 million kilowatts. Furthermore, the transmission power of each river-crossing channel is severely uneven, with some channels frequently exceeding their limits, leading to a sharp increase in transmission pressure. Specifically, the maximum transmission power of the central channel is 5.04 million kilowatts, slightly exceeding the channel's maximum transmission capacity; the maximum transmission power of the eastern channel is 14.14 million kilowatts, exceeding the channel's transmission capacity by approximately 6 million kilowatts. Therefore, it is urgent to strengthen the transmission capacity of the central and eastern channels.
[0004] Traditional power system transmission corridors consist only of AC lines, exhibiting a natural power distribution. When the power flow of one line reaches its limit, it restricts the power transmission of the entire section. To enhance the transmission capacity of AC sections, flexible and controllable power flow distribution methods are needed to maximize transmission capacity. Taking the Jiangsu power grid as an example, the Yanhuai, Xitai, Jiansu, and Yangzhen DC lines have been put into operation. Utilizing the flexible and controllable power characteristics of DC lines, the power flow of heavily loaded lines is transferred, further enhancing the power transmission capacity of transmission sections within the AC system where generation centers and load centers are mismatched. This is an important means to improve the power transmission capacity of AC / DC hybrid transmission corridors. However, there is currently limited research on power flow control technologies for enhancing the transmission capacity of AC / DC hybrid power grids. Existing research mainly focuses on local power flow optimization control based on FACTS power electronic devices.
[0005] Therefore, how to fully explore and utilize limited power transmission channel resources, improve the level of new energy consumption, and enhance the transmission capacity of important sections are urgent problems that need to be solved in current dispatching, operation, and safe power supply. Summary of the Invention
[0006] To address the above problems, this invention provides a method and system for improving the transmission capacity of AC / DC hybrid power transmission corridors.
[0007] The technical solution of this invention is: a method for improving the transmission capacity of an AC / DC hybrid power transmission corridor, comprising the following steps:
[0008] Step 1): Select the key transmission corridor for the AC / DC hybrid system;
[0009] Step 2): Identify controllable equipment for power transmission in key transmission corridors of AC / DC hybrid systems;
[0010] Step 3): Perform sensitivity analysis on controllable devices to screen out high-sensitivity control devices;
[0011] Step 4): Based on steps 2)-3), construct a Markov decision process model;
[0012] Step 5): Based on the maximum entropy secure deep reinforcement learning method, solve the Markov decision process model to obtain the optimal control strategy.
[0013] In step 1), a key transmission corridor is selected, and the AC lines L within that corridor are determined. AC (i,j) and DC line L DC (m,n) and AC line capacity P AC max(i,j), DC line capacity P DCmax(m,n), where i and j are the starting and ending busbars connected to the AC line, respectively; and m and n are the starting and ending busbars connected to the DC line, respectively.
[0014] In step 2), the controllable equipment includes: a DC line and a controllable conventional generator set, wherein P gen (k) represents the active power output of a controllable conventional generator set;
[0015] All controllable DC lines are selected as control measures to construct a DC line control set P. DC .
[0016] In step 3), historical power flow data files are collected, and sensitivity calculations are performed on the set of controllable conventional generator sets, including:
[0017] 3.1) Using historical power flow data files, the total power (P) of the selected critical transmission corridor is calculated using AC power flow calculation methods. 走廊 =P 交流 +P 直流 ,
[0018] Among them, P 交流 It is the sum of the power of all AC lines in the transmission corridor, P. 交流 =∑ i,j P AC (i,j); P AC (i,j) represents the active power of the AC line;
[0019] P 直流 It is the sum of the power of all DC lines in the transmission corridor, P. 直流 =∑ m,n P DC (m,n); P DC (m,n) represents the active power of the DC line;
[0020] 3.2) For the k-th controllable conventional generator set, increase the active power output ΔP based on the current active power output. gen (k), resolve the AC power flow equations, and calculate the applied ΔP gen (k) Total power of the transmission corridor The sensitivity of the active power of the k-th controllable conventional generator unit to changes in the power of the transmission corridor is derived from the following formula:
[0021]
[0022] 3.3) Following step 3.2), iterate through all controllable conventional generator sets to obtain a sensitivity list; set a threshold Sensitivity_Tgen, remove generators with a sensitivity lower than this value, and thus obtain a list Ggen of controllable generators with a sensitivity higher than Sensitivity_Tgen.
[0023] In step 4), the Markov decision process model includes:
[0024] The state space is a one-dimensional vector of system information obtained after solving the power flow problem for the power grid operation mode, S = [V m V a ,P ac Q ac ,P dc ,P gen ,P renewable ], where V m It is the AC bus voltage magnitude vector, V a It is the phase angle vector of the AC bus voltage, P ac It is the active power vector of AC lines, Q ac It is the reactive power vector of AC lines, P dc It is the active power vector of DC lines, P gen It is the active power vector of a conventional generator set, P renewable It is the active power vector of new energy generator sets;
[0025] The action space of multi-type intelligent agents is:
[0026] Type 1: The intelligent agent of a conventional generator set is the vector of active power change ΔP of the conventional generator set. gen =[ΔP gen(1) ,ΔP gen(2) ,…,ΔP gen(g) ], g∈Ggen;
[0027] Type 2: The intelligent agent of a DC line is the vector of active power change ΔP of the DC line. dc =[ΔP dc(1) ,ΔP dc(2) ,…,ΔP dc(h) ], h∈P DC ;
[0028] Reward function:
[0029] Reward=β1r1+β2r2-β3r3-β4r4-β5r5
[0030] in:
[0031] r1 represents the power value that increases the transmission capacity of the power transmission corridor;
[0032] r2 represents the increase in the total power of DC lines within the transmission corridor;
[0033] r3 represents the total number of AC line power exceeding limits within the power system;
[0034] r4 represents the total number of transformer power exceeding limits within the power system;
[0035] r5 represents the non-convergence penalty for AC power flow calculation;
[0036] β1, β2, β3, β4, and β5 are weighting coefficients used to balance different types of reward values.
[0037] In step 5), for the DC line agent and the conventional generator agent, the reinforcement learning agent model is trained according to the state space, action space and reward function described in step 4) to obtain the optimal control strategy.
[0038] The training process of the reinforcement learning agent model includes:
[0039] First, collect historical operating condition power flow files and perform power flow solution;
[0040] Secondly, the state space is extracted, and based on the state vector, the agent outputs control actions with the goal of maximizing the reward function value.
[0041] Then, control actions are applied to the current system state, the AC power flow is solved and the agent reward function is calculated, while the samples saved in tuple form are stored in the cache.
[0042] Finally, determine the termination condition for agent training. If the termination condition is met, exit the training process and save the agent; if the termination condition is not met, extract the next flow file and continue training the agent.
[0043] In a reinforcement learning agent model, the weights and biases are represented by θ. The agent learns to update θ in order to obtain the maximum reward value.
[0044] When solving the problem, the maximum entropy safe reinforcement learning algorithm is adopted. Based on the maximum entropy reinforcement learning, the safe reinforcement learning introduces a safety constraint g(s)≤0, where g(s) is a function that measures the safety level of state s.
[0045] The objective function of maximum entropy secure reinforcement learning is expressed as:
[0046] max_πJ(π)+αH(π)stE[g(s)]≤0
[0047] In the formula, π is the strategy, J(π) is the expected reward obtained when applying strategy π, α is the coefficient of the entropy regularization term, H(π) is the entropy of strategy π, and E is the mathematical expectation function.
[0048] Transform the constrained optimization problem into Lagrangian form:
[0049] L(π,λ)=J(π)+αH(π)-λE[g(s)]
[0050] In the formula, λ is the Lagrange multiplier;
[0051] Using the policy gradient method, the policy update rule is obtained:
[0052]
[0053] In the formula, L is the objective function after transforming the constrained optimization problem into Lagrange form, logπ represents the logarithmic form of policy π, θ(a|s) represents the probability of policy choosing action a in state s, and Q^π(s,a) represents the expected reward obtained after taking action a from state s under policy π.
[0054] It also includes step 6), which is as follows:
[0055] During the real-time dispatch and operation of the power grid, system data under the current power grid operation mode is extracted, power flow is solved and intelligent agent input vector is constructed. The reinforcement learning intelligent agent model is called to give the optimal decision values of DC lines and conventional generator units. Power flow calculation is performed to verify the effectiveness of the intelligent agent decision results.
[0056] If the control strategy verification is successful, the data will be sent to the DC power station and generator set.
[0057] If the control strategy is not verified, the data will not be sent to the station or generator set.
[0058] A system for enhancing the transmission capacity of an AC / DC hybrid power transmission corridor includes:
[0059] The selection module is used to select key transmission corridors for AC / DC hybrid systems;
[0060] Controllable modules are used to identify controllable devices for power transmission in critical transmission corridors of AC / DC hybrid systems.
[0061] The analysis module is used to perform sensitivity analysis on controllable devices and screen out high-sensitivity control devices.
[0062] The decision module is used to build Markov decision process models;
[0063] The solution module is used to solve Markov decision process models based on the maximum entropy secure deep reinforcement learning method to obtain the optimal control strategy.
[0064] In this invention, key transmission corridors are selected, controllable equipment is identified, and more effective control equipment for enhancing the transmission capacity of these corridors is screened. A Markov decision process model is constructed and solved to derive the optimal control strategy. This method provides an effective solution to the limited transmission capacity of key transmission corridors. The trained agent, based on DC line power and generator active power, offers advantages in speed and adaptability compared to traditional control strategies derived from complete system mechanism models.
[0065] This invention utilizes both DC line power and generator power in the cross-river section to comprehensively enhance the transmission capacity of the power transmission corridor. Attached Figure Description
[0066] Figure 1 This is a flowchart of the method of the present invention.
[0067] Figure 2 This is a diagram illustrating the training process of a reinforcement learning agent model.
[0068] Figure 3 This is a schematic diagram of the reward value during the agent training process in this embodiment.
[0069] Figure 4 This is a schematic diagram illustrating the control effect of the intelligent agent on increasing the total power of the power transmission corridor in the embodiment.
[0070] Figure 5 This is a schematic diagram illustrating the control effect of the intelligent agent on the improvement of the power transmission capacity of the power transmission corridor in the embodiment. Detailed Implementation
[0071] like Figure 1 As shown, the present invention provides a method for improving the transmission capacity of an AC / DC hybrid power transmission corridor, comprising the following steps:
[0072] Step 1): Select the key transmission corridor of the AC / DC hybrid system; thus, the key transmission corridor is defined as the controlled object.
[0073] Step 2): Identify controllable equipment for power transmission in key transmission corridors of the AC / DC hybrid system; for the transmission corridors defined in Step 1), specify the set of controllable equipment, including DC line power and generator power output;
[0074] Step 3): Perform sensitivity analysis on controllable equipment to screen out high-sensitivity control equipment; by calculating sensitivity information, screen out more effective control equipment for improving the transmission capacity of key power transmission corridors.
[0075] Step 4): Based on steps 2)-3), construct a Markov decision process model;
[0076] Step 5): Based on the maximum entropy secure deep reinforcement learning method, solve the Markov decision process model to obtain the optimal control strategy.
[0077] This invention utilizes both DC line power and generator power in the river-crossing section to comprehensively enhance the transmission capacity of the power transmission corridor. Compared to traditional methods, the method proposed in this invention can instantly provide an optimization strategy, demonstrating its superiority.
[0078] The specific steps are described below:
[0079] Step 1): Select the key transmission corridor of the AC / DC hybrid system and determine the AC line L within the corridor. AC (i,j) and DC line L DC (m,n) and AC line capacity P AC max(i,j), DC line capacity P DC max(m,n). Where L AC (i,j) represents the AC line, where i and j are the starting and ending busbars connected to the AC line, respectively; L DC (m,n) represents a DC line, where m and n are the starting and ending busbars of the DC line, respectively. AC stands for alternating current; DC stands for direct current.
[0080] Step 2): Identify controllable equipment for power transmission in key transmission corridors of the AC / DC hybrid system, including: DC lines and controllable conventional generator sets, where P gen (k) represents the active power output of a controllable conventional generator set. All controllable DC lines are selected as control measures to construct a DC line control set P. DC .
[0081] Step 3): For conventional generator sets, sensitivity analysis is required to screen out high-sensitivity units for control measures. Historical power flow data files are collected, and sensitivity calculations are performed on the set of controllable conventional generator sets, as follows:
[0082] 3.1) Using historical power flow data files, the total power (P) of the selected critical transmission corridor is calculated using AC power flow calculation methods. 走廊 =P 交流 +P 直流 , where P 交流 It is the sum of the power of all AC lines in the transmission corridor, i.e., P 交流 =∑ i,j P AC (i,j); P AC (i,j) represents the active power of the AC line; P 直流 It is the sum of the power of all DC lines in the transmission corridor, P. 直流 =∑ m,nP DC (m,n); P DC (m,n) represents the active power of a DC line.
[0083] 3.2) For the k-th controllable conventional generator set, increase the active power output ΔP based on the current active power output. gen (k), resolve the AC power flow equations, and calculate the applied ΔP gen (k) Total power of the transmission corridor The sensitivity of the active power of the k-th controllable conventional generator unit to changes in the power of the transmission corridor is derived from the following formula:
[0084]
[0085] 3.3) Following step 3.2), iterate through all controllable conventional generator sets to obtain a sensitivity list. Set a threshold Sensitivity_Tgen, and remove generators with a sensitivity lower than this value to obtain a list Ggen of controllable generators with a sensitivity higher than Sensitivity_Tgen.
[0086] Step 4): The problem of improving the power transmission capacity of key transmission corridors in the AC / DC hybrid system is constructed as a Markov decision process, and a cooperative intelligent agent model for multiple types of control equipment is designed. An optimization scheduling model considering resource operation constraints and system security constraints is constructed, and a constrained multi-agent cooperative control model is designed. The above Markov decision process model with a continuous action space transforms the real-time optimization scheduling of the system into a multi-agent optimization problem. Each agent uses its own policy function as the decision variable, maximizing its own value function while satisfying security and cost constraints. The Markov decision process model includes the following main parts:
[0087] The state space is a one-dimensional vector of system information obtained after solving the power flow problem for the power grid operation mode, S = [V m V a ,P ac Q ac ,P dc ,P gen ,P renewable ], where V m It is the AC bus voltage magnitude vector, V a It is the phase angle vector of the AC bus voltage, P ac It is the active power vector of AC lines, Q ac It is the reactive power vector of AC lines, P dc It is the active power vector of DC lines, P gen It is the active power vector of a conventional generator set, P renewable It is the active power vector of the new energy generator set.
[0088] The action space of multi-type intelligent agents is:
[0089] Type 1: The intelligent agent of a conventional generator set is the vector of active power change ΔP of the conventional generator set. gen =[ΔP gen(1) ,ΔP gen(2) ,…,ΔP gen(g) ], g∈Ggen.
[0090] Type 2: The intelligent agent of a DC line is the vector of active power change ΔP of the DC line. dc =[ΔP dc(1) ,ΔP dc(2) ,…,ΔP dc(h) ], h∈P DC .
[0091] Reward Function: To effectively improve the power transmission capacity of critical transmission corridors, the reward function design includes rewarding increased transmission capacity, rewarding increased DC power, penalizing line overruns, penalizing non-convergence in power flow calculations, and penalizing equipment overload. The agent reward function design is as follows:
[0092] Reward=β1r1+β2r2-β3r3-β4r4-β5r5
[0093] in:
[0094] r1=ΔC transfer This represents the power increase in the transmission capacity of the power transmission corridor. The transmission capacity of a power transmission corridor is defined as:
[0095]
[0096] r2=ΔP 直流 This represents the increase in the total power of DC lines within the transmission corridor.
[0097] This represents the total number of AC lines that exceed their power limits within the power system. Here, H is the number of lines that exceed their power limits, Pline represents the active power of the AC lines, and Plinemax represents the upper limit of the active power of the AC lines.
[0098] This represents the total number of transformers exceeding their power limits within the power system, where K is the number of transformers exceeding their power limits, Xfm represents the active power of the transformers, and Xfm_max represents the upper limit of the active power of the transformers.
[0099] r5 represents the non-convergence penalty for AC power flow calculation.
[0100] β1, β2, β3, β4, and β5 are weighting coefficients used to balance different types of reward values.
[0101] Step 5): A secure deep reinforcement learning method based on maximum entropy is used to solve the above Markov decision process model and derive the optimal control strategy.
[0102] When multiple agents can simultaneously satisfy multiple control objectives and meet safety constraints, the agent joint strategy can achieve online optimization decision-making for improving power transmission capacity.
[0103] Train the DC line agent and the conventional generator set agent. Based on the state space, action space, and reward function described in step 4), train the reinforcement learning agent.
[0104] like Figure 2 As shown, the training process of the reinforcement learning agent model includes:
[0105] Historical power flow data files under operating conditions are collected. For each power flow data problem, AC power flow calculation is first performed. After the calculation iteratively converges, the state space is extracted, S = [V m V a ,P ac Q ac ,P dc ,P gen ,P renewable Based on the state vector, the agent outputs a control action with the goal of maximizing the reward function value, [ΔP]. gen ,ΔP dc The control action is applied to the current system state, the AC power flow is solved and the agent's reward function is calculated. At the same time, the samples saved in tuple form are stored in the cache to record the agent's control effect. The neural network parameters in the agent are updated. Furthermore, the termination condition of the agent training is determined. If the termination condition is met, the training process is exited and the agent is saved. If the termination condition is not met, the next power flow file is extracted and the agent continues to be trained.
[0106] Step 6): During the real-time dispatch operation of the power grid, extract system data under the current power grid operation mode, solve the power flow problem, construct the agent input vector S, call the agent model to provide optimized decision values for DC lines and conventional generator sets, and perform power flow calculations to verify the effectiveness of the agent's decision results. If the control strategy verification is passed, the decision is sent to the DC substations and generator sets; if the control strategy verification is not passed, the decision is not sent to the substations or generator sets.
[0107] This invention proposes a method for enhancing the transmission capacity of power transmission corridors based on maximum entropy deep reinforcement learning. Reinforcement learning, a field of machine learning, emphasizes how to act based on the environment to maximize expected benefits and can be used for online control decisions of various devices in complex dynamic systems. Successful applications of reinforcement learning have been reported in other fields, such as game theory, cybernetics, operations research, information theory, simulation optimization, multi-agent system learning, swarm intelligence, statistics, and genetic algorithms.
[0108] Deep reinforcement learning is a branch of artificial intelligence that aims to maximize the accumulated reward value of a trained agent through continuous interaction with its environment to achieve a predetermined control objective. In this process, given a set of states and actions, the transition probability of obtaining a new state s′ from state s given action a is denoted as p(s,a,s′), and the reward value r obtained from action a and state s is denoted as r(s,a). Together with the discount factor γ for future reward values, they are represented as a tuple <s,a,p,r>. The goal of training the agent is to find the optimal policy π(a|s), which maps behavior to a given state to maximize the long-term accumulation of reward value.
[0109]
[0110] V π (s)=E(G t |s t =s;π)
[0111] Q π (s,a)=E(G t |s t =s,a t =a;π)
[0112] Among them, G t It's about accumulating reward values, where t controls the number of iterations, and r... t Let V be the reward value at step t, γ be the discount factor, and E be the expected value function. Two important concepts in deep reinforcement learning algorithms are the state-value function V. π (s) and Q-value function Q π (s,a). V π (s) Evaluates the quality of a state by calculating the expected reward value of following an action starting from a given state; while Q π (s,a) evaluates the reward value of a state by calculating the expected reward value of following action a from state 0.
[0113] In a reinforcement learning agent's neural network, the weights and biases are denoted by θ. The policy can be represented as πθ(a|s). The agent learns to update θ to obtain the maximum reward value.
[0114] This invention employs the maximum entropy safe reinforcement learning algorithm (safe SAC, Soft Actor Critic).
[0115] The process of updating the optimal policy in the traditional SAC algorithm is as follows:
[0116]
[0117] The principle of maximum entropy reinforcement learning encourages the policy to explore more possibilities, increasing the randomness of the policy. Where H(π·|s t ) represents the control policy in state s t The entropy value at time t is H(π) = -Σπ(a|s)logπ(a|s); the α coefficient controls the balance between exploring new control strategies and adopting existing control strategies, R(s). t ,a t ) represents the state s at time t. t The following control action a t The subsequent reward value, In probability distribution ρ π State s at time t in space t The following control action a t The expected value of the reward.
[0118] Based on traditional maximum entropy reinforcement learning, this invention introduces a safety constraint g(s)≤0 into safe reinforcement learning, where g(s) is a function that measures the safety level of state s.
[0119] The objective function of maximum entropy secure reinforcement learning is expressed as:
[0120] max_πJ(π)+αH(π)stE[g(s)]≤0
[0121] In the formula, π is the strategy, J(π) is the expected return obtained when applying strategy π, α is the coefficient of the entropy regularization term, H(π) is the entropy of strategy π, which is used to measure the randomness or uncertainty of the strategy, and E is the mathematical expectation function.
[0122] Transform the constrained optimization problem into Lagrangian form:
[0123] L(π,λ)=J(π)+αH(π)-λE[g(s)]
[0124] In the formula, λ is the Lagrange multiplier;
[0125] Using the policy gradient method, the policy update rule is obtained:
[0126]
[0127] In the formula, L is the objective function after transforming the constrained optimization problem into Lagrange form, logπ represents the logarithmic form of policy π, θ(a|s) represents the probability of policy choosing action a in state s, and Q^π(s,a) represents the expected reward obtained after taking action a from state s under policy π.
[0128] Maximum entropy-safe reinforcement learning comprehensively considers reward maximization, entropy maximization, and safety constraints. By iteratively optimizing this objective, a policy can be obtained that achieves high rewards, is exploratory, and remains safe.
[0129] The present invention also provides a system for enhancing the transmission capacity of an AC / DC hybrid power transmission corridor, comprising:
[0130] The selection module is used to select key transmission corridors for AC / DC hybrid systems;
[0131] Controllable modules are used to identify controllable devices for power transmission in critical transmission corridors of AC / DC hybrid systems.
[0132] The analysis module is used to perform sensitivity analysis on controllable devices and screen out high-sensitivity control devices.
[0133] The decision module is used to build Markov decision process models;
[0134] The solution module is used to solve Markov decision process models based on the maximum entropy secure deep reinforcement learning method to obtain the optimal control strategy.
[0135] Based on the dispatch and control system, this invention optimizes power flow to enhance the transmission capacity of AC / DC hybrid transmission corridors. By utilizing the flexible and rapid power control characteristics of DC transmission systems, it improves the power transmission capacity of regional power grids and the capacity for renewable energy absorption.
[0136] Example: Taking the Jiangsu power grid's cross-river transmission corridor as an example, this corridor includes the Yangzhen DC (±200kV) and 11 AC transmission lines (9 500kV and 2 1000kV). Time-series power flow files were collected to train a maximum entropy security reinforcement learning agent for improving the transmission capacity of the Jiangsu cross-river corridor.
[0137] Among them, the reward value during the training process of the agent is as follows: Figure 3 As shown, with the increase in the number of training iterations, the agent's reward value accumulates continuously. After more than 2000 training iterations, the agent's reward value stabilizes at around 1.4, proving the effectiveness of the agent training. Compared with traditional methods based on a complete system model and optimization solution, the method proposed in this invention can instantly provide an optimization strategy, demonstrating its superiority.
[0138] like Figure 4 and Figure 5 As shown, the strategies provided by the intelligent agent can effectively improve the transmission power and power transmission capacity of the Jiangsu power grid's cross-river corridors, with the total transmission power increasing by an average of 473.73MW and the corridor transmission capacity increasing by an average of 646.06MW.
[0139] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for improving the transmission capacity of an AC / DC hybrid power transmission corridor, characterized in that, The method comprises the following steps: Step 1): selecting a key AC-DC hybrid power transmission corridor; Step 2): determining controllable devices for power transmission of the key AC-DC hybrid power transmission corridor; Step 3): performing sensitivity analysis on the controllable devices to screen out high-sensitivity control devices; Step 4): based on steps 2) and 3), constructing a Markov decision process model; Step 5): based on a maximum entropy safety depth reinforcement learning method, solving the Markov decision process model to obtain an optimal control strategy; In step 2), the controllable devices include: a direct current line and a controllable conventional generator set, wherein the active power output of the controllable conventional generator set is P gen (k); Select all controllable DC lines as control measures to construct the DC line control set P DC .
2. The method of claim 1, wherein In step 1), a key power transmission corridor is selected, and the AC line L AC (i,j) and the DC line L DC (m,n) and the AC line capacity P AC max(i,j), the DC line capacity P DC max(m,n), where i and j are the first and last busbars connected by the AC line, respectively, and m and n are the first and last busbars connected by the DC line, respectively.
3. The method of claim 1, wherein In step 3), historical tidal data files are collected, and sensitivity calculation is performed on the controllable conventional generator set, including: 3.1) For the historical flow data files, the total power, P is derived for the selected key transmission corridors using the AC flow calculation method 走廊 = P 交流 + P 直流 , where P 交流 is the sum of the power of all AC lines in the power transmission corridor, P 交流 =∑ i,j P AC (i,j); P AC (i,j) represents the active power of the AC line; P 直流 is the sum of the power of all DC lines of the power transmission corridor, P 直流 =∑ m,n P DC (m,n);P DC (m,n) represents the active power of the DC line; 3.2) For the kth controllable conventional generator unit, increase the active power output ΔP based on the current active power output gen (k), re-solve the AC power flow equations to calculate the total power flow in the transmission corridor after applying ΔP gen (k) The sensitivity of the active power of the kth controllable conventional generator unit with respect to the power flow in the transmission corridor is given by: 3.3) According to step 3.2), a sensitivity list is obtained, a threshold Sensitivity_Tgen is set, and generators below the threshold are removed, thereby obtaining a list Ggen of controllable generators with sensitivity higher than Sensitivity_Tgen.
4. The method of claim 3, wherein In step 4), the Markov decision process model comprises: State space is the system information one-dimensional vector S = [V m ,V a ,P ac ,Q ac ,P dc ,P gen ,P renewable ] obtained after power flow solution of power grid operation mode, wherein, V m is AC bus voltage amplitude vector, V a is AC bus voltage phase angle vector, P ac is AC line active power vector, Q ac is AC line reactive power vector, P dc is DC line active power vector, P gen is conventional generator set active power vector, P renewable is new energy generator set active power vector; The action space of the multi-type agent is: Type 1: The conventional generator set agent is a conventional generator set active power change vector ΔP gen = [ΔP gen(1) , ΔP gen(2) , …, ΔP gen(g) ], g ∈ Ggen; Type 2: The DC line agent is a DC line active power variation vector ΔP dc = [ΔP dc(1) , ΔP dc(2) , …, ΔP dc(h) ], h ∈ P DC ; The reward function is: Reward=β1r1+β2r2-β3r3-β4r4-β5r5 Wherein: r1 represents the power value of the power transmission corridor transmission capacity improvement; r2 represents the increase in the total DC line power in the power transmission corridor; r3 represents the total AC line power over-limit in the power system; r4 represents the total transformer power over-limit in the power system; r5 represents the AC power flow calculation divergence penalty; β1, β2, β3, β4, β5 are weight coefficients for balancing different types of reward values.
5. The method of claim 4, wherein In step 5), for the DC line agent and the conventional generator set agent, the reinforcement learning agent model is trained according to the state space, action space and reward function in step 4) to obtain the optimal control strategy.
6. The method of claim 5, wherein The training process of the reinforcement learning agent model comprises: First, collect historical operating condition tidal files and perform tidal calculation; Second, extract the state space, and the agent outputs control actions to maximize the reward function value according to the state vector; Then, apply the control actions to the current system state, perform AC tidal calculation and calculate the agent reward function, and store the samples in the form of tuples in the cache; Finally, determine the termination condition of the agent training, if the termination condition is met, exit the training process and save the agent; if the termination condition is not met, extract the next tidal file and continue training the agent.
7. The AC / DC hybrid power transmission corridor power transmission capacity improvement method according to claim 6, characterized in that, the weight and bias parameters in the reinforcement learning agent model are denoted by θ, and the agent updates θ by learning to obtain the maximum reward value; when solving, a maximum entropy safety reinforcement learning algorithm is used, and on the basis of the maximum entropy reinforcement learning, the safety reinforcement learning introduces a safety constraint g(s)≤0, wherein g(s) is a function for measuring the safety degree of state s; the objective function of the maximum entropy safety reinforcement learning is expressed as: max_πJ(π)+αH(π)s.t.E[g(s)]≤0 wherein π is a policy, J(π) is the expected return obtained by applying the policy π, α is the coefficient of the entropy regularization term, H(π) is the entropy of the policy π, and E is a mathematical expectation function; the constraint optimization problem is converted into a Lagrange form: L(π,λ)=J(π)+αH(π)-λE[g(s)] wherein λ is a Lagrange multiplier; using a policy gradient method, the policy update rule is obtained: wherein L is the objective function after the constraint optimization problem is converted into the Lagrange form, logπ represents the logarithmic form of the policy π, θ(a|s) represents the probability of the policy selecting an action a under a state s, and Q^π(s,a) represents the expected return obtained by taking the action a from the state s under the policy π.
8. The AC / DC hybrid power transmission corridor power transmission capacity improvement method according to claim 1, characterized in that, it further comprises step 6), which is specifically: in the process of real-time dispatching and operation of the power grid, system data under the current power grid operation mode are extracted, power flow is solved, and an agent input vector is constructed, an optimization decision value of the DC line and the conventional generator set is given by calling the reinforcement learning agent model, power flow calculation is performed to verify the effectiveness of the agent decision result; if the control strategy is checked, it is issued to the DC field station and the generator set; if the control strategy is not checked, it is not issued to the field station or the generator set.
9. A system for enhancing the transmission capacity of an AC / DC hybrid power transmission corridor, characterized in that, it comprises: a selection module for selecting a key power transmission corridor of an AC / DC hybrid system; a controllable module for determining controllable equipment for power transmission of the key power transmission corridor of the AC / DC hybrid system; an analysis module for performing sensitivity analysis on the controllable equipment and screening out high-sensitivity control equipment; a decision module for constructing a Markov decision process model; a solving module for solving the Markov decision process model based on a maximum entropy safety deep reinforcement learning method to obtain an optimal control strategy. The controllable device comprises a direct current line and a controllable conventional generator set, wherein the active power output of the controllable conventional generator set is P gen (k); All controllable DC lines are selected as control measures to construct the DC line control set P DC .
Citation Information
Patent Citations
Interval optimal power flow method for AC-DC hybrid power transmission system based on confidence transformation
CN107221935A
Calculation method and system for security constrained unit commitment of alternating current and direct current hybrid power grid
CN109904856A