A power control method and system for damping low frequency oscillations in an ac power grid

By connecting a DC branch in the low-frequency oscillation region of the AC power grid and optimizing the unit output and DC branch transmission power, and by using reinforcement learning algorithms to improve the damping ratio, the problem of low-frequency oscillation of the AC power grid was solved, achieving higher stability and lower losses.

CN112186781BActive Publication Date: 2026-01-16CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011036575.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-23
Publication Date
2026-01-16
Estimated Expiration
2040-09-23

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively suppress low-frequency oscillations in AC power grids, especially in large-scale grid oscillation modes, where traditional methods such as power system stabilizers (PSS) are difficult to implement online adjustments.

Method used

By connecting a DC branch in the low-frequency oscillation region of the AC power grid and optimizing the unit output and DC branch transmission power through reinforcement learning algorithms, the damping ratio of the AC power grid is improved and the oscillation degree is reduced.

Benefits of technology

By optimizing the unit output and DC branch transmission power, the operational stability of the AC power grid is significantly improved, power transmission losses are reduced, and short-circuit current of the AC power grid is limited.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112186781B_ABST
    Figure CN112186781B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of power control method and system for inhibiting low-frequency oscillation of alternating current network, when low-frequency oscillation occurs in alternating current network, direct current branch is connected between any two busbars in low-frequency oscillation region of alternating current network;The initial output of unit and the initial transmission power of direct current branch of alternating current network are obtained;The initial output of unit and the initial transmission power of direct current branch are substituted as initial state information into power optimization model, and the optimal unit output and the optimal transmission power of direct current branch of alternating current network are obtained by optimizing and solving model.The present application improves the damping ratio when low-frequency oscillation occurs in alternating current network by connecting direct current branch in alternating current network and optimizing the unit output and the transmission power of direct current branch in alternating current network when low-frequency oscillation occurs in alternating current network, and then reduces the oscillation degree of alternating current network, improves the operating stability of alternating current network.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of power system safety and stability analysis and control, in particular to a power control method and system for suppressing low-frequency oscillation of an alternating current power grid. BACKGROUND

[0002] Regional interconnection has become an inevitable trend of power grid development. The adjustment of power output and load in the power system is conducive to coping with unexpected accidents and ensuring the stable operation of the power grid. Low-frequency oscillation problems rarely occur in the early stage of power grid establishment, but with the development of industrialization and urbanization, the demand for electricity is increasing, and the entire power grid is often passively in a limit operation state, and low-frequency power fluctuation problems begin to plague the power industry. The operation mode of the power system is variable, and various factors such as the increase of disturbance sources, the increase of unit capacity, and long-distance power transmission are threatening the safe and stable operation of the power grid.

[0003] Generally, low-frequency oscillation of the power system refers to the relative swing between the rotors of the generator set when the generator set is running in parallel through the transmission line under disturbance, and it will cause continuous oscillation in the absence of damping. At this time, the power on the transmission line will also oscillate accordingly. Since the oscillation frequency is low, about 0.2HZ to 2.5HZ, it is called low-frequency oscillation.

[0004] Low-frequency oscillation of the power system is generally divided into two categories: one is local oscillation, and the other is regional oscillation. Local oscillation refers to the oscillation between the units in the station or between the units in the stations with close electrical distance. This oscillation is limited within the region, has a small influence range and is easy to eliminate. The oscillation frequency is about 0.7HZ to 2.5HZ. Regional oscillation refers to the oscillation of a part of the generator group relative to another part of the generator group. In a weakly interconnected system, low-frequency oscillation often occurs between two or more coupled generator groups. Due to the large electrical distance and the large inertia time constant of the equivalent generator group of the generator group, the oscillation frequency is about 0.2HZ to 0.7HZ. Here, we mainly aim at regional low-frequency oscillation.

[0005] Because the oscillation damping ratio of the region is inversely proportional to the oscillation degree, the oscillation damping ratio of the region can be improved to suppress the low-frequency oscillation of the power system. The oscillation mode of the region where the low-frequency oscillation occurs can be many, and the oscillation damping ratio of each oscillation mode can be different. We believe that improving the weakest oscillation damping ratio corresponding to various oscillation modes can achieve the effect of suppressing the low-frequency oscillation of the power system.

[0006] In the prior art, a power system stabilizer (PSS) is used to improve the oscillation mode damping ratio of the large power grid, but this method has limitations and is difficult to realize online adjustment.

[0007] At present, no other technology except the above-mentioned technology has been proposed. SUMMARY

[0008] In view of the deficiencies of the prior art, the purpose of the present application is to provide a power control method and system for suppressing low-frequency oscillation of an alternating current power grid, which improves the damping ratio when the alternating current power grid is in low-frequency oscillation by connecting a direct current branch in the alternating current power grid and optimizing the unit output and the direct current branch power transmission in the alternating current power grid, thereby reducing the oscillation degree of the alternating current power grid and improving the operation stability of the alternating current power grid.

[0009] The purpose of the present application is achieved by using the following technical solutions:

[0010] The present application provides a power optimization method for suppressing low-frequency oscillation of an alternating current power grid, which improves in that the method comprises:

[0011] When the alternating current power grid is in low-frequency oscillation, a direct current branch is connected between any two buses in the low-frequency oscillation region of the alternating current power grid;

[0012] The initial unit output and the initial direct current branch power transmission of the alternating current power grid are obtained;

[0013] The initial unit output and the initial direct current branch power transmission are substituted into a power optimization model as initial state information, and the model is optimized and solved to obtain the optimal unit output of the alternating current power grid and the optimal direct current branch power transmission.

[0014] The power optimization model is an initial state set and a preset action set based on the historical unit output and the direct current branch power transmission of the alternating current power grid with the connected direct current branch, and is obtained by training a reinforcement learning network using a reinforcement learning algorithm.

[0015] Preferably, the preset action set comprises actions corresponding to each unit in the alternating current power grid with the connected direct current branch and actions corresponding to each direct current branch in the alternating current power grid with the connected direct current branch.

[0016] The action corresponding to the s-th unit in the alternating current power grid with the connected direct current branch is an increase in output ΔP s and a decrease in output ΔP s .

[0017] The action corresponding to the z-th direct current branch in the alternating current power grid with the connected direct current branch is an increase in power transmission ΔP z and a decrease in power transmission ΔP z .

[0018] Wherein, s∈(1~ξ s ), ξ sthe total number of units included in the AC power grid connected with the DC branch, z e (1~ξ z ), ξ z the total number of DC branches included in the AC power grid connected with the DC branch, ΔP s a preset output change value of the s-th unit included in the AC power grid connected with the DC branch, ΔP2 a preset transmission power change value of the z-th DC branch included in the AC power grid connected with the DC branch.

[0019] Preferably, the power optimization model is obtained by training a reinforcement learning network based on an initial state set and a preset action set constructed based on historical unit output and DC branch transmission power of the AC power grid connected with the DC branch, and includes:

[0020] taking the i-th state information in the initial state set as initial state information of the AC power grid connected with the DC branch in the reinforcement learning network;

[0021] substituting state information and an action set at the beginning of each iteration process corresponding to the initial state information into a pre-constructed action selection model to obtain an optimal action corresponding to the state information at the beginning of each iteration process corresponding to the initial state information, and a weakest oscillation damping ratio corresponding to the optimal action;

[0022] based on the weakest oscillation damping ratio corresponding to the optimal action, calculating a Q value of each iteration process corresponding to the initial state information by using a Q learning algorithm;

[0023] updating the state information based on the optimal action, and taking the updated state information as state information at the beginning of the next iteration process;

[0024] when a preset iteration number threshold is reached, taking unit output and DC line transmission power in the updated state information in the iteration process with the largest Q value as output unit optimization output and DC line optimization transmission power corresponding to the initial state information in the reinforcement learning network taking the i-th state information in the initial state set as the initial state information;

[0025] repeating all the above steps until unit optimization output and DC line optimization transmission power corresponding to initial state information in the reinforcement learning network taking each state information in the initial state set as the initial state information are obtained, taking the trained reinforcement learning network at this time as the power optimization model, and outputting the power optimization model;

[0026] wherein, i e (1~N), N is the number of state information included in the initial state set.

[0027] Further, the obtaining process of the pre-constructed action selection model includes:

[0028] The initial state set of each piece of state information and the preset action set are taken as input data of the initial neural network, the optimal action corresponding to each piece of state information in the initial state set and the weakest oscillation damping ratio corresponding to the optimal action are taken as output data, the initial neural network is trained, and a pre-constructed action selection model is obtained.

[0029] Further, the process of obtaining the optimal action corresponding to each piece of state information in the initial state set and the weakest oscillation damping ratio corresponding to the optimal action comprises:

[0030] Step A: initialization f = 1;

[0031] Step B: based on the i-th piece of state information in the initial state set, the f-th action of the preset action set is applied to obtain state information after the action is executed;

[0032] Step C: the disturbance response information corresponding to the unit output and the DC branch transmission power in the state information is obtained, the weakest oscillation damping ratio corresponding to the low-frequency oscillation region of the AC power grid in various possible oscillation modes is obtained from the disturbance response information, and the weakest oscillation damping ratio is taken as the weakest oscillation damping ratio corresponding to the f-th action;

[0033] Step D: if f = δ, the action corresponding to the maximum one of the weakest oscillation damping ratios corresponding to the first action to the δ-th action is taken as the optimal action corresponding to the i-th piece of state information in the initial state set, and the optimal action and the weakest oscillation damping ratio corresponding to the optimal action are output;

[0034] Wherein, f ∈ (1 ~ δ), δ is the number of actions contained in the preset action set, i ∈ (1 ~ N), N is the number of pieces of state information contained in the initial state set.

[0035] The application also provides a power control method for suppressing low-frequency oscillation of an AC power grid, and the improvement thereof lies in that the method comprises:

[0036] The optimal unit output of the AC power grid and the optimal transmission power of the DC branch are obtained by using the above power optimization step;

[0037] The unit output of the AC power grid connected with the DC branch is controlled to be the optimal output, and the transmission power of the DC branch of the AC power grid connected with the DC branch is controlled to be the optimal transmission power.

[0038] The application provides a power optimization system for suppressing low-frequency oscillation of an AC power grid, and the improvement thereof lies in that the system comprises:

[0039] The access module is configured to access a DC branch between any two buses in a low-frequency oscillation region of the AC power grid when the low-frequency oscillation occurs in the AC power grid.

[0040] The acquisition module is configured to acquire initial unit output and initial transmission power of the DC branch of the AC power grid.

[0041] The optimization solving module is configured to substitute the initial unit output and the initial transmission power of the DC branch as initial state information into a power optimization model, and perform optimization solving on the model to obtain optimal unit output and optimal transmission power of the DC branch of the AC power grid.

[0042] The power optimization model is obtained by training a reinforcement learning network based on an initial state set and a preset action set constructed based on historical unit output and transmission power of the DC branch of the AC power grid to which the DC branch is accessed.

[0043] Compared with the closest prior art, the present application has the following beneficial effects:

[0044] The technical scheme provided by the present application has the following beneficial effects: when the low-frequency oscillation occurs in the AC power grid, the DC branch is accessed between any two buses in the low-frequency oscillation region of the AC power grid; the initial unit output and the initial transmission power of the DC branch of the AC power grid are acquired; the initial unit output and the initial transmission power of the DC branch are substituted as initial state information into a power optimization model, and the model is optimized and solved to obtain optimal unit output and optimal transmission power of the DC branch of the AC power grid; the power optimization model is obtained by training a reinforcement learning network based on an initial state set and a preset action set constructed based on historical unit output and transmission power of the DC branch of the AC power grid to which the DC branch is accessed. In this way, when the low-frequency oscillation occurs in the AC power grid, the damping ratio when the low-frequency oscillation occurs in the AC power grid is improved by accessing the DC branch in the AC power grid and optimizing the unit output and the transmission power of the DC branch in the AC power grid, thereby reducing the oscillation degree of the AC power grid and improving the operation stability of the AC power grid.

[0045] The technical scheme provided by the present application has the following beneficial effects: when the low-frequency oscillation occurs in the AC power grid, the DC branch is accessed between any two buses in the low-frequency oscillation region of the AC power grid; the initial unit output and the initial transmission power of the DC branch of the AC power grid are acquired; the initial unit output and the initial transmission power of the DC branch are substituted as initial state information into a power optimization model, and the model is optimized and solved to obtain optimal unit output and optimal transmission power of the DC branch of the AC power grid; the power optimization model is obtained by training a reinforcement learning network based on an initial state set and a preset action set constructed based on historical unit output and transmission power of the DC branch of the AC power grid to which the DC branch is accessed. In this way, when the low-frequency oscillation occurs in the AC power grid, the damping ratio when the low-frequency oscillation occurs in the AC power grid is improved by accessing the DC branch in the AC power grid and optimizing the unit output and the transmission power of the DC branch in the AC power grid, thereby reducing the oscillation degree of the AC power grid and improving the operation stability of the AC power grid.

[0046] The technical scheme provided by the present application has the following beneficial effects: when the low-frequency oscillation occurs in the AC power grid, the DC branch is accessed between any two buses in the low-frequency oscillation region of the AC power grid; the initial unit output and the initial transmission power of the DC branch of the AC power grid are acquired; the initial unit output and the initial transmission power of the DC branch are substituted as initial state information into a power optimization model, and the model is optimized and solved to obtain optimal unit output and optimal transmission power of the DC branch of the AC power grid; the power optimization model is obtained by training a reinforcement learning network based on an initial state set and a preset action set constructed based on historical unit output and transmission power of the DC branch of the AC power grid to which the DC branch is accessed. In this way, when the low-frequency oscillation occurs in the AC power grid, the damping ratio when the low-frequency oscillation occurs in the AC power grid is improved by accessing the DC branch in the AC power grid and optimizing the unit output and the transmission power of the DC branch in the AC power grid, thereby reducing the oscillation degree of the AC power grid and improving the operation stability of the AC power grid.

[0047] The technical scheme provided by the present application has the following beneficial effects: when the low-frequency oscillation occurs in the AC power grid, the DC branch is accessed between any two buses in the low-frequency oscillation region of the AC power grid; the initial unit output and the initial transmission power of the DC branch of the AC power grid are acquired; the initial unit output and the initial transmission power of the DC branch are substituted as initial state information into a power optimization model, and the model is optimized and solved to obtain optimal unit output and optimal transmission power of the DC branch of the AC power grid; the power optimization model is obtained by training a reinforcement learning network based on an initial state set and a preset action set constructed based on historical unit output and transmission power of the DC branch of the AC power grid to which the DC branch is accessed. In this way, when the low-frequency oscillation occurs in the AC power grid, the damping ratio when the low-frequency oscillation occurs in the AC power grid is improved by accessing the DC branch in the AC power grid and optimizing the unit output and the transmission power of the DC branch in the AC power grid, thereby reducing the oscillation degree of the AC power grid and improving the operation stability of the AC power grid. BRIEF DESCRIPTION OF DRAWINGS

[0048] Figure 1It is a power optimization method for inhibiting low-frequency oscillation of an alternating current power grid.

[0049] Figure 2 It is a power optimization system structure diagram for inhibiting low-frequency oscillation of an alternating current power grid.

[0050] Figure 3 It is a 4-machine 2-area alternating current power grid structure diagram in which a direct current branch is accessed in embodiment 3 of the present application.

[0051] Figure 4 It is a 4-machine 2-area alternating current power grid structure diagram in which a direct current branch and a generator GEN5 are accessed in embodiment 4 of the present application. DETAILED DESCRIPTION

[0052] The specific embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0053] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of the present application.

[0054] Embodiment 1

[0055] The present application provides a power optimization method for inhibiting low-frequency oscillation of an alternating current power grid, as shown in the figure, the method comprises the following steps: Figure 1

[0056] Step 101, when low-frequency oscillation occurs in the alternating current power grid, a direct current branch is accessed between any two buses in the low-frequency oscillation area of the alternating current power grid;

[0057] Step 102, obtaining initial unit output and initial transmission power of the direct current branch of the alternating current power grid;

[0058] Step 103, substituting the initial unit output and the initial transmission power of the direct current branch as initial state information into a power optimization model, and optimizing and solving the model to obtain optimal unit output of the alternating current power grid and optimal transmission power of the direct current branch;

[0059] The power optimization model is an initial state set and a preset action set based on historical unit output and transmission power of the direct current branch of the alternating current power grid in which the direct current branch is accessed, and is obtained by training a reinforcement learning network using a Q-learning reinforcement learning algorithm.

[0060] The direct current branch is connected with a rectifier and an inverter. ​

[0061] In the embodiments of the present application, the low-frequency oscillation region of the alternating current power grid in step 101 is a region surrounded by the units with the oscillation damping ratio greater than 1 and the oscillation frequency greater than 0 and less than 1 in the alternating current power grid, and the type of low-frequency oscillation occurring in the region is regional low-frequency oscillation.

[0062] In the embodiments of the present application, any two buses in step 101 can be preferentially selected according to the following rules:

[0063] 1) The two buses connecting different field groups in the low-frequency oscillation region of the alternating current power grid;

[0064] 2) The two buses with the farthest electrical distance in the same field group in the low-frequency oscillation region of the alternating current power grid.

[0065] In the embodiments of the present application, in addition to using the Q-learning reinforcement learning algorithm to train the reinforcement learning network, the DQN, Policy Gradient and other reinforcement learning algorithms can also be used to train the reinforcement learning network.

[0066] In the embodiments of the present application, each alternating current power grid low-frequency oscillation region prone to low-frequency oscillation in the alternating current power grid can be processed as follows:

[0067] According to the low-frequency oscillation working condition that can occur, the bus position of each working condition for accessing the DC branch is set in advance, and the alternating current power grid structure after accessing the DC branch corresponding to each working condition is obtained;

[0068] The power optimization model corresponding to the alternating current power grid structure after accessing the DC branch corresponding to each working condition is trained respectively;

[0069] When the alternating current power grid occurs low-frequency oscillation, the power optimization model is selected according to the alternating current power grid low-frequency oscillation region and the low-frequency oscillation working condition occurring in the region;

[0070] According to the alternating current power grid structure corresponding to the power optimization model, the real alternating current power grid structure is reconstructed (i.e. accessing the DC branch);

[0071] The unit output and DC branch transmission power of the reconstructed alternating current power grid are optimized by using the selected power optimization model;

[0072] The DC branch is connected with a circuit breaker, and the connection and disconnection of the circuit breaker can be flexibly controlled.

[0073] Specifically, the process of obtaining the initial state set constructed by the unit output and DC branch transmission power of the alternating current power grid with the accessed DC branch includes:

[0074] The disturbance response information corresponding to each historical unit output and DC branch transmission power of the alternating current power grid with the accessed DC branch during normal operation is obtained.

[0075] When the disturbance response information corresponding to the historical unit output and DC branch transmission power of each history exists in the low-frequency oscillation region of the alternating current grid, the unit output and DC branch transmission power of the corresponding history are filled in the initial state set as a state information.

[0076] Further, the disturbance response information corresponding to the historical unit output and DC branch transmission power of each history when the alternating current grid accessing the DC branch is in normal operation is obtained, comprising:

[0077] A power flow file is constructed by using the network parameters of the alternating current grid accessing the DC branch, the Kth set of unit output and DC branch transmission power;

[0078] The power flow file is imported into the BPA simulation platform for power flow calculation, and the power flow calculation result is obtained;

[0079] The power flow calculation result is imported into the BPA simulation platform for small disturbance calculation, and the disturbance response information corresponding to the Kth set of unit output and DC branch transmission power is obtained;

[0080] The disturbance response information comprises a low-frequency oscillation region of the alternating current grid and a weakest oscillation damping ratio in various oscillation modes in the region, k ∈ (1 ~ M), and M is the number of historical unit output and DC branch transmission power when the alternating current grid accessing the DC branch is in normal operation.

[0081] In the embodiment of the application, the power flow calculation and small disturbance calculation are performed on the initial generator output and DC branch transmission power by using the power system simulation software BPA, and the damping ratio calculation formula The damping ratio of the low-frequency oscillation region of the alternating current grid in various oscillation modes is obtained, and the weakest damping ratio in all oscillation modes can be selected from the damping ratio; that is, each initial generator output and DC branch transmission power corresponds to a weakest damping ratio.

[0082] Further, the preset action set comprises an action corresponding to each unit in the alternating current grid accessing the DC branch and an action corresponding to each DC branch in the alternating current grid accessing the DC branch.

[0083] The action corresponding to the s th unit in the alternating current grid accessing the DC branch is an output increase ΔP s and an output reduction ΔP s .

[0084] The action corresponding to the z th DC branch in the alternating current grid accessing the DC branch is a transmission power increase ΔP z and a transmission power reduction ΔP z .

[0085] s ∈ (1 ~ ξs ), ξ s is the total number of units contained in the AC power grid connected with the DC branch, z e (1 ~ ξ z ), ξ z is the total number of DC branches contained in the AC power grid connected with the DC branch, ΔP s is the preset output change value of the s-th unit contained in the AC power grid connected with the DC branch, and ΔP2 is the preset transmission power change value of the z-th DC branch contained in the AC power grid connected with the DC branch.

[0086] Specifically, the power optimization model is obtained by training a reinforcement learning network based on an initial state set and a preset action set constructed based on historical unit output and DC branch transmission power of the AC power grid connected with the DC branch, and includes:

[0087] taking the i-th state information in the initial state set as initial state information of the AC power grid connected with the DC branch in the reinforcement learning network;

[0088] substituting state information and an action set at the beginning of each iteration process corresponding to the initial state information into a pre-constructed action selection model to obtain an optimal action corresponding to the state information at the beginning of each iteration process corresponding to the initial state information, and a weakest oscillation damping ratio corresponding to the optimal action;

[0089] based on the weakest oscillation damping ratio corresponding to the optimal action, calculating a Q value of each iteration process corresponding to the initial state information by using a Q learning algorithm;

[0090] updating the state information based on the optimal action, and taking the updated state information as state information at the beginning of the next iteration process;

[0091] when a preset iteration number threshold is reached, taking the unit output and the DC line transmission power in the updated state information in the iteration process with the largest Q value as the output unit optimization output and the DC line optimization transmission power corresponding to the initial state information of the AC power grid connected with the DC branch in the reinforcement learning network;

[0092] repeating all the above steps until the unit optimization output and the DC line optimization transmission power corresponding to each state information in the initial state set as the initial state information of the AC power grid connected with the DC branch in the reinforcement learning network are obtained, taking the trained reinforcement learning network at this time as the power optimization model, and outputting the power optimization model;

[0093] wherein, i e (1 ~ N), N is the number of state information contained in the initial state set.

[0094] In the embodiment of the application, the initial generator output and the DC branch transmission power are taken as the initial point of the deep Q network optimization calculation, the weakest damping ratio corresponding to the initial generator output and the DC branch transmission power is taken as the optimization target, the difference between the weakest damping ratio and the target damping ratio is taken as the reward value, the Q value is updated by the reward value, and the initial generator output and the DC branch transmission power are adjusted in each iteration process of the deep Q network (the weakest damping ratio corresponding to each adjusted generator output and DC branch transmission power is changed), so that the optimal generator output and DC branch transmission power corresponding to the optimal weakest damping ratio are found.

[0095] That is, the initial generator output and the DC branch transmission power are taken as the state input of the agent, the change of the active power each time is taken as the action of the agent, the reward of the agent corresponding to each action is calculated according to the weakest damping ratio and the target damping ratio each time, and the generator output and the DC branch transmission power corresponding to the maximum value of the system weakest damping ratio are obtained through the programmed iterative calculation.

[0096] A function composed of the active power (action) and the target value in each iteration process can be established.

[0097] F = max (ζ min (P1, P2, …, P n , P dc ))

[0098] The optimization of the agent is encouraged according to the reward mechanism of the Q network, and the Q value is updated according to the following formula:

[0099] Q (s t+1 , a t+1 ) = Q (s t , a t ) + α [r t + γmax a Q (s t+1 , a t+1 ) - Q (s t , a t )]

[0100] We hope that the action of each iteration is optimal (that is, the weakest damping ratio corresponding to each action is maximum), and this process needs to obtain the weakest damping ratio corresponding to various actions in the action set, but this process needs to be realized by means of the BPA simulation platform, which is time-consuming and has a large amount of calculation and is difficult to realize, so that the current state and the action set are substituted into the action selection model in each iteration, the optimal action and the weakest damping ratio corresponding to the optimal action are predicted by the model, so as to ensure the training speed.

[0101] The deep Q network is optimized according to a reward mechanism under a certain probability of a greedy rate. The Q network is continuously optimized while the agent is optimized. Finally, the agent finds an optimal solution, and records the unit output and DC transmission power at this time. The optimal solution obtained through calculation is the maximum value of the weakest damping ratio. The stability of the power system after optimization of the unit output and DC branch transmission power is simulated and calculated, and the stability of the power system after optimization is compared with the stability of the power system before optimization. It can be seen that the stability of the power system after optimization is significantly improved.

[0102] Further, the obtaining process of the pre-constructed action selection model comprises:

[0103] The initial state set, the state information and the preset action set are used as input data of the initial neural network, the optimal action corresponding to the state information in the initial state set and the weakest oscillation damping ratio corresponding to the optimal action are used as output data, and the initial neural network is trained to obtain the pre-constructed action selection model.

[0104] In the embodiment of the application, the action selection model is trained by using a deep reinforcement learning network structure comprising an input layer, double hidden layers and an output layer, so that the training accuracy of the action selection model is ensured.

[0105] Further, the obtaining process of the optimal action corresponding to the state information in the initial state set and the weakest oscillation damping ratio corresponding to the optimal action comprises:

[0106] Step A: initializing f = 1;

[0107] Step B: applying the fth action of the preset action set to the ith state information in the initial state set to obtain state information after execution of the action;

[0108] Step C: obtaining disturbance response information corresponding to the unit output and the DC branch transmission power in the state information, obtaining the weakest oscillation damping ratio corresponding to the oscillation damping ratio in the low-frequency oscillation region of the alternating current power grid under various possible oscillation modes from the disturbance response information, and taking the weakest oscillation damping ratio as the weakest oscillation damping ratio corresponding to the fth action;

[0109] Step D: if f = δ, taking the action corresponding to the maximum one of the weakest oscillation damping ratios corresponding to the first action to the δth action as the optimal action corresponding to the ith state information in the initial state set, and outputting the optimal action and the weakest oscillation damping ratio corresponding to the optimal action;

[0110] wherein f is an element of (1-δ), δ is the number of actions included in the preset action set, i is an element of (1-N), and N is the number of state information included in the initial state set.

[0111] Further, the Q-learning algorithm is used to calculate the Q value of each iteration process corresponding to the initial state information, and the calculation comprises the following steps.

[0112] The Q value Q calculated in the t+1th iteration process is determined by the following formula t+1 :

[0113] Q t+1 = Q t + α [r t + γ (Q t+1 - Q t )]

[0114] In the formula, Q t is the Q value calculated in the tth iteration process, r t is the reward value of the optimal action corresponding to the state information at the beginning of the tth iteration process, α is a first preset parameter, γ is a second preset parameter, t ∈ (1 ~ T), and T is a preset iteration number threshold value.

[0115] The reward value r t of the optimal action corresponding to the state information at the beginning of the tth iteration process is determined by the following formula

[0116]

[0117] In the formula, Δη t is the difference between the target weakest oscillation damping ratio and the weakest oscillation damping ratio corresponding to the optimal action corresponding to the state information at the beginning of the tth iteration.

[0118] In the embodiment of the application, i represents the running algebra of the agent, and when the agent is iterated to a certain number of steps or power flow calculation does not converge, the iteration calculation of this generation is stopped, and the iteration calculation of the next generation is started; when the agent optimization reaches 100 generations, the agent stops optimization, and outputs the optimal power system scheduling combination and the maximum value of the weakest damping ratio.

[0119] In the embodiment of the application, when low-frequency oscillation occurs in the alternating current power grid, the direct current branch is selected to be connected in the low-frequency oscillation region of the alternating current power grid, and the low-frequency oscillation of the alternating current power grid is suppressed by optimizing the unit output and the direct current branch power transmission.

[0120] 1) The direct current transmission is more stable than the alternating current transmission, and the connection of the direct current branch can positively affect the stability of the alternating current power grid.

[0121] 2) When the low-frequency oscillation is serious, the direct current branch can be used to transmit direct current power from the outside to maintain the stability of the alternating current power grid.

[0122] Based on the same inventive concept, the application also provides a power control method for inhibiting low-frequency oscillation of an alternating current power grid, the method comprising:

[0123] By the above method, the optimal unit output of the alternating current power grid and the optimal transmission power of the DC branch are obtained.

[0124] The unit output of the alternating current power grid connected with the DC branch is controlled to be the respective optimal output, and the DC branch transmission power of the alternating current power grid connected with the DC branch is controlled to be the respective optimal transmission power.

[0125] Embodiment 2:

[0126] The application also provides a power optimization system for inhibiting low-frequency oscillation of an alternating current power grid, as shown in the accompanying drawings, the system comprising: Figure 2

[0127] An access module is configured to access a DC branch between any two buses in a low-frequency oscillation region of the alternating current power grid when low-frequency oscillation occurs in the alternating current power grid.

[0128] An obtaining module is configured to obtain initial unit output and initial DC branch transmission power of the alternating current power grid.

[0129] An optimization solving module is configured to substitute the initial unit output and the initial DC branch transmission power as initial state information into a power optimization model, and to optimize and solve the model to obtain optimal unit output of the alternating current power grid and optimal DC branch transmission power.

[0130] The power optimization model is obtained by training a reinforcement learning network based on an initial state set and a preset action set constructed by historical unit output and DC branch transmission power of the alternating current power grid connected with the DC branch.

[0131] Specifically, the system further comprises a construction module for pre-construction of the power optimization model, and the construction module comprises:

[0132] A pre-construction unit is configured to construct an initial state set and a preset action set based on historical unit output and DC branch transmission power of the alternating current power grid connected with the DC branch.

[0133] A first training unit is configured to train a reinforcement learning network based on the initial state set and the preset action set by using a reinforcement learning algorithm to obtain the power optimization model.

[0134] Specifically, the pre-construction unit comprises:

[0135] A first obtaining subunit is configured to obtain interference response information corresponding to each historical unit output and DC branch transmission power of the alternating current power grid connected with the DC branch during normal operation. ​

[0136] The filling unit is configured to fill the unit output and the DC branch transmission power of the corresponding history as a piece of state information into the initial state set when the AC power grid low-frequency oscillation region exists in the interference response information corresponding to the unit output and the DC branch transmission power of each history.

[0137] Further, the obtaining subunit comprises:

[0138] The constructing sub-module is configured to construct a power flow file by using the network parameters of the AC power grid connected with the DC branch, the Kth set of unit output and the DC branch transmission power;

[0139] The first obtaining sub-module is configured to import the power flow file into the BPA simulation platform to perform power flow calculation and obtain a power flow calculation result;

[0140] The second obtaining sub-module is configured to import the power flow calculation result into the BPA simulation platform to perform small disturbance calculation and obtain the interference response information corresponding to the Kth set of unit output and the DC branch transmission power;

[0141] The interference response information comprises an AC power grid low-frequency oscillation region and a weakest oscillation damping ratio in various oscillation modes in the region, and k∈(1~M) and M is the number of pieces of unit output and DC branch transmission power in the history of the AC power grid connected with the DC branch in normal operation.

[0142] Further, the preset action set comprises an action corresponding to each unit in the AC power grid connected with the DC branch and an action corresponding to each DC branch in the AC power grid connected with the DC branch;

[0143] The action corresponding to the s th unit in the AC power grid connected with the DC branch is an increase in output ΔP s and a decrease in output ΔP s .

[0144] The action corresponding to the z th DC branch in the AC power grid connected with the DC branch is an increase in transmission power ΔP z and a decrease in transmission power ΔP z .

[0145] The s∈(1~ξ s ), ξ s is the total number of units in the AC power grid connected with the DC branch, z∈(1~ξ z ), ξ z is the total number of DC branches in the AC power grid connected with the DC branch, ΔP s is a preset output change value of the s th unit in the AC power grid connected with the DC branch, and ΔP2 is a preset transmission power change value of the z th DC branch in the AC power grid connected with the DC branch.

[0146] Specifically, the first training unit comprises:

[0147] an input subunit configured to take the i-th state information in the initial state set as initial state information of the AC power grid connected with the DC branch in the reinforcement learning network;

[0148] an optimal action selection subunit configured to input the state information and the action set at the beginning of each iteration process corresponding to the initial state information into a pre-constructed action selection model to obtain an optimal action corresponding to the state information at the beginning of each iteration process corresponding to the initial state information and a weakest oscillation damping ratio corresponding to the optimal action;

[0149] a Q value calculation subunit configured to calculate a Q value of each iteration process corresponding to the initial state information based on the weakest oscillation damping ratio corresponding to the optimal action by using a Q learning algorithm;

[0150] an updating unit configured to update the state information based on the optimal action and take the updated state information as the state information at the beginning of the next iteration process;

[0151] a defining subunit configured to take the unit output and the DC line transmission power in the updated state information in the iteration process with the maximum Q value as the unit optimal output and the DC line optimal transmission power corresponding to the initial state information in the initial state set when a preset iteration number threshold is reached, the initial state information being taken as initial state information of the AC power grid connected with the DC branch in the reinforcement learning network;

[0152] a first output subunit configured to repeat all the above steps until the unit optimal output and the DC line optimal transmission power corresponding to each state information in the initial state set are obtained, the trained reinforcement learning network at this time being taken as a power optimization model, and the power optimization model being output;

[0153] wherein i is an integer in (1, N), and N is the number of state information included in the initial state set.

[0154] Specifically, the system further comprises a second construction module configured to pre-construct an action selection model, and the second construction module is configured to:

[0155] an obtaining unit configured to obtain the optimal action corresponding to each state information in the initial state set and the weakest oscillation damping ratio corresponding to the optimal action;

[0156] The second training unit is used to take each state information in the initial state set and the preset action set as input data of the initial neural network, and the optimal action corresponding to each state information in the initial state set and the weakest oscillation damping ratio corresponding to the optimal action as output data to train the initial neural network and obtain the pre-constructed action selection model.

[0157] Furthermore, the acquisition unit includes:

[0158] Initialize the subunit to initialize f = 1;

[0159] An action execution unit is used to apply the f-th action of a preset action set based on the i-th state information in the initial state set, and to obtain the state information after the action is executed.

[0160] The second acquisition subunit is used to acquire the interference response information corresponding to the unit output and DC branch transmission power in the status information, acquire the weakest oscillation damping ratio among the oscillation damping ratios of the AC power grid low-frequency oscillation region in its various possible oscillation modes from the interference response information, and take the weakest oscillation damping ratio as the weakest oscillation damping ratio corresponding to the f-th action.

[0161] The second output subunit is used to, if f = δ, take the action corresponding to the weakest oscillation damping ratio with the largest among the weakest oscillation damping ratios corresponding to the first action to the δth action as the optimal action corresponding to the i-th state information in the initial state set, and output the optimal action and the weakest oscillation damping ratio corresponding to the optimal action.

[0162] Where f∈(1~δ), δ is the number of actions contained in the preset action set, i∈(1~N), and N is the number of state information entries contained in the initial state set.

[0163] Furthermore, the Q-value calculation subunit is used for:

[0164] The Q value Q calculated during the (t+1)th iteration is determined by the following formula. t+1 :

[0165] Q t+1 =Q t +α[r t +γ(Q t+1 -Q t )]

[0166] In the formula, Q t Let r be the Q value calculated during the t-th iteration. t Let α be the reward value of the optimal action corresponding to the state information at the beginning of the t-th iteration, γ be the first preset parameter, t ∈ (1~T), and T be the preset iteration number threshold.

[0167] wherein the reward value r of the optimal action corresponding to the state information at the beginning of the tth iteration process is determined according to the following formula t :

[0168]

[0169] wherein Δη t is the difference between the target weakest oscillation damping ratio and the weakest oscillation damping ratio corresponding to the optimal action corresponding to the state information at the beginning of the tth iteration.

[0170] Embodiment 3:

[0171] This embodiment takes the 4-machine 2-area AC power grid as an example to elaborate the step 101 in Embodiment 1 in detail.

[0172] The 4-machine 2-area AC power grid refers to that two power generation areas and four generator groups are covered in the low-frequency oscillation area of a certain AC power grid.

[0173] In the step 101, when the low-frequency oscillation occurs in the AC power grid, a DC branch is connected between any two buses in the low-frequency oscillation area of the AC power grid, which can include the following steps:

[0174] As shown in FIG. 7, a DC branch is connected between the bus 7 and the bus 9 of the 4-machine 2-area AC power grid. Figure 3

[0175] The power optimization model corresponding to the 4-machine 2-area AC power grid with the connected DC branch is selected (the model is trained based on the initial state set constructed based on the historical unit output and DC branch transmission power of the 4-machine 2-area AC power grid with the connected DC branch, and the environment of the reinforcement learning network is also built based on the 4-machine 2-area AC power grid with the connected DC branch), to obtain the optimized unit output and the optimized DC branch transmission power of the 4-machine 2-area AC power grid with the connected DC branch.

[0176] Embodiment 4:

[0177] This embodiment takes the 4-machine 2-area AC power grid as an example to elaborate the step 101 in Embodiment 1 in detail.

[0178] The 4-machine 2-area AC power grid refers to that two power generation areas and four generator groups are covered in the low-frequency oscillation area of a certain AC power grid.

[0179] In the step 101, in addition to connecting a DC branch between any two buses of the AC power grid, a sending end for transmitting DC power to the AC power grid can also be externally connected to one of the buses with the connected DC branch, which can include the following steps:

[0180] As shown in FIG. 7, a DC branch is connected between the bus 7 and the bus 9 of the 4-machine 2-area AC power grid. Figure 4 ​As shown, the DC branch is connected between the bus 9 and the bus 12 of the 4-machine 2-area AC power grid, and a generator GEN5 is added to transmit DC power to the 4-machine 2-area AC power grid, and the generator parameters are the same as those of the generator GEN1, and at this time, the optimal unit output and DC branch transmission power can be obtained in the same way.

[0181] From the simulation experiments of Embodiment 3 and Embodiment 4, it can be concluded that the stability of the system is quite different when different unit outputs are matched with the same DC transmission power, and the ability of the system to withstand DC transmission power under different unit outputs is also quite different. By using the deep Q network to jointly optimize the DC transmission power and the generator active power, and comparing the fixed DC optimization of the generator active power, the damping ratio of the interval oscillation mode and the oscillation curve of the generator active power after the fault, it can be concluded that appropriate DC transmission power can improve the stability of the system, and the optimization of the combination of DC transmission power and generator output can greatly improve the stability of the system.

[0182] From the experimental data, it can also be concluded that the 4-machine 2-area system is the receiving end, and the generator output combination of the 4-machine 2-area system injected with DC is optimized and adjusted, and it is concluded that the optimized unit output can greatly improve the stability of the AC / DC system.

[0183] Embodiment 5

[0184] The power optimization model in Embodiment 1 of the application can also be trained by selecting another Q learning algorithm, and the following is the processing process thereof:

[0185] The data processing link of the deep Q network includes:

[0186] (1) The initial state is calculated by using the power system simulation software BPA for power flow calculation and small disturbance calculation, and the damping ratio calculated is All oscillation modes of the system are obtained in this way, and the weakest damping ratio of the interval oscillation mode is selected as the optimization target. The unit output combination at this time and the weakest damping ratio are used as the initialization origin of the deep Q network optimization calculation, and the weakest contrast group with the optimal weakest damping ratio.

[0187] The generator active output and the DC transmission power are taken as the quantities to be adjusted, and other quantities remain unchanged. Through the analysis of the operation state of the power system, the characteristic relationship is obtained, the damping ratio under all electromechanical oscillation modes is calculated, and the local oscillation mode and the interval oscillation mode are distinguished, and the weakest damping ratio of the interval oscillation mode is taken as the optimization target.

[0188] The generator output and DC transmission power are taken as the state input of the agent by the deep Q network, the change of active power each time is taken as the action of the agent, the change of damping ratio each time is taken as the reward of the agent (the reward value each time has positive and negative), and the maximum value of the weakest damping ratio of the system is obtained through programming iteration calculation.

[0189] A function composed of active power and target value can be established:

[0190] F = max (ζ min (P1, P2, …, P n , P dc ))

[0191] The optimization of the agent is encouraged according to the reward mechanism of the Q network, and the Q value is updated according to the following formula:

[0192] Q (s t+1 , a t+1 ) = Q (s t , a t ) + α [r t + γmax a Q (s t+1 , a t+1 ) - Q (s t , a t )]

[0193] The random calculation data accumulation link includes:

[0194] (1) Use python programming to call BPA to randomly modify the generator output, and perform power flow calculation and small disturbance calculation in BPA, record the changed unit output, DC transmission power and the weakest damping ratio, and input into the neural network

[0195] (2) The results of random calculation are input into the neural network, and the neural system prediction system is built under Tensorflow, and the possible data in the power system is predicted according to the past data.

[0196] The generator output and DC transmission power value, generator output change and DC transmission power change, the weakest damping ratio and the actual value of Q are taken as the input, and the predicted value of Q is taken as the output of the neural network. The difference between the actual value of Q and the predicted value of the neural network in the neural network can obtain the cost function, and according to the minimization of the cost function, the prediction accuracy of the neural network is improved.

[0197] The deep Q network comprises two neural networks, namely a latest data network and a historical data network. The two networks are neural networks with the same structure and are used to store generator output values, direct current transmission power states, and reward values (increase / decrease values of the weakest damping ratio) obtained by increasing / decrease of each variable in the current state. The historical data network is a frozen network that stores the behavior of the power system data with a lag; and the latest data network stores the latest neural network parameters of the generator output change and the reaction of the agent. That is, the historical data network is a historical version of the latest data network, and the historical data network exists to freeze the historical parameters, so as to provide a more accurate comparison tool when the latest data network predicts. Through learning of each item of data obtained by the latest generator output change and each item of data obtained by the generator unit output change with a lag, the latest data network prediction can be more accurate.

[0198] The iterative calculation link of the optimization algorithm comprises:

[0199] After a certain number of steps of random calculation, the increase / decrease of the agent output and the change of the direct current transmission power value are randomly selected, and the behavior, state, Q value, and reward value of each step are transmitted into the neural network for learning and training. When the training of the agent accumulates to a certain extent, the random output increase / decrease is stopped, the neural network is used to predict the result after each output change of the agent, and the next output selection of the agent is performed according to the greedy probability to select the behavior with the maximum reward value under the requirement of a certain low probability random selection behavior of the agent, and the Q learning is updated. Of course, after the execution process, the agent improves the prediction accuracy according to the update of the neural network.

[0200] Finally, the agent stops calculation and records the obtained data after the agent is optimized to a certain stage according to the above steps and meets certain conditions. Of course, the following flowchart is programmed in python, the BPA is mobilized for basic calculation through the program, and the required data is selected according to the calculation result. When the generator active power output and the direct current transmission power are changed, no adjustment is made to other parameters to ensure the accuracy of the control variable.

[0201] The left side of the flowchart is the random optimization of the agent for a certain number of steps. Through this optimization, the agent has a preliminary database, and can make a preliminary prediction in the subsequent iterative calculation. The middle part is the process of iterative optimization of the agent, and i represents the running algebra of the agent. When the agent iterates to a certain step or the power flow calculation does not converge, the iterative calculation of this generation is stopped, and the iterative calculation of the next generation is started. When the agent optimization reaches 100 generations, the agent stops optimization and outputs the optimal power system dispatching combination and the maximum value of the weakest damping ratio.

[0202] Those skilled in the art will appreciate that embodiments of the application can be devised for a method, a system, or a computer program product. Accordingly, the present application can be embodied in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code.

[0203] The present application is described in reference to the flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.

[0204] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.

[0205] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.

[0206] Finally, it should be noted that the above-mentioned embodiments are merely intended for describing the technical solutions of the present application, but not for limiting it. Although the present application has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that the technical solutions of the present application can still be modified or equivalent replaced without departing from the spirit and scope of the present application, and any modification or equivalent replacement should be covered in the protection scope of the claims of the present application.

Claims

1. A power optimization method for damping low frequency oscillations of an alternating current power grid, characterized in that, The method comprises: When low-frequency oscillation occurs in the alternating current power grid, a direct current branch is connected between any two buses in the low-frequency oscillation region of the alternating current power grid; Obtain initial unit output and initial direct current branch transmission power of the alternating current power grid; Substitute the initial unit output and the initial direct current branch transmission power into a power optimization model as initial state information, optimize and solve the model to obtain optimal unit output and optimal direct current branch transmission power of the alternating current power grid; The power optimization model is obtained by training a reinforcement learning network based on an initial state set and a preset action set constructed based on historical unit output and direct current branch transmission power of the alternating current power grid with the direct current branch connected thereto, and using a reinforcement learning algorithm. The power optimization model is obtained by training a reinforcement learning network based on an initial state set and a preset action set constructed based on historical unit output and direct current branch transmission power of the alternating current power grid with the direct current branch connected thereto, and using a reinforcement learning algorithm, and comprises: Take the i th state information in the initial state set as initial state information of the alternating current power grid with the direct current branch connected thereto in the reinforcement learning network; Substitute the state information and the action set at the beginning of each iteration process corresponding to the initial state information into a pre-constructed action selection model to obtain optimal action corresponding to the state information at the beginning of each iteration process, and the weakest oscillation damping ratio corresponding to the optimal action; Based on the weakest oscillation damping ratio corresponding to the optimal action, calculate the Q value of each iteration process corresponding to the initial state information using a Q learning algorithm; Update the state information based on the optimal action, and take the updated state information as the state information at the beginning of the next iteration process; When the preset iteration number threshold is reached, take the unit output and the direct current line transmission power in the updated state information in the iteration process with the maximum Q value as the optimal unit output and the optimal direct current line transmission power corresponding to the i th state information in the initial state set as the initial state information of the alternating current power grid with the direct current branch connected thereto in the reinforcement learning network; Repeat all the above steps until the optimal unit output and the optimal direct current line transmission power corresponding to each state information in the initial state set as the initial state information of the alternating current power grid with the direct current branch connected thereto in the reinforcement learning network are obtained, take the trained reinforcement learning network at this time as the power optimization model, and output the power optimization model; Wherein, i∈(1~N), N is the number of state information contained in the initial state set.

2. The method of claim 1, wherein, The process of obtaining the initial state set constructed based on historical unit output and direct current branch transmission power of the alternating current power grid with the direct current branch connected thereto comprises: Obtain the interference response information corresponding to each historical unit output and direct current branch transmission power of the alternating current power grid with the direct current branch connected thereto when the alternating current power grid is normally operated; When the interference response information corresponding to each historical unit output and direct current branch transmission power of the alternating current power grid with the direct current branch connected thereto contains a low-frequency oscillation region of the alternating current power grid, then the unit output and the direct current branch transmission power of the corresponding history are filled as a state information in the initial state set.

3. The method of claim 2, wherein, The interference response information corresponding to each history of unit output and DC branch transmission power of the AC power grid with the DC branch connected in normal operation is obtained, and the interference response information includes: A power flow file is constructed by using network parameters of the AC power grid with the DC branch connected, the Kth set of unit output and DC branch transmission power; The power flow file is imported into a BPA simulation platform to perform power flow calculation, and a power flow calculation result is obtained; The power flow calculation result is imported into the BPA simulation platform to perform small interference calculation, and the interference response information corresponding to the Kth set of unit output and DC branch transmission power is obtained; The interference response information includes: an AC power grid low-frequency oscillation region and a weakest oscillation damping ratio in various oscillation modes of the region, k∈(1~M), and M is the number of histories of unit output and DC branch transmission power of the AC power grid with the DC branch connected in normal operation.

4. The method of claim 1, wherein, The preset action set includes: an action corresponding to each unit in the AC power grid with the DC branch connected and an action corresponding to each DC branch in the AC power grid with the DC branch connected; Wherein, the action corresponding to the s-th unit in the AC power grid with DC branch is output increase ΔP s and output decrease ΔP s ; The action corresponding to the zth DC branch in the AC power grid connected with the DC branch is to increase the transmission power ΔP z and decrease the transmission power ΔP z ; wherein s∈(1~ξ s ), ξ s is the total number of units included in the AC power grid with the DC branch connected, z∈(1~ξ z ), ξ z is the total number of DC branches included in the AC power grid with the DC branch connected, ΔP s is the preset output change value of the s th unit included in the AC power grid with the DC branch connected, and ΔP2 is the preset transmission power change value of the z th DC branch included in the AC power grid with the DC branch connected.

5. The method of claim 1, wherein, The obtaining process of the pre-constructed action selection model includes: Each piece of state information in the initial state set and the preset action set are taken as input data of an initial neural network, and the optimal action corresponding to each piece of state information in the initial state set and the weakest oscillation damping ratio corresponding to the optimal action are taken as output data, so that the initial neural network is trained to obtain the pre-constructed action selection model.

6. The method of claim 5, wherein, The obtaining process of the optimal action corresponding to each piece of state information in the initial state set and the weakest oscillation damping ratio corresponding to the optimal action includes: Step A: initializing f=1; Step B: the fth action in the preset action set is applied on the basis of the ith piece of state information in the initial state set to obtain state information after the action is executed; Step C: interference response information corresponding to unit output and DC branch transmission power in the state information is obtained, the weakest oscillation damping ratio in oscillation damping ratios corresponding to the AC power grid low-frequency oscillation region in various oscillation modes of the region is obtained from the interference response information, and the weakest oscillation damping ratio is taken as the weakest oscillation damping ratio corresponding to the fth action; Step D: if f=δ, an action corresponding to the largest one of the weakest oscillation damping ratios corresponding to the first action to the δth action is taken as the optimal action corresponding to the ith piece of state information in the initial state set, and the optimal action and the weakest oscillation damping ratio corresponding to the optimal action are output; Wherein, f∈(1~δ), δ is the number of actions included in the preset action set, i∈(1~N), and N is the number of state information included in the initial state set.

7. The method of claim 1, wherein, The Q value of each iteration process corresponding to the initial state information is calculated by using a Q learning algorithm. The Q value Q calculated in the (t+1)th iteration process is determined by the following expression t+1 : Q t+1 = Q t + α[r t + γ(Q t+1 - Q t )] In the formula, Q t is the Q value calculated in the tth iteration process, r t is the reward value of the optimal action corresponding to the state information at the beginning of the tth iteration process, a is a first preset parameter, g is a second preset parameter, t e (1~T), and T is a preset iteration number threshold. wherein the reward value r of the optimal action corresponding to the state information at the beginning of the tth iteration process is determined according to the following formula t : where Δη t is the difference between the target least damped mode damping ratio and the least damped mode damping ratio corresponding to the optimal action corresponding to the state information at the beginning of the tth iteration.

8. A power control method for damping low frequency oscillations of an alternating current power grid, characterized in that, The method includes: The optimal unit output of the AC power grid and the optimal transmission power of the DC branch are obtained by using the method in any one of claims 1 to 7. The unit output of the AC power grid with the DC branch connected is controlled to be the optimal output, and the DC branch transmission power of the AC power grid with the DC branch connected is controlled to be the optimal transmission power.

9. A power optimization system for damping low frequency oscillations in an alternating current power grid, characterized by, The system includes: The access module is configured to access the DC branch between any two buses in a low-frequency oscillation region of the AC power grid when low-frequency oscillation occurs in the AC power grid; The acquisition module is configured to acquire initial unit output and initial DC branch transmission power of the AC power grid; The optimization solving module is configured to substitute the initial unit output and the initial DC branch transmission power as initial state information into a power optimization model, and to perform optimization solving on the model to obtain optimal unit output and optimal DC branch transmission power of the AC power grid; The power optimization model is obtained by training a reinforcement learning network based on an initial state set and a preset action set, and is constructed based on historical unit output and DC branch transmission power of the AC power grid with the DC branch accessed, and specifically includes the following steps: The i th state information in the initial state set is taken as initial state information of the AC power grid with the DC branch accessed in the reinforcement learning network; The state information and the action set at the beginning of each iteration process corresponding to the initial state information are substituted into a pre-constructed action selection model to obtain optimal action corresponding to the state information at the beginning of each iteration process and weakest oscillation damping ratio corresponding to the optimal action; Based on the weakest oscillation damping ratio corresponding to the optimal action, a Q learning algorithm is used to calculate the Q value of each iteration process corresponding to the initial state information; The state information is updated based on the optimal action, and the updated state information is taken as the state information at the beginning of the next iteration process; When a preset iteration number threshold is reached, the unit output and the DC line transmission power in the updated state information in the iteration process with the maximum Q value are taken as the optimal unit output and the optimal DC line transmission power corresponding to the initial state information in the i th state information in the initial state set as input by the reinforcement learning network with the DC branch accessed in the AC power grid; All the above steps are repeated until the optimal unit output and the optimal DC line transmission power corresponding to the initial state information in each state information in the initial state set as input by the reinforcement learning network with the DC branch accessed in the AC power grid are obtained, the trained reinforcement learning network at this time is taken as the power optimization model, and the power optimization model is output; Wherein, i∈(1~N), N is the number of state information contained in the initial state set.

Citation Information

Patent Citations

  • Direct-current transmission additional control method in frequency domain analysis

    CN102801178A