A method, device, electronic device and storage medium for primary frequency modulation of power grid
By applying the Dueling DQN algorithm in primary frequency regulation of the power grid, the problem of difficulty in establishing a reliable model in the prior art is solved, and more efficient and accurate grid frequency regulation is achieved.
Patent Information
- Application Number
- CN202310122079.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-13
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-02-13
AI Technical Summary
Due to the influence of data information, it is difficult to establish a reliable model for the actual unit's primary frequency regulation due to the influence of the amount of data.
The Dueling DQN algorithm in the deep reinforcement learning method is used to view the frequency regulation process of the power grid as the Markov decision-making process. By initializing the environment and network parameters, setting the discount factor and learning rate, action space and state initialization, sampling the Markov quadruple and updating the network parameters until the grid frequency fluctuates within the set range.
The efficiency and accuracy of the primary frequency regulation of the power grid is significantly improved. Compared with the DQN algorithm, stable fluctuations in the grid frequency are achieved within a shorter number of steps within the set range.
Smart Images

Figure CN116191466B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power grid frequency modulation, and in particular to a power grid primary frequency modulation method, device, electronic equipment and storage medium. Background Art
[0002] Grid frequency is an important indicator for ensuring power quality. If the grid frequency deviates from the rated value, the control system of the unit in the grid will automatically control the increase or decrease of the unit's active power, limit the change of grid frequency, and finally maintain a stable automatic control process for the grid frequency.
[0003] As new energy sources such as distributed power sources are connected to the power grid, their output may be affected by natural weather or the instability of power generation, which may lead to the emergence of a large number of harmonics, affecting the load of the power grid, reducing the stability of the power grid, and causing the line transmission power to exceed the thermal stability limit. Due to the complexity of the structure of the power grid itself, the primary frequency regulation of the generator set is an important means to maintain a dynamic balance between power generation and power consumption, as well as to ensure power quality and power grid safety. Therefore, it is necessary to provide a highly reliable primary frequency regulation method for the power grid.
[0004] At present, the existing primary frequency regulation methods of power grids mainly focus on establishing a transfer function model of primary frequency regulation and governor parameters. The governor controls the steam flow of the steam turbine by changing the opening of the turbine regulating valve, thereby achieving control of its speed. The steam turbine achieves closed-loop control through speed negative feedback. The speed is measured and amplified and fed back to the input end, and compared with the speed set value to obtain the speed deviation and adjust the steam valve opening accordingly, thereby changing the steam flow and achieving speed control. Due to the existence of negative feedback, if the actual output value is not equal to the set value, the regulation system will work until the actual output value is basically equal to the set value. However, due to the existence of speed deviation, the reliability of its model is not high.
[0005] The above-mentioned existing primary frequency regulation method of the power grid is difficult to establish a reliable model of the primary frequency regulation of the actual unit due to the influence of the amount of data information, which is a technical problem that needs to be solved urgently. Summary of the invention
[0006] The present invention mainly solves the technical problem that the existing primary frequency regulation method of a generator set is affected by the amount of data information and it is difficult to establish a reliable model of the primary frequency regulation of an actual generator set.
[0007] In order to solve this technical problem, the technical solution adopted by the present invention is to solve the above technical problem by using the Dueling DQN algorithm in the deep reinforcement learning method.
[0008] In a first aspect, the present invention provides a primary frequency regulation method for a power grid, wherein the primary frequency regulation process of the power grid is regarded as a Markov decision process, and comprises the following steps:
[0009] Initialize the environment; initialize the parameters ω in the current Q network and the parameters ω' in the target Q network; set the discount factor γ, learning rate μ, exploration rate ε, and the number of batch training samples;
[0010] Set the action space A in the Markov decision process as A = {a1, a2, ..., a m}, a m represents the mth action;
[0011] Initialize the current state S, i.e. the change in the active output of the generator set, the load power consumption and the system frequency, and use its state as the input of the Q network to obtain the Q value output corresponding to all actions in the network;
[0012] The action space A={a1,a2,...,a m}The actual action of transforming the power grid;
[0013] Randomly select action a according to the greedy strategy t , and the current state s t Input to the current Q network, the action interacts with the environment to obtain the state s at time t+1 t+1 And the reward function value r t , and then sample to obtain the Markov quadruple (s t ,a t ,r t ,s t+1 );
[0014] The sampled Markov quadruple (s t ,a t ,r t ,s t+1 ) is stored in the experience pool and the updated state s is obtained t+1 ;
[0015] Update the network parameters and continue training until the grid frequency fluctuates within the set range.
[0016] Preferably, the action space A = {a1, a2, a3, a4}.
[0017] Preferably, the action space A={a1, a2, ..., a n The actual actions of transforming the power grid include:
[0018] The action space A = {a1, a2, a3, a4} is transformed into the actual action of the power grid: a1 is the increase in the active output value of generator set 1, a2 is the decrease in the active output value of generator set 1, a3 is the increase in the active output value of generator set 2, and a4 is the decrease in the active output value of generator set 2.
[0019] Furthermore, in the primary frequency regulation process of the power grid, the relationship between the active load and the frequency in the power grid needs to be considered as follows:
[0020]
[0021] P f is the total active power consumed by the grid when the frequency is f, α0, α1, α2, …α n is the ratio coefficient of each type of load to the rated load, and satisfies α0+α1+α2+…α n =1;P fd is the load under rated grid power;
[0022] Under this condition, the objective function is set as:
[0023]
[0024] i is the current number of loads, n is the current number of generator sets, I is the total number of loads, N is the total number of generator sets, P ci is the total power consumption of the current load, is the total active output value of the current generator set;
[0025] The active output value constraint of the generator set is:
[0026]
[0027] is the active output value of the nth generator set, and is the active output value of the nth generator set.
[0028] The upper and lower limits of load power consumption are:
[0029]
[0030] P c is the load power consumption, and P c They are the upper and lower limit constraints of load power consumption respectively.
[0031] Preferably, the reward function in the Markov decision process is set to:
[0032]
[0033] α, β, γ are constants, and f is the frequency of the power grid operation.
[0034] Preferably, calculating the reward function value according to the reward function specifically includes:
[0035] If the load power consumption in the power grid is equal to the sum of the active output values of the generator sets, a positive reward β is given; otherwise, a negative reward -αf is given;
[0036] If the active output value of the generator set and the load power consumption exceed the set upper and lower limits, a negative reward γ will be given.
[0037] Preferably, the setting range is 50±0.2 Hz.
[0038] In a second aspect, the present invention provides a primary frequency modulation device for a power grid, comprising the following modules:
[0039] The environment initialization module is used to initialize the environment; initialize the parameters ω in the current Q network and the parameters ω' in the target Q network; set the discount factor γ, learning rate μ, exploration rate ε, and the number of batch training samples.
[0040] The action space setting module is used to set the action space A={a1,a2,...,a n}, a n Indicates the nth action;
[0041] The state initialization module is used to initialize the current state S, that is, the change value of the active output of the generator set, the load power consumption and the frequency of the system, and use its state as the input of the Q network to obtain the Q value output corresponding to all actions in the network;
[0042] The action space conversion module is used to convert the action space A = {a1, a2, ..., a n}The actual action of transforming the power grid;
[0043] The four-tuple sampling module is used to randomly select action a according to the greedy strategy t , and the current state s t Input to the current Q network, the action interacts with the environment to obtain the state s at time t+1 t+1 And the reward function value r t , and then sample to obtain the Markov quadruple (s t ,a t ,r t ,s t+1 );
[0044] The state update module is used to update the sampled Markov quadruple (s t ,a t ,r t ,s t+1 ) is stored in the experience pool and the updated state s is obtained t+1 ;
[0045] The network update module is used to update the network parameters and continuously perform training until the system frequency fluctuates within the set range.
[0046] In a third aspect, the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the primary frequency modulation method of the power grid when executing the program.
[0047] In a fourth aspect, the present invention provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the primary frequency modulation method of the power grid.
[0048] Compared with the prior art, the present invention has the following beneficial technical effects:
[0049] This application constructs a Markov decision process and a simple mathematical model of the primary frequency modulation process of the power grid, and uses the DuelingDQN algorithm to directly perform primary frequency modulation on the power grid. It is verified through experimental simulation that the DQN algorithm tends to be stable at about 4000 steps, but the frequency value fluctuates back and forth between 48.99HZ and 50.84HZ, and the interval where the fluctuation value is located exceeds the frequency standard range; when the Dueling DQN algorithm adjusts the primary frequency of the power grid, at 2000 steps, the power grid frequency fluctuates back and forth between 49.82HZ and 50.17HZ and tends to be stable. It proves that the efficiency and accuracy of primary frequency modulation using the Dueling DQN algorithm in this invention are significantly improved compared with the DQN algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:
[0051] Figure 1 A flow chart of a primary frequency modulation method for a power grid according to the present invention;
[0052] Figure 2 This is a framework diagram of a power grid primary frequency regulation program based on Dueling DQN algorithm of the present invention;
[0053] Figure 3 This is a frequency change diagram of different iteration times under the DQN algorithm in an embodiment of the present invention;
[0054] Figure 4 This is a frequency change diagram of different iteration times under the Dueling DQN algorithm in an embodiment of the present invention;
[0055] Figure 5 It is a structural schematic diagram of a primary frequency modulation device of a power grid according to the present invention;
[0056] Figure 6The figure is a schematic structural diagram of an electronic device of the present invention. DETAILED DESCRIPTION
[0057] In order to have a clearer understanding of the technical features, purposes and effects of the present invention, specific embodiments of the present invention are now described in detail with reference to the accompanying drawings.
[0058] The present invention considers the following problems with the Q-learning algorithm: (1) the amount of memory required to save and update the table will increase as the number of states increases; (2) it takes too long to explore each state to create the required Q-table, so some scholars have introduced a method of approximating the Q-value function using a neural network to solve this problem. However, the DQN algorithm will produce the problem of overestimating the Q value, so the Dueling DQN algorithm is used to solve the above problem.
[0059] First, let me introduce the principle of Dueling DQN algorithm:
[0060] The DQN algorithm uses a neural network to approximate the optimal action value function, that is, the function Q(s, a; w) to approximate the action value Q * (s, a) to get the optimal strategy. The Dueling DQN algorithm introduces a dual network on its basis, that is, using the neural network A(s, a; w, α) to approximate the optimal advantage function A * (s, a), use the neural network V(s; w, β) to approximate the optimal state value function V * (s). The input of the network is the state, and the output is the state value V of the state and the advantage value A of each action, that is, Q(s,a;w,α,β)=V(s;w,β)+A(s,a;w,α), where w is the network parameter, and α, β represent the parameters of the two fully connected layer networks respectively.
[0061] DuelingDQN algorithm uses experience replay mechanism, which can break the correlation between data, that is, the four-tuple (s t ,a t ,r t ,s t+1 ) is stored in the memory unit, where s t is the state at time t, a t is the action at time t, r t is the reward value obtained at time t, s t+1 is the state at time t+1. The network weights are updated using the stochastic gradient descent (SGD) method, that is, the root mean square error between the Q network and the target Q network is estimated to update the network weights. The loss function of Q network training is:
[0062]
[0063] Use SGD to update the weight value:
[0064]
[0065] Where γ is the discount factor, w i is the estimated network parameter at the i-th iteration, w i - is the target network parameter at the i-th iteration.
[0066] Target network parameters w i - The update formula is:
[0067] w i - ←μw i +(1-μ)w i - (3)
[0068] where μ is the learning rate, μ∈(0,1).
[0069] Based on the Dueling DQN algorithm, the present invention provides a primary frequency modulation method for a power grid, such as Figure 1 , Figure 2 As shown, the following steps are included:
[0070] S1: Initialize the environment; initialize the parameters ω in the current Q network and ω' in the target Q network; set the discount factor γ, learning rate μ, exploration rate ε, and the number of batch training samples;
[0071] S2: Set the action space A in the Markov decision process = {a1, a2, ..., a m}, a m represents the mth action;
[0072] S3: Initialize the current state S, i.e. the change in the active output of the generator set, the load power consumption and the system frequency, and use its state as the input of the Q network to obtain the Q value output corresponding to all actions in the network;
[0073] S4: Set the action space A = {a1, a2, ..., a m}The actual action of transforming the power grid;
[0074] S5: Randomly select action a according to the greedy strategy t , and the current state s t Input to the current Q network, the action interacts with the environment to obtain the state s at time t+1 t+1 And the reward function value r t , and then sample to obtain the Markov quadruple (s t ,a t,r t ,s t+1 );
[0075] S6: The sampled Markov quadruple (s t ,a t ,r t ,s t+1 ) is stored in the experience pool and the updated state s is obtained t+1 ;
[0076] S7: Update network parameters and continue training until the grid frequency fluctuates within the set range.
[0077] The primary frequency modulation method of the power grid described in the embodiment of the present invention is performed by an electronic device. The electronic device can be any type of electronic device; for example, the electronic device can be but is not limited to at least one of the following: a server, a computer, a tablet computer or other electronic devices.
[0078] It should be noted that the relationship between active load and frequency in the power grid needs to be considered during the primary frequency regulation process:
[0079]
[0080] P f is the total active power consumed by the grid when the frequency is f, α0, α1, α2, …α n is the ratio coefficient of each type of load to the rated load, and satisfies α0+α1+α2+…α n =1;P fd is the load under rated grid power;
[0081] Under this condition, the objective function is set as:
[0082]
[0083] i is the current number of loads, n is the current number of generator sets, I is the total number of loads, N is the total number of generator sets, P ci is the total power consumption of the current load, It is the total active output value of the current generator set.
[0084] Active output value constraints of generator sets:
[0085]
[0086] is the active output value of the nth generator set, and is the active output value of the nth generator set.
[0087] Upper and lower limits of load power consumption:
[0088]
[0089] P c is the load power consumption, and P c They are the upper and lower limit constraints of load power consumption respectively.
[0090] The reward function is set as:
[0091]
[0092] α and β are constants, and f is the frequency of the power grid operation.
[0093] If the power consumption in the power grid is equal to the sum of the active output values of the generator sets, a positive reward β is given; otherwise, a negative reward -αf is given; if the active output value of the generator sets and the power consumption end exceed the set upper and lower limits, a negative reward γ is given. In the simulation verification, the value of α is set to -0.1, the value of β is 10, and the value of γ is -1.
[0094] Reinforcement learning may cause dimensionality explosion when facing high-dimensional action state space. Usually, the DQN algorithm uses a neural network to approximate the action value function to represent the Q function in a high-dimensional continuous space. Using the target network and the evaluation network, the target network obtains the evaluation value of the reward. In the Dueling DQN algorithm, each action value function is divided into a state value function and an advantage function. Its input is the same as the DQN algorithm, which is the state. The output is the state value V of the state and the advantage value A of each action, and finally the action value of each action is output.
[0095] Example:
[0096] The process of primary frequency regulation of the power grid is regarded as a Markov decision process. Taking the three-machine nine-node system as an example, generator set 1 is a balanced unit, and the switch action that controls the active output of the unbalanced generator set is regarded as the action space A = {a1, a2, a3, a4}, which is set into four actions respectively. The action space A = {a1, a2, a3, a4} is converted into the actual action of the power grid. a1 is the increase of the active output value of generator set 1, a2 is the decrease of the active output value of generator set 1, a3 is the increase of the active output value of generator set 2, and a4 is the decrease of the active output value of generator set 2. The frequency change of the power grid is the state space. The reward function is set as (8). If the frequency change of the power grid is 50 ± 0.2 Hz, it means that the load consumption causes the frequency fluctuation, and it is adjusted to the frequency value under the normal operation of the power grid.
[0097] Based on the above embodiment 1, the Dueling DQN algorithm and the DQN algorithm in the present invention are respectively applied to the primary frequency regulation of the power grid. The comparison results of the frequency regulation effects are as follows: Figure 3 , 4 As shown, Figure 3 The frequency change diagram corresponding to the Dueling DQN algorithm at different iteration times, Figure 4 The frequency change diagram corresponding to different iteration numbers under Dueling DQN algorithm. By comparing the DQN algorithm and Dueling DQN algorithm, it can be obtained that the DQN algorithm tends to be stable at about 4000 steps, but the frequency value fluctuates between 48.99HZ and 50.84HZ, and the interval of the fluctuation value exceeds the frequency standard range. When the DuelingDQN algorithm adjusts the primary frequency of the power grid, at 2000 steps, the power grid frequency fluctuates between 49.82HZ and 50.17HZ and tends to be stable, verifying the efficiency and accuracy of the primary frequency regulation of the DuelingDQN algorithm of the present invention.
[0098] It should be noted that the above two embodiments of the generator sets are only preferred embodiments of the present invention. The primary frequency regulation method of the power grid described in the present invention is also applicable to a power grid system including multiple generator sets. Different generator sets only have different numbers of action spaces set, and the rest are consistent with the steps of the above embodiments, and can achieve corresponding beneficial effects.
[0099] A power grid primary frequency regulation device provided by the present invention is described below. The power grid primary frequency regulation device described below and the power grid primary frequency regulation method described above can be referenced to each other.
[0100] like Figure 5 As shown, a primary frequency regulation device for a power grid includes the following modules:
[0101] The environment initialization module 510 is used to initialize the environment; initialize the parameters ω in the current Q network and the parameters ω' in the target Q network; set the discount factor γ, the learning rate μ, the exploration rate ε and the number of batch training samples,
[0102] The action space setting module 520 is used to set the action space A in the Markov decision process = {a1, a2, ..., a n}, a n Indicates the nth action;
[0103] The state initialization module 530 is used to initialize the current state S, i.e., the change value of the active output of the generator set, the load power consumption and the frequency of the system, and use its state as the input of the Q network to obtain the Q value output corresponding to all actions in the network;
[0104] The action space conversion module 540 is used to convert the action space A = {a1, a2, ..., a n}The actual action of transforming the power grid;
[0105] The four-tuple sampling module 550 is used to randomly select action a according to the greedy strategy t , and the current state s t Input to the current Q network, the action interacts with the environment to obtain the state s at time t+1 t+1 And the reward function value r t , and then sample to obtain the Markov quadruple (s t ,a t ,r t ,s t+1 );
[0106] The state updating module 560 is used to update the sampled Markov quadruple (s t ,a t ,r t ,s t+1 ) is stored in the experience pool and the updated state s is obtained t+1 ;
[0107] The network update module 570 is used to update the network parameters and continuously perform training until the system frequency fluctuates within a set range.
[0108] like Figure 6 As shown, an example of a physical structure diagram of an electronic device, the electronic device may include: a processor (processor) 610, a communication interface (Communications Interface) 620, a memory (memory) 630 and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call the logic instructions in the memory 630 to execute the steps of the above-mentioned power grid primary frequency regulation method, specifically including: initializing the environment; initializing the parameters ω in the current Q network and the parameters ω' in the target Q network; setting the discount factor γ, the learning rate μ, the exploration rate ε and the number of batch training samples; setting the action space A={a1,a2,...,a m}, a m represents the mth action; initialize the current state S, i.e., the change in the active output of the generator set, the load power consumption, and the frequency of the system, and use its state as the input of the Q network to obtain the Q value output corresponding to all actions in the network; convert the action space A = {a1, a2, ..., a m}Convert the actual action of the power grid; randomly select action a according to the greedy strategy t , and the current state st Input to the current Q network, the action interacts with the environment to obtain the state s at time t+1 t+1 And the reward function value r t , and then sample to obtain the Markov quadruple (s t ,a t ,r t ,s t+1 ); The sampled Markov quadruple (s t ,a t ,r t ,s t+1 ) is stored in the experience pool and the updated state s is obtained t+1 ; Update network parameters and continue training until the grid frequency fluctuates within the set range.
[0109] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random15 Access Memory), disk or optical disk and other media that can store program codes.
[0110] On the other hand, an embodiment of the present invention further provides a storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of implementing the above-mentioned power grid primary frequency regulation method specifically include: initializing the environment; initializing the parameter ω in the current Q network and the parameter ω' in the target Q network; setting the discount factor γ, the learning rate μ, the exploration rate ε and the number of batch training samples; setting the action space A in the Markov decision process = {a1, a2, ..., a m}, a m represents the mth action; initialize the current state S, i.e., the change in the active output of the generator set, the load power consumption, and the frequency of the system, and use its state as the input of the Q network to obtain the Q value output corresponding to all actions in the network; convert the action space A = {a1, a2, ..., a m}Convert the actual action of the power grid; randomly select action a according to the greedy strategy t , and the current state s tInput to the current Q network, the action interacts with the environment to obtain the state s at time t+1 t+1 And the reward function value r t , and then sample to obtain the Markov quadruple (s t ,a t ,r t ,s t+1 ); The sampled Markov quadruple (s t ,a t ,r t ,s t+1 ) is stored in the experience pool and the updated state s is obtained t+1 ; Update network parameters and continue training until the grid frequency fluctuates within the set range.
[0111] The embodiment of the present invention provides a method, device, electronic device and storage medium for primary frequency modulation of power grid, which initializes the environment; initializes network parameters; sets parameters such as discount factor, learning rate, action space, etc.; initializes the current state; converts the action space into the actual action of the power grid; randomly selects actions according to the greedy strategy, and inputs the current state into the current Q network, obtains the state and reward function value of the next moment after the action interacts with the environment, and then samples to obtain the Markov quadruple; stores the Markov quadruple in the experience pool, and obtains the updated state; updates the network parameters, and continuously trains until the power grid frequency fluctuates within the set range. The problem of low reliability of the mechanism model established in the prior art is solved, and compared with the DQN algorithm, the efficiency and accuracy of primary frequency modulation of power grid are significantly improved.
[0112] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.
[0113] The serial numbers of the embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. In a unit claim that lists several means, several of these means may be embodied by the same hardware item. The use of the words first, second, and third, etc. does not indicate any order and these words may be interpreted as identifiers.
[0114] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for primary frequency modulation of a power grid, characterized in that: The primary frequency regulation process of the power grid is regarded as a Markov decision process, which includes the following steps: Initialize the environment; initialize the parameters ω in the current Q network and the parameters ω' in the target Q network; set the discount factor γ, learning rate μ, exploration rate ε, and the number of batch training samples; Set the action space A in the Markov decision process as A = {a1, a2, ..., a m }, a m represents the mth action; Initialize the current state S, i.e. the change in the active output of the generator set, the load power consumption and the system frequency, and use its state as the input of the Q network to obtain the Q value output corresponding to all actions in the network; The action space A={a1,a2,...,a m }The actual action of transforming the power grid; Randomly select action a according to the greedy strategy t , and the current state s t Input to the current Q network, the action interacts with the environment to obtain the state s at time t+1 t+1 And the reward function value r t , and then sample to obtain the Markov quadruple (s t ,a t ,r t ,s t+1 ); The sampled Markov quadruple (s t ,a t ,r t ,s t+1 ) is stored in the experience pool and the updated state s is obtained t+1 ; Update the network parameters and continue training until the grid frequency fluctuates within the set range.
2. The primary frequency modulation method of the power grid according to claim 1, characterized in that: The action space A = {a1, a2, a3, a4}.
3. The primary frequency modulation method of the power grid according to claim 2, characterized in that: The action space A={a1,a2,...,a n The actual actions of transforming the power grid include: The action space A = {a1, a2, a3, a4} is transformed into the actual action of the power grid: a1 is the increase in the active output value of generator set 1, a2 is the decrease in the active output value of generator set 1, a3 is the increase in the active output value of generator set 2, and a4 is the decrease in the active output value of generator set 2.
4. The primary frequency modulation method of the power grid according to claim 1, characterized in that: In the primary frequency regulation process of the power grid, it is necessary to consider the relationship between the active load and the frequency in the power grid: P f is the total active power consumed by the grid when the frequency is f, α0, α1, α2, …α n is the ratio coefficient of each type of load to the rated load, and satisfies α0+α1+α2+…α n =1;P fd is the load under rated grid power; Under this condition, the objective function is set as: i is the current number of loads, n is the current number of generator sets, I is the total number of loads, N is the total number of generator sets, P ci is the total power consumption of the current load, is the total active output value of the current generator set; The active output value constraint of the generator set is: is the active output value of the nth generator set, and is the active output value of the nth generator set; The upper and lower limits of load power consumption are: P c is the load power consumption, and P c They are the upper and lower limit constraints of load power consumption respectively.
5. The primary frequency modulation method of the power grid according to claim 4, characterized in that: The reward function in the Markov decision process is set as: α, β, γ are constants, and f is the frequency of the power grid operation.
6. The method for primary frequency modulation of a power grid according to claim 5, characterized in that: The reward function value is calculated according to the reward function, including: If the load power consumption in the power grid is equal to the sum of the active output values of the generator sets, a positive reward β is given; otherwise, a negative reward -αf is given; If the active output value of the generator set and the load power consumption exceed the set upper and lower limits, a negative reward γ will be given.
7. The primary frequency modulation method of a power grid according to claim 1, characterized in that: The setting range is 50±0.2 Hz.
8. A primary frequency modulation device for a power grid, characterized in that: Includes the following modules: The environment initialization module is used to initialize the environment; initialize the parameters ω in the current Q network and the parameters ω' in the target Q network; set the discount factor γ, learning rate μ, exploration rate ε, and the number of batch training samples. The action space setting module is used to set the action space A={a1,a2,...,a n }, a n Indicates the nth action; The state initialization module is used to initialize the current state S, that is, the change value of the active output of the generator set, the load power consumption and the frequency of the system, and use its state as the input of the Q network to obtain the Q value output corresponding to all actions in the network; The action space conversion module is used to convert the action space A = {a1, a2, ..., a n }The actual action of transforming the power grid; The four-tuple sampling module is used to randomly select action a according to the greedy strategy t , and the current state s t Input to the current Q network, the action interacts with the environment to obtain the state s at time t+1 t+1 And the reward function value r t , and then sample to obtain the Markov quadruple (s t ,a t ,r t ,s t+1 ); The state update module is used to update the sampled Markov quadruple (s t ,a t ,r t ,s t+1 ) is stored in the experience pool and the updated state s is obtained t+1 ; The network update module is used to update the network parameters and continuously perform training until the system frequency fluctuates within the set range.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the power grid primary frequency regulation method according to any one of claims 1 to 7 are implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the primary frequency regulation method of a power grid as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Power distribution network voltage regulation method based on deep reinforcement learning algorithm
CN111884213A
AGC unit dynamic optimization method based on deep reinforcement learning
CN112186811A