A frequency control method and device based on model-based reinforcement learning
By employing a frequency control method based on model-based reinforcement learning, and utilizing the DQL model and a stochastic greedy strategy to control the output power of generator units, the frequency stability problem of traditional power grids after the integration of new energy sources is solved, thereby improving the stability of distributed power grids and the capacity for new energy absorption.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN POWER SUPPLY BUREAU
- Filing Date
- 2022-12-16
- Publication Date
- 2026-07-31
AI Technical Summary
Traditional centralized automatic generation control mode is difficult to meet the development and operation conditions of the power grid, while frequency regulation of large-scale new energy access to the power system is difficult to maintain frequency stability in various areas of the distributed power grid under highly random load conditions.
A frequency control method based on model-based reinforcement learning is adopted. The target ACE state data, initial control power command value and target environment parameters are collected by the model-based controller to generate target environment information. The matrix is updated using a DQL model and a random greedy strategy is used to determine the power command difference to control the output power of the generator set.
Maintaining frequency stability in various regions of the distributed power grid under highly random load conditions improves the distributed power grid's ability to absorb intermittent energy sources such as wind power and photovoltaic power.
Smart Images

Figure CN115882510B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of frequency control technology for distributed power grids, and in particular to a frequency control method and apparatus based on model-based reinforcement learning. Background Technology
[0002] With the supply of fossil energy becoming increasingly tight, the development of new energy sources can solve the environmental degradation caused by the combustion of fossil fuels. Integrated energy systems that combine multiple distributed energy sources such as power generation, load, gas, heat, and storage are imperative. However, the grid connection of large-scale distributed new energy sources will bring strong random disturbances and frequency instability problems caused by reduced inertia of traditional units, lack of auxiliary frequency support, and insufficient frequency regulation capacity, which will affect the operation and control of traditional power systems.
[0003] Currently, the traditional centralized automatic generation control mode is difficult to meet the development and operation conditions of the power grid, while the existing frequency regulation of large-scale new energy access to the power system has the problem of difficulty in maintaining the frequency stability of various areas of the distributed power grid under highly random load conditions. Summary of the Invention
[0004] This invention provides a frequency control method and apparatus based on model-based reinforcement learning, which solves the technical problem that traditional centralized automatic generation control mode is difficult to meet the development and operation conditions of the power grid, while the existing frequency regulation of large-scale new energy access to the power system has the problem of difficulty in maintaining the frequency stability of various areas of the distributed power grid under highly random load conditions.
[0005] The first aspect of this invention provides a frequency control method based on model-based reinforcement learning, comprising: In response to the received trigger information, the model controller in the new energy power system to be controlled is invoked to collect the target ACE status data, initial control power command value, target parameters and target environmental parameters corresponding to the new energy power system to be controlled; Using the target ACE state data, the target parameters, and the target environment parameters, the corresponding initial environment information is determined and stored in a preset information pool; Extract a first preset amount of first environmental information from the preset information pool, and combine it with the initial environmental information to generate a second preset amount of target environmental information; The initial matrices are updated using the target environment information through a preset DQL model to generate multiple corresponding target matrices. A random greedy strategy is used to select associated second target actions from each of the target matrices, and the corresponding target power command difference is determined based on the power difference associated with each second target action. Based on the calculation result of the difference between the initial control power command value and the target power command, the output power of the generator set in the new energy power system to be controlled is controlled.
[0006] Optionally, the step of responding to the received trigger information and invoking the model controller within the new energy power system to be controlled to collect the target ACE state data, initial control power command value, target parameters, and target environmental parameters corresponding to the new energy power system to be controlled includes: In response to the received trigger information, the model controller in the new energy power system to be controlled is invoked to collect the initial ACE state data, initial ACE control power command value, target parameters and target environmental parameters corresponding to the new energy power system to be controlled; The initial ACE state data is used as input to a preset ACE state model, and the corresponding target ACE state data is output.
[0007] Optionally, the target environment parameters include a first environment parameter and a second environment parameter. The step of determining the corresponding initial environment information using the target ACE state data, the target parameters, and the target environment parameters, and storing the initial environment information in a preset information pool, includes: Based on the second environmental parameters, any initial matrix is randomly selected from the target parameters; The first target action is selected from the target action set data within the target parameters using a random greedy strategy. The first target action is executed using the target ACE state data, and a corresponding target reward value is generated through a preset ACE state model; Using the first environmental parameters, the first target action, the target reward value, and the second environmental parameters, corresponding initial environmental information is generated. The initial environmental information is stored in a preset information pool.
[0008] Optionally, the step of extracting a first preset amount of first environmental information from the preset information pool and combining it with the initial environmental information to generate a second preset amount of target environmental information includes: Extract a first preset amount of first environmental information from the information pool; Using the initial environmental information and the first preset amount of first environmental information, a second preset amount of target environmental information is generated.
[0009] Optionally, the step of selecting associated second target actions from each of the target matrices using a random greedy strategy, and determining the corresponding target power command difference based on the power difference associated with each of the second target actions, includes: A random greedy strategy is used to select associated target actions from each of the target matrices to generate multiple second target actions; Multiple initial power command differences are generated by matching associated power differences from the target action set data using multiple second target actions; The sum of the multiple initial power command differences is calculated to generate the target power command difference.
[0010] Optionally, the step of controlling the output power of the generator sets in the new energy power system to be controlled based on the calculation result of the difference between the initial control power command value and the target power command includes: Calculate the sum of the difference between the initial control power command value and the target power command value to generate the target total power command value; The target total power is associated with the output of the target total power command value by the generator set within the new energy power system to be controlled.
[0011] Optionally, it also includes: Obtain the frequency deviation data and tie-line power deviation data of the new energy power system to be controlled at the current moment; Use the target total power at the current moment as the new initial control power command value; Based on the frequency deviation data and the tie-line power deviation data, new target ACE state data is determined, and the process jumps to execute the step of using the target ACE state data, the target parameters, and the target environmental parameters to determine the corresponding initial environmental information and store the initial environmental information in a preset information pool.
[0012] A second aspect of the present invention provides a frequency control device based on model-based reinforcement learning, comprising: The response module is used to respond to the received trigger information and call the model controller in the new energy power system to be controlled to collect the target ACE status data, initial control power command value, target parameters and target environmental parameters corresponding to the new energy power system to be controlled; The initial environment information module is used to determine the corresponding initial environment information using the target ACE state data, the target parameters, and the target environment parameters, and to store the initial environment information in a preset information pool. The target environment information module is used to extract a first preset amount of first environment information from the preset information pool, and combine it with the initial environment information to generate a second preset amount of target environment information. The target matrix module is used to update all the initial matrices using the target environment information through a preset DQL model, thereby generating multiple corresponding target matrices. The target power command difference module is used to select associated second target actions from each of the target matrices using a random greedy strategy, and determine the corresponding target power command difference based on the power difference associated with each second target action. The control module is used to control the output power of the generator set in the new energy power system to be controlled based on the calculation result of the difference between the initial control power command value and the target power command value.
[0013] A third aspect of the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the frequency control method based on model reinforcement learning as described in any of the preceding claims.
[0014] The fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed, implements the frequency control method based on model-based reinforcement learning as described in any of the preceding claims.
[0015] As can be seen from the above technical solutions, the present invention has the following advantages: In response to the received trigger information, the model controller within the new energy power system to be controlled is invoked to collect the target ACE state data, initial control power command value, target parameters, and target environmental parameters corresponding to the new energy power system to be controlled. Using the target ACE state data, target parameters, and target environmental parameters, the corresponding initial environmental information is determined and stored in a preset information pool. A first preset amount of first environmental information is extracted from the preset information pool, and combined with the initial environmental information, a second preset amount of target environmental information is generated. The target environmental information is then used to update all initial matrices using a preset DQL model, generating multiple corresponding target matrices. A random greedy strategy is employed to select from each... The target matrix selects the associated second target action, and determines the corresponding target power command difference based on the power difference associated with each second target action. Based on the calculation result of the initial control power command value and the target power command difference, the output power of the generator units in the new energy power system to be controlled is controlled. This solves the technical problem that the traditional centralized automatic generation control mode is difficult to meet the development and operation conditions of the power grid, and that the frequency regulation of the existing large-scale new energy access power system has difficulty in maintaining the frequency stability of various areas of the distributed power grid under strong random load conditions. It realizes the maintenance of frequency stability of various areas of the distributed power grid under strong random load conditions and improves the ability of the distributed power grid to absorb intermittent energy such as wind power and photovoltaic power. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating the steps of a frequency control method based on model-based reinforcement learning, provided in Embodiment 1 of the present invention. Figure 2 This is a flowchart of the steps of a frequency control method based on model-based reinforcement learning provided in Embodiment 2 of the present invention; Figure 3 This is a schematic diagram of the new energy power system to be controlled in Embodiment 2 of the present invention; Figure 4 This is a schematic diagram of the process for generating the target total power in Embodiment 2 of the present invention; Figure 5 This is a schematic diagram of the generator set in Embodiment 2 of the present invention; Figure 6 This is a schematic diagram of the first control performance index of a two-region power grid using a model controller and existing Dyna-QL and Q controllers in Embodiment 2 of the present invention. Figure 7 This is a schematic diagram of the second control performance index of a two-region power grid using a model controller and existing Dyna-QL and Q controllers in Embodiment 2 of the present invention; Figure 8 This is a schematic diagram of the third control performance index of a two-region power grid using a model controller and existing Dyna-QL and Q controllers in Embodiment 2 of the present invention; Figure 9 This is a structural block diagram of a frequency control device based on model-based reinforcement learning provided in Embodiment 3 of the present invention. Detailed Implementation
[0018] This invention provides a frequency control method and apparatus based on model-based reinforcement learning, which addresses the technical problem that traditional centralized automatic generation control modes cannot meet the development and operation conditions of the power grid, while existing frequency regulation for large-scale new energy access to the power system has the problem of difficulty in maintaining frequency stability in various regions of the distributed power grid under highly random load conditions.
[0019] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0020] Please see Figure 1 , Figure 1 The flowchart illustrates the steps of a frequency control method based on model-based reinforcement learning, as provided in Embodiment 1 of the present invention.
[0021] This invention provides a frequency control method based on model-based reinforcement learning, comprising: Step 101: In response to the received trigger information, call the model controller in the new energy power system to be controlled to collect the target ACE status data, initial control power command value, target parameters and target environmental parameters corresponding to the new energy power system to be controlled.
[0022] The term "new energy power system under control" refers to a new type of power system that is based on new energy sources, uses a robust and intelligent power grid as its hub platform, and is supported by the interaction of power generation, grid, load, and storage, as well as multi-energy complementarity. It is characterized by being clean and low-carbon, safe and controllable, flexible and efficient.
[0023] A model-based controller refers to a Dyna-DQL reinforcement learning controller that incorporates a model.
[0024] A generator set refers to a thermal power turbine generator set and a superconducting magnetic energy storage unit that takes into account generator output rate control (GRC). The thermal power generator set includes a governor module, a GRC module, a turbine module, and a power limiting module.
[0025] Triggering information refers to receiving information for frequency regulation of the new energy power system.
[0026] Target ACE state data refers to one of the parameters used for initial environmental information. It is the ACE state data generated by inputting the initial ACE state data corresponding to the current ACE value of the new energy power system to be controlled into a preset ACE state model.
[0027] The initial control power command value refers to the output power of the generator unit in the new energy power system to be controlled at the current moment, according to the power command sent by the model controller at the previous moment.
[0028] The target parameters refer to the environmental state S and the target action. Target action set data A, two initial matrices of the same type Initial matrix Initialization parameters: learning rate α, discount factor γ, number of environmental information N.
[0029] The target environmental parameters refer to the first environmental parameter and the second environmental parameter. It is worth mentioning that the first environmental parameter refers to the environmental state at the current moment. The second environmental parameter refers to the environmental state at the next moment. .
[0030] In this embodiment of the invention, upon receiving information about frequency modulation for a new energy power system, the system calls upon the model controller within the new energy power system to collect the target ACE state data, initial control power command value, target parameters, and target environmental parameters corresponding to the new energy power system to be controlled.
[0031] Step 102: Using the target ACE state data, target parameters, and target environment parameters, determine the corresponding initial environment information and store the initial environment information in a preset information pool.
[0032] An information pool refers to a data repository used to store initial environment information.
[0033] In this embodiment of the invention, based on the acquired target ACE state data, target parameters, and target environment parameters, corresponding initial environment information is generated, and the acquired initial environment information is input into a preset information pool for storage.
[0034] Step 103: Extract a first preset amount of first environmental information from the preset information pool, and combine it with the initial environmental information to generate a second preset amount of target environmental information.
[0035] The first preset quantity refers to the quantity of first environmental information extracted from the information pool. The specific quantity is set according to the requirements.
[0036] The second preset quantity refers to the quantity generated by adding the first environmental information of the first preset quantity to the current initial environmental information. In other words, if the first preset quantity is n, then the second preset quantity is n+1.
[0037] In this embodiment of the invention, a first preset number of first environmental information is extracted from a preset information pool, and a second preset number of target environmental information is generated by combining the currently generated initial environmental information and the associated environmental information.
[0038] Step 104: Update all initial matrices using the target environment information through the preset DQL model to generate multiple corresponding target matrices.
[0039] DQL model refers to a reinforcement learning model used to update an initial matrix.
[0040] In this embodiment of the invention, the acquired target environment information is used to update all the initial matrices using a preset DQL model, and the updated initial matrices are used to generate the corresponding target matrix.
[0041] Step 105: Use a random greedy strategy to select the associated second target action from each target matrix, and determine the corresponding target power command difference based on the power difference associated with each second target action.
[0042] A stochastic greedy policy refers to a stochastic policy in reinforcement learning, meaning that the probability of selecting the action that maximizes the action-value function is... The probabilities of other actions are equal, all being 1 / 2. .
[0043] The second target action refers to the action selected from the target matrix, with each target matrix corresponding to one target action. Each target action corresponds to a power value in the target action set data, and this power value is used as the power difference.
[0044] In this embodiment of the invention, a random greedy strategy is used to select associated second target actions from each target matrix. It is worth mentioning that each target matrix corresponds to a second target action, and each second target action corresponds to a power value in the target action set data. This power value is used as a power difference to generate multiple power differences. All power differences are summed to generate the corresponding target power command difference.
[0045] Step 106: Based on the calculation result of the difference between the initial control power command value and the target power command value, control the output power of the generator set in the new energy power system to be controlled.
[0046] In this embodiment of the invention, the difference between the initial control power command value and the target power command is summed to generate the corresponding target total power, and the output power of the generator set in the new energy power system to be controlled is controlled. The output of the generator set is adjusted according to a preset gradient until the output power of the generator set is equal to the target total power.
[0047] In this embodiment of the invention, in response to the received trigger information, a model controller within the new energy power system to be controlled is invoked to collect target ACE state data, initial control power command values, target parameters, and target environmental parameters corresponding to the new energy power system to be controlled. Using the target ACE state data, target parameters, and target environmental parameters, corresponding initial environmental information is determined and stored in a preset information pool. A first preset amount of first environmental information is extracted from the preset information pool, and combined with the initial environmental information, a second preset amount of target environmental information is generated. A preset DQL model is used to update all initial matrices using the target environmental information, generating multiple corresponding target matrices. A random greedy strategy is employed to distribute the target matrix. The method selects associated second target actions from each target matrix and determines the corresponding target power command difference based on the power difference associated with each second target action. Based on the calculation results of the initial control power command value and the target power command difference, it controls the output power of the generator units within the new energy power system to be controlled. This solves the technical problem that traditional centralized automatic generation control modes cannot meet the development and operation conditions of the power grid, and that existing frequency regulation for large-scale new energy access to the power system has difficulty maintaining frequency stability in various areas of the distributed power grid under highly random load conditions. It achieves frequency stability in various areas of the distributed power grid under highly random load conditions and improves the distributed power grid's ability to absorb intermittent energy sources such as wind power and photovoltaics.
[0048] Please see Figure 2 , Figure 2 The flowchart illustrates the steps of a frequency control method based on model-based reinforcement learning, as provided in Embodiment 2 of the present invention.
[0049] This invention provides a frequency control method based on model-based reinforcement learning, comprising: Step 201: In response to the received trigger information, call the model controller in the new energy power system to be controlled to collect the target ACE status data, initial control power command value, target parameters and target environmental parameters corresponding to the new energy power system to be controlled.
[0050] like Figure 3 As shown, Figure 3 This is a schematic diagram of the new energy power system to be controlled in an embodiment of the present invention.
[0051] It is worth mentioning that the model controller is applicable to multi-regional distributed power grids. The new energy power system to be controlled includes the model controller and generator units. Among them, the regional power tie line power deviation Δ P tie Load power Δ P L and the frequency deviation Δ of power grid information feedback f. T gLet be the time delay constant of the governor of the thermal power unit. T t The time constant of the thermal power unit. T 1. T 2. T 3. T 4 and T s The time constant of the superconducting magnetic energy storage unit. K s This represents the gain coefficient of the superconducting magnetic energy storage unit. T p Let be the time constant of the frequency response function. K p These are the coefficients of the frequency response function. T 12 Here, is the tie-line time constant. ACE is the area control deviation, and P is the power. f The frequency is [not specified]. It's worth noting that the Dyna-DQL controller is a model-based controller.
[0052] like Figure 4 As shown, Figure 4 A schematic diagram of the process for generating the target total power.
[0053] The model controller is given a load disturbance as input, and then the parameters are initialized using two identical initial matrices. Initial matrix State transition matrix P, target action set data A, learning rate α Adjust parameters β Discount factor γ, Quantity of environmental information N ; Figure 4 The ACE (Entity Control Parameter) is used to update the initial matrix. A random-greedy strategy is applied to select actions from the action set A. The selected actions are used to calculate the total power command output, ultimately providing the specific power generation value of the unit to maintain grid frequency stability and the Control Performance Standard (CPS) index. Simultaneously, collected environmental information is stored in an experience pool. Real-time feedback and experience pool data are used together to update the initial matrix, controlling the controller's output. The experience pool is a collection of historical data, a two-dimensional matrix of system parameters (first environmental parameter, first target action, target reward value, second environmental parameter). The amount of historical data extracted affects the accuracy of Q-value updates.
[0054] After initializing the parameters, our first step is to obtain a real environment parameter. ,based on Randomly select one of two initial matrices of the same type, and use a random-greedy strategy to select the first target action from the action set A. aAnd the target reward value r, followed by environmental information ( , ,r, The information is stored in the information pool, and N pieces of information are extracted from the information pool. In addition to the real-time information, there are a total of N+1 pieces of information. The two identical Q matrices are updated a total of N+1 times using the DQL formula. Then, the total power command is calculated based on the selected action and output to the unit. Finally, the next iteration begins.
[0055] Further, step 201 may include the following sub-steps: S11. In response to the received trigger information, call the model controller in the new energy power system to be controlled to collect the initial ACE state data, initial ACE control power command value, target parameters and target environmental parameters corresponding to the new energy power system to be controlled.
[0056] Initial ACE status data refers to tie-line switching power, primary frequency regulation coefficient, and grid information feedback frequency deviation.
[0057] In this embodiment of the invention, upon receiving information about frequency modulation for a new energy power system, the system calls upon the model controller within the new energy power system to collect the initial ACE state data, initial control power command value, target parameters, and target environmental parameters corresponding to the new energy power system to be controlled.
[0058] S12. Input the initial ACE state data into the preset ACE state model and output the corresponding target ACE state data.
[0059] In practical implementation, to facilitate the method's implementation, the above process can be converted into a formulaic encapsulation, where the preset ACE state model can be as follows:
[0060] In the formula, This represents the target ACE state data for the i-th time. This represents the power exchanged on the i-th tie line. This represents the frequency modulation coefficient for the i-th time. This represents the frequency deviation of the i-th power grid information feedback.
[0061] In this embodiment of the invention, the tie-line switching power, primary frequency regulation coefficient and grid information feedback frequency deviation are used as inputs to a preset ACE state model, and the corresponding target ACE state data are output.
[0062] Step 202: Using the target ACE state data, target parameters, and target environment parameters, determine the corresponding initial environment information and store the initial environment information in a preset information pool.
[0063] Furthermore, the target environmental parameters include a first environmental parameter and a second environmental parameter, and step 202 may include the following sub-steps: S21. Based on the second environment parameter, randomly select any initial matrix from the target parameters.
[0064] It is worth mentioning that the objective parameters contain two initial matrices of the same type. and initial matrix .
[0065] In this embodiment of the invention, based on the obtained real second environmental parameters In two initial matrices of the same type and initial matrix An initial matrix is randomly selected from the given matrix.
[0066] S22. Select the corresponding first target action from the target action set data within the target parameters using a random greedy strategy.
[0067] Random greedy strategy refers to The greedy strategy balances exploitation and exploration, selecting the portion of the action-value function that maximizes exploitation, while other non-optimal actions still have a probability of being considered for exploration. Its specific formula is:
[0068] In the formula, Indicated in the second environmental parameter The probability of choosing action 'a'; ε represents the random factor; A represents the target action set data; Indicated in the second environmental parameter Action-state value of 'a'.
[0069] In this embodiment of the invention, a first target action is selected from the target action set data within the target parameters using a random greedy strategy.
[0070] S23. Execute the first target action using the target ACE state data, and generate the corresponding target reward value through the preset ACE state model.
[0071] It is worth mentioning that, The target reward value.
[0072] In this embodiment of the invention, the first target action is performed using target ACE state data, and the corresponding target reward value is generated through a preset ACE state model.
[0073] S24. Using the first environmental parameters, the first target action, the target reward value, and the second environmental parameters, generate the corresponding initial environmental information.
[0074] Initial environment information refers to the information derived from the first environment parameters. First target action Target reward value Second environmental parameters Composition of environmental information, initial environmental information ( , , , ).
[0075] In this embodiment of the invention, the first environmental parameter, the first target action, the target reward value, and the second environmental parameter are used to generate the corresponding initial environmental information.
[0076] S25. Store the initial environment information through a preset information pool.
[0077] In this embodiment of the invention, the acquired initial environmental information is input into a preset information pool for storage.
[0078] Step 203: Extract a first preset amount of first environmental information from the preset information pool, and combine it with the initial environmental information to generate a second preset amount of target environmental information.
[0079] Furthermore, step 203 may include the following sub-steps: S31. Extract a first preset amount of first environmental information from the information pool.
[0080] In this embodiment of the invention, a first preset amount of first environmental information is extracted from the information pool, such as... Figure 4 As shown, N pieces of first environmental information are extracted.
[0081] S32. Using the initial environmental information and the first preset amount of first environmental information, generate a second preset amount of target environmental information.
[0082] In this embodiment of the invention, the initial environmental information generated in S24 is integrated with the first environmental information extracted from the information pool in a first preset quantity to obtain N+1 target environmental information items, and the second preset quantity is N+1.
[0083] Step 204: Update all initial matrices using the target environment information through the preset DQL model to generate multiple corresponding target matrices.
[0084] In this embodiment of the invention, the target environment information is used to update all initial matrices using a preset DQL model. The number of updates is a second preset number, which is N+1 random matrix selection updates.
[0085] In practical implementation, to facilitate the method's implementation, the above process can be converted into a formulaic encapsulation. The pre-defined DQL model can be as follows:
[0086] In the formula, Q A Q B Let be an initial Q-matrix of the same type, s be the current environment state, and s' be the environment state for the next iteration. a Let α be the target action and α be the learning rate. This is the discount factor.
[0087] Step 205: Use a random greedy strategy to select the associated second target action from each target matrix, and determine the corresponding target power command difference based on the power difference associated with each second target action.
[0088] Furthermore, step 205 may include the following sub-steps: S41. Using a random greedy strategy, select associated target actions from each target matrix to generate multiple second target actions.
[0089] In this embodiment of the invention, a random greedy strategy is used to select associated second target actions from each target matrix. It is worth mentioning that each target matrix corresponds to one second target action.
[0090] S42. Use multiple second target actions to match the associated power difference from the target action set data to generate multiple initial power command differences.
[0091] In this embodiment of the invention, each second target action corresponds to a power value in the target action set data, and this power value is used as a power difference to generate multiple power differences.
[0092] S43. Calculate the sum of multiple initial power command differences to generate the target power command difference.
[0093] In this embodiment of the invention, all power differences are summed to generate the corresponding target power command difference.
[0094] Step 206: Based on the calculation result of the difference between the initial control power command value and the target power command value, control the output power of the generator set in the new energy power system to be controlled.
[0095] Furthermore, step 206 may include the following sub-steps: S51. Calculate the sum of the difference between the initial control power command value and the target power command value to generate the target total power command value.
[0096] In this embodiment of the invention, the sum of the difference between the initial control power command value and the target power command value is calculated to generate the target total power command value.
[0097] S52, the target total power is associated with the target total power command value output by the generator set in the new energy power system to be controlled.
[0098] In this embodiment of the invention, the target total power is associated with the target total power command value output by the generator set in the new energy power system to be controlled.
[0099] It is worth mentioning that the target total power is ΔP .
[0100] Step 207: Obtain the frequency deviation data and tie-line power deviation data of the new energy power system to be controlled at the current moment.
[0101] In this embodiment of the invention, frequency deviation data and tie-line power deviation data of the new energy power system to be controlled at the current moment are obtained.
[0102] Step 208: Use the current target total power as the new initial control power command value.
[0103] In this embodiment of the invention, the target total power at the current moment is used as the new initial control power command value.
[0104] Step 209: Based on the frequency deviation data and tie-line power deviation data, determine the new target ACE state data, jump to the step of using the target ACE state data, target parameters and target environmental parameters to determine the corresponding initial environmental information and store the initial environmental information through a preset information pool.
[0105] In this embodiment of the invention, based on frequency deviation data and tie-line power deviation data, new target ACE state data is determined, and the process proceeds to step 202.
[0106] Please see Figure 5 , Figure 5 For thermal power turbine units and superconducting magnetic energy storage units that take into account generator output rate control (GRC), the superconducting magnetic energy storage unit uses a frequency deviation Δ f As input, the unit outputs power Δ based on the input power. P s T 1, T 2, T 3, T4 is the unit time constant, T gThe time delay constant of the generator unit is denoted as . The thermal power generating unit includes a governor module, a GRC module, a turbine module, and a power limiting module. These four modules are used to simulate the actual operating conditions of a reheat turbine unit. Regional power tie lines are AC power tie lines between different regions. When there is an imbalance between power generation and consumption in one region, power from other regions is drawn through these tie lines to maintain the stability of the entire power grid. The response time and capacity of the superconducting magnetic energy storage unit are superior to those of traditional units. To verify the controller's control capability, the load is a randomly generated load signal using MATLAB / Simulink.
[0107] Please see Figures 6-8 The model controller will be compared with existing Dyna-QL controllers and Q controllers.
[0108] Please see Figure 6 , Figure 6 To ensure that the average absolute value of the regional frequency deviation can always be kept within 0.0005 Hz under sinusoidal disturbances, the Dyna-QL controller achieves 0.0011 Hz and the Q controller achieves 0.0025 Hz.
[0109] Please see Figure 7 , Figure 7 To ensure that the average absolute value of ACE can always be maintained within 0.2 MW under sinusoidal disturbances, the Dyna-DQL-based controller achieves 0.4 MW, while the Q controller achieves 8 MW.
[0110] Please see Figure 8 , Figure 8 Under sinusoidal disturbances, the controller proposed based on Dyna-DQL can always maintain the 10-minute CPS average above 200%. The Dyna-QL controller is slightly weaker, with the Q controller achieving 199.7%.
[0111] In this embodiment of the invention, in response to the received trigger information, a model controller within the new energy power system to be controlled is invoked to collect target ACE state data, initial control power command values, target parameters, and target environmental parameters corresponding to the new energy power system to be controlled. Using the target ACE state data, target parameters, and target environmental parameters, corresponding initial environmental information is determined and stored in a preset information pool. A first preset amount of first environmental information is extracted from the preset information pool, and combined with the initial environmental information, a second preset amount of target environmental information is generated. A preset DQL model is used to update all initial matrices using the target environmental information, generating multiple corresponding target matrices. A random greedy strategy is employed to distribute the target matrix. The method selects associated second target actions from each target matrix and determines the corresponding target power command difference based on the power difference associated with each second target action. Based on the calculation results of the initial control power command value and the target power command difference, it controls the output power of the generator units within the new energy power system to be controlled. This solves the technical problem that traditional centralized automatic generation control modes cannot meet the development and operation conditions of the power grid, and that existing frequency regulation for large-scale new energy access to the power system has difficulty maintaining frequency stability in various areas of the distributed power grid under highly random load conditions. It achieves frequency stability in various areas of the distributed power grid under highly random load conditions and improves the distributed power grid's ability to absorb intermittent energy sources such as wind power and photovoltaics.
[0112] Please see Figure 9 , Figure 9 This is a structural block diagram of a frequency control device based on model-based reinforcement learning provided in Embodiment 3 of the present invention.
[0113] This invention provides a frequency control device based on model-based reinforcement learning, comprising: The response module 301 is used to respond to the received trigger information and call the model controller in the new energy power system to be controlled to collect the target ACE status data, initial control power command value, target parameters and target environmental parameters corresponding to the new energy power system to be controlled.
[0114] The initial environment information module 302 is used to determine the corresponding initial environment information by using the target ACE state data, target parameters and target environment parameters, and store the initial environment information in a preset information pool.
[0115] The target environment information module 303 is used to extract a first preset amount of first environment information from a preset information pool, and combine it with the initial environment information to generate a second preset amount of target environment information.
[0116] The target matrix module 304 is used to update all initial matrices using target environment information through a preset DQL model, generating multiple corresponding target matrices.
[0117] The target power command difference module 305 is used to select the associated second target action from each target matrix using a random greedy strategy, and determine the corresponding target power command difference based on the power difference associated with each second target action.
[0118] The control module 306 is used to control the output power of the generator set in the new energy power system to be controlled based on the calculation result of the difference between the initial control power command value and the target power command value.
[0119] Furthermore, the response module 301 includes: The parameter acquisition submodule is used to respond to the received trigger information and call the model controller in the new energy power system to be controlled to collect the initial ACE state data, initial ACE control power command value, target parameters and target environmental parameters corresponding to the new energy power system to be controlled.
[0120] The ACE state model submodule is used to input the initial ACE state data into the preset ACE state model and output the corresponding target ACE state data.
[0121] Furthermore, the target environmental parameters include a first environmental parameter and a second environmental parameter, and the initial environmental information module 302 includes: The initial matrix selection submodule is used to randomly select any initial matrix from the target parameters based on the second environment parameter.
[0122] The first target action submodule is used to select the corresponding first target action from the target action set data within the target parameters using a random greedy strategy.
[0123] The target reward value submodule is used to execute the first target action using the target ACE state data and generate the corresponding target reward value through a preset ACE state model.
[0124] The initial environment information submodule is used to generate corresponding initial environment information using the first environment parameters, the first target action, the target reward value, and the second environment parameters.
[0125] The storage module is used to store the initial environment information through a preset information pool.
[0126] Furthermore, the target environment information module 303 includes: The first environmental information extraction submodule is used to extract a first preset amount of first environmental information from the information pool.
[0127] The target environment information generation submodule is used to generate a second preset number of target environment information by using the initial environment information and a first preset number of first environment information.
[0128] Furthermore, the target power command difference module 305 includes: The second target action generation submodule is used to select associated target actions from each target matrix using a random greedy strategy to generate multiple second target actions.
[0129] The initial power command difference generation submodule is used to generate multiple initial power command differences by matching associated power differences from target action set data using multiple second target actions.
[0130] The target power command difference generation submodule is used to calculate the sum of multiple initial power command differences and generate the target power command difference.
[0131] Furthermore, the control module 306 includes: The target total power command value generation submodule is used to calculate the sum of the difference between the initial control power command value and the target power command value, and generate the target total power command value.
[0132] The target total power submodule is used to associate the target total power with the output target total power command value of the generator set in the new energy power system to be controlled.
[0133] Furthermore, it also includes: The deviation data acquisition module is used to acquire the frequency deviation data and tie-line power deviation data of the new energy power system to be controlled at the current moment.
[0134] The update module is used to take the target total power at the current moment as the new initial control power command value.
[0135] The jump module is used to determine new target ACE state data based on frequency deviation data and tie-line power deviation data, jump to execute the steps of using target ACE state data, target parameters and target environmental parameters to determine the corresponding initial environmental information and store the initial environmental information through a preset information pool.
[0136] In this embodiment of the invention, in response to the received trigger information, a model controller within the new energy power system to be controlled is invoked to collect target ACE state data, initial control power command values, target parameters, and target environmental parameters corresponding to the new energy power system to be controlled. Using the target ACE state data, target parameters, and target environmental parameters, corresponding initial environmental information is determined and stored in a preset information pool. A first preset amount of first environmental information is extracted from the preset information pool, and combined with the initial environmental information, a second preset amount of target environmental information is generated. A preset DQL model is used to update all initial matrices using the target environmental information, generating multiple corresponding target matrices. A random greedy strategy is employed to distribute the target matrix. The method selects associated second target actions from each target matrix and determines the corresponding target power command difference based on the power difference associated with each second target action. Based on the calculation results of the initial control power command value and the target power command difference, it controls the output power of the generator units within the new energy power system to be controlled. This solves the technical problem that traditional centralized automatic generation control modes cannot meet the development and operation conditions of the power grid, and that existing frequency regulation for large-scale new energy access to the power system has difficulty maintaining frequency stability in various areas of the distributed power grid under highly random load conditions. It achieves frequency stability in various areas of the distributed power grid under highly random load conditions and improves the distributed power grid's ability to absorb intermittent energy sources such as wind power and photovoltaics.
[0137] An electronic device according to an embodiment of the present invention includes: a memory and a processor, wherein the memory stores a computer program; when the computer program is executed by the processor, the processor performs a frequency control method based on model reinforcement learning as described in any of the above embodiments.
[0138] The memory can be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. The memory has storage space for program code used to perform any of the method steps described above. For example, the storage space for program code may include individual program codes for implementing the various steps in the methods described above. This program code can be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact discs (CDs), memory cards, or floppy disks. The program code may be compressed, for example, in a suitable form. When run by a computing processing device, this code causes the computing processing device to perform the various steps in the methods described above.
[0139] This invention provides a computer-readable storage medium storing a computer program thereon, which, when executed, implements a frequency control method based on model-based reinforcement learning as described in any embodiment of this invention.
[0140] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0141] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0142] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0143] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0144] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0145] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A frequency control method based on model-based reinforcement learning, characterized by, include: In response to the received trigger information, the model controller in the new energy power system to be controlled is invoked to collect the target ACE state data, initial control power command value, target parameters and target environmental parameters corresponding to the new energy power system to be controlled. The target parameters contain two identical initial matrices. Using the target ACE state data, the target parameters, and the target environment parameters, the corresponding initial environment information is determined and stored in a preset information pool; Extracting a first preset amount of first environmental information from the preset information pool, and combining it with the initial environmental information to generate a second preset amount of target environmental information, including: Extract a first preset number of N environmental information from the preset information pool, and integrate the generated initial environmental information with the first preset number of N environmental information extracted from the preset information pool to generate a second preset number of N+1 target environmental information. The initial matrices within the target parameters are updated using the target environment information through a preset DQL model, generating multiple corresponding target matrices. A random greedy strategy is used to select associated second target actions from each of the target matrices, and the corresponding target power command difference is determined based on the power difference associated with each second target action. Based on the calculation result of the difference between the initial control power command value and the target power command, the output power of the generator set in the new energy power system to be controlled is controlled.
2. The model-based reinforcement learning based frequency control method of claim 1, wherein, The steps of responding to the received trigger information and invoking the model controller within the new energy power system to be controlled to collect the target ACE state data, initial control power command value, target parameters, and target environmental parameters corresponding to the new energy power system to be controlled include: In response to the received trigger information, the model controller in the new energy power system to be controlled is invoked to collect the initial ACE state data, initial ACE control power command value, target parameters and target environmental parameters corresponding to the new energy power system to be controlled; The initial ACE state data is used as input to a preset ACE state model, and the corresponding target ACE state data is output.
3. The model-based reinforcement learning based frequency control method of claim 2, wherein, The target environmental parameters include a first environmental parameter and a second environmental parameter. The step of determining the corresponding initial environmental information using the target ACE state data, the target parameters, and the target environmental parameters, and storing the initial environmental information in a preset information pool, includes: Based on the second environmental parameters, any one of the initial matrices is randomly selected from the target parameters; The first target action is selected from the target action set data within the target parameters using a random greedy strategy. The first target action is executed using the target ACE state data, and a corresponding target reward value is generated through a preset ACE state model; Using the first environmental parameters, the first target action, the target reward value, and the second environmental parameters, corresponding initial environmental information is generated. The initial environmental information is stored in a preset information pool.
4. The frequency control method based on model-based reinforcement learning according to claim 1, characterized in that, The step of selecting associated second target actions from each of the target matrices using a random greedy strategy, and determining the corresponding target power command difference based on the power difference associated with each of the second target actions, includes: A random greedy strategy is used to select associated target actions from each of the target matrices to generate multiple second target actions; Multiple initial power command differences are generated by matching associated power differences from the target action set data using multiple second target actions; The sum of the multiple initial power command differences is calculated to generate the target power command difference.
5. The frequency control method based on model-based reinforcement learning according to claim 1, characterized in that, The step of controlling the output power of the generator sets in the new energy power system to be controlled based on the calculation result of the difference between the initial control power command value and the target power command includes: Calculate the sum of the difference between the initial control power command value and the target power command value to generate the target total power command value; The target total power is associated with the output of the target total power command value by the generator set within the new energy power system to be controlled.
6. The frequency control method based on model-based reinforcement learning according to claim 1, characterized in that, Also includes: Obtain the frequency deviation data and tie-line power deviation data of the new energy power system to be controlled at the current moment; Use the target total power at the current moment as the new initial control power command value; Based on the frequency deviation data and the tie-line power deviation data, new target ACE state data is determined, and the process jumps to execute the step of using the target ACE state data, the target parameters, and the target environmental parameters to determine the corresponding initial environmental information and store the initial environmental information in a preset information pool.
7. A frequency control device based on model-based reinforcement learning, characterized by, include: The response module is used to respond to the received trigger information and call the model controller in the new energy power system to be controlled to collect the target ACE state data, initial control power command value, target parameters and target environmental parameters corresponding to the new energy power system to be controlled. The target parameters contain two identical initial matrices. The initial environment information module is used to determine the corresponding initial environment information using the target ACE state data, the target parameters, and the target environment parameters, and to store the initial environment information in a preset information pool. The target environment information module is used to extract a first preset amount of first environment information from the preset information pool, and combine it with the initial environment information to generate a second preset amount of target environment information, including: Extract a first preset number of N environmental information from the preset information pool, and integrate the generated initial environmental information with the first preset number of N environmental information extracted from the preset information pool to generate a second preset number of N+1 target environmental information. The target matrix module is used to update all the initial matrices in the target parameters using the target environment information through a preset DQL model, thereby generating multiple corresponding target matrices. The target power command difference module is used to select associated second target actions from each of the target matrices using a random greedy strategy, and determine the corresponding target power command difference based on the power difference associated with each second target action. The control module is used to control the output power of the generator set in the new energy power system to be controlled based on the calculation result of the difference between the initial control power command value and the target power command value.
8. An electronic device, comprising: The system includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor causes the processor to perform the steps of the frequency control method based on model-based reinforcement learning as described in any one of claims 1-6.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed, it implements the frequency control method based on model-based reinforcement learning as described in any one of claims 1-6.