Reinforcement learning based multi-energy power system load frequency collaborative control method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHWESTERN POLYTECHNICAL UNIV
- Filing Date
- 2026-04-03
- Publication Date
- 2026-08-07
AI Technical Summary
[0007]本发明的目的是提供一种通信约束下基于强化学习的多能源电力系统负荷频率协同控制方法,以解决异质调频资源在负荷频率控制中协同利用不足、功率分配缺乏自适应性以及通信受限条件下控制信号更新冗余的问题,从而实现多能源电力系统频率稳定控制,降低通信与控制资源消耗,提高系统协调调频性能
本发明实现了通信约束下基于强化学习的多能源电力系统负荷频率协同控制,通过在传统负荷频率控制模型中引入风力发电单元、光伏发电单元及储能系统等多种异质调频资源,实现了多能源电力系统统一负荷频率控制建模,使不同调频资源能够协同参与系统频率调节,从而提升系统整体调频能力和可再生能源参与频率调节的灵活性。通过构建基于强化学习的功率分配优化机制,设计包含系统频率稳定、风电和光伏可用调节容量约束以及储能功率变化平滑性的综合奖励函数,使强化学习算法能够根据系统运行状态自适应调整各调频资源的功率分配系数,从而提高风电和光伏等可再生能源调频能力利用率,并抑制储能系统功率的剧烈波动,提高多能源调频资源的协调控制能力。此外,带记忆感知的触发机制通过综合利用当前系统输出信息与历史触发时刻输出信息构造触发条件,同时引入时变记忆长度来确定触发条件所涉及的历史状态数量。从而降低通信资源消耗、减少控制器更新频率并降低系统计算开销,提高网络化电力系统的控制效率和运行可靠性。通过强化学习功率分配优化机制与带记忆触发机制控制策略的协同设计,本发明能够在通信资源受限条件下实现多能源调频资源的高效协同控制,在保证系统频率稳定的同时提升资源利用效率和控制性能,拓展了负荷频率控制方法在通信受限电力系统中的应用范围,并为多能源电力系统的实际工程应用提供了一种有效的控制方案。
Smart Images

Figure CN122532987A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of automatic generation control technology of power systems, and relates to a method for coordinated load frequency control of multi-energy power systems based on reinforcement learning under communication constraints. Background Technology
[0002] With the widespread integration of renewable energy and energy storage systems into power systems, the types of resources involved in frequency regulation are becoming increasingly diversified. Load frequency control is no longer limited to traditional generating units, but has gradually evolved into a multi-energy coordinated frequency regulation problem involving synchronous generator units, wind power units, photovoltaic power units, and energy storage units. The document points out that different frequency regulation resources differ significantly in terms of available regulation capacity, dynamic response characteristics, and operational constraints, making their coordinated participation in load frequency control more complex. Therefore, it is necessary to construct an effective coordination mechanism to improve the utilization efficiency of heterogeneous frequency regulation resources and the overall frequency regulation capability of the system.
[0003] Existing multi-resource load frequency control methods typically achieve frequency regulation task sharing by setting fixed power allocation coefficients. However, these methods are difficult to adapt to changes in system operating conditions and fail to reflect the dynamic differences in the operating characteristics and available capacity of different frequency regulation resources. This can easily lead to underutilization of the regulation capacity of some resources or uneven resource participation, thereby affecting the system's frequency control performance. To address this issue, the document explicitly points out the need to develop a power allocation scheme with adaptive decision-making capabilities, in order to dynamically adjust the participation level of various frequency regulation resources according to the operating status.
[0004] On the other hand, reinforcement learning, due to its strong adaptive decision-making capabilities, has been gradually introduced into the field of load frequency control to improve system performance. However, existing reinforcement learning methods have focused more on controller parameter optimization and have not been fully applied to power allocation coordination among heterogeneous frequency regulation resources, thus making it difficult to achieve efficient and balanced coordinated frequency regulation among different resources. The document further points out that in heterogeneous frequency regulation resource scenarios, it is necessary to incorporate the changes in available regulation capacity and differences in operating characteristics of different resources into the power allocation optimization framework, and improve coordination capabilities by dynamically adjusting the power allocation coefficients.
[0005] Furthermore, in networked load frequency control systems, control signals and status information typically rely on communication networks for transmission. Using a fixed-period transmission method can easily generate a large amount of redundant data updates, increasing the communication burden and potentially reducing control efficiency. The document points out that event-triggered mechanisms can reduce unnecessary data transmission while maintaining control performance, and are therefore widely used in load frequency control under communication-constrained conditions; however, most existing triggering mechanisms primarily make triggering decisions based on the current system state, making it difficult to fully reflect dynamic changes in the system, and potentially leading to inaccurate trigger judgments when the system state fluctuates significantly. Therefore, a memory-triggered mechanism capable of incorporating historical state information is needed to further improve control accuracy and communication efficiency.
[0006] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0007] The purpose of this invention is to provide a load frequency coordinated control method for multi-energy power systems based on reinforcement learning under communication constraints. This method addresses the problems of insufficient coordinated utilization of heterogeneous frequency regulation resources in load frequency control, lack of adaptive power allocation, and redundancy in control signal updates under communication-limited conditions. As a result, it achieves stable frequency control of multi-energy power systems, reduces the consumption of communication and control resources, and improves the coordinated frequency regulation performance of the system.
[0008] Other features and advantages of the invention will become apparent from the following detailed description, or may be learned in part by practice of the invention.
[0009] According to a first aspect of the present invention, a method for coordinated load frequency control of a multi-energy power system based on reinforcement learning is provided, the method comprising: Step 1: Establish a load frequency control model for a multi-regional interconnected power system, dividing the multi-regional interconnected power system into multiple interconnected control regions; wherein each control region includes synchronous generator sets, wind power generation units, photovoltaic power generation units, and energy storage units, which serve as heterogeneous frequency regulation resources participating in frequency regulation; Step 2: Design a power allocation optimization mechanism based on reinforcement learning. Based on the dynamic changes in the available adjustment capacity of different frequency modulation resources and the differences in frequency adjustment characteristics, optimize the power allocation coefficient of each frequency modulation resource using reinforcement learning methods. Step 3: Design a load frequency controller with a memory trigger mechanism. Construct trigger conditions by combining the output information of the current control area with the output information of the historical trigger time of the control area, determine whether the control signal is updated, and generate control inputs that act on the load frequency control model based on the trigger result.
[0010] In some exemplary embodiments, step 1.1: Define the first i State vector, output vector, and system disturbance for each control region:
[0011]
[0012]
[0013] in, Let be the system state vector. Indicates the output vector. For system disturbance; Step 1.2: Establish the first i State-space equations of a regional load frequency control model:
[0014]
[0015] ,
[0016] in, and They represent the first Region and the Frequency deviation in the region and These represent the turbine output power deviation and the governor output power deviation, respectively. as well as These represent the incremental regulation power provided by the wind turbine generator, the photovoltaic power generation unit, and the energy storage system, respectively. and These represent tie line power deviation and load disturbance, respectively. Indicates the area control error. and These represent the time constants of the steam turbine and the governor, respectively. This is the frequency offset coefficient. This is the speed adjustment coefficient. Indicates the load damping coefficient. Represents the inertia constant, parameter Characterizes the power exchange characteristics of the interconnected areas. and These represent the dynamic time constants of the wind turbine generator, the photovoltaic power generation unit, and the energy storage system, respectively. and Let represent the power allocation coefficients for different frequency modulation resources, and satisfy . .
[0017] In some exemplary embodiments, step 2 specifically includes: Step 2.1: Construct a reinforcement learning power allocation optimization framework. Use the multi-energy power system load frequency control model as the reinforcement learning environment. Use the system frequency deviation, the output power of each frequency regulation resource, and the actual available power of each frequency regulation resource as reinforcement learning state variables. Use the power allocation coefficient of each frequency regulation resource as the action space of the reinforcement learning agent. Step 2.2: Construct a comprehensive reward function that integrates frequency stability reward, wind power frequency regulation utilization reward, photovoltaic frequency regulation utilization reward, and energy storage power smoothing reward. Step 2.3: Train the power allocation strategy using the near-end strategy optimization algorithm, calculate the advantage function of the state transition samples, construct the truncated surrogate objective function to update the actor network parameters, update the critic network parameters according to the temporal difference error, and iteratively output the optimal power allocation coefficient.
[0018] In some exemplary embodiments, the frequency stability reward item specifically includes:
[0019] in, and These are the weighting coefficients. This indicates the frequency deviation threshold.
[0020] In some exemplary embodiments, the wind power frequency regulation utilization incentive specifically includes:
[0021] in, , This indicates the available power of the wind turbine. and These are the weighting coefficients.
[0022] In some exemplary embodiments, the photovoltaic frequency regulation utilization incentive specifically includes:
[0023] in, , This indicates the available power of the wind turbine. and These are the weighting coefficients.
[0024] In some exemplary embodiments, the energy storage power smoothing reward item specifically includes: , in, This indicates the current regulation power of the energy storage system. This indicates the most recently sent power command, parameters. It is a weighting coefficient.
[0025] In some exemplary embodiments, step 3 specifically includes: Step 3.1, Select the area control error Design a PI-type load frequency controller for the control signal, with the following expression:
[0026] in, This is the gain matrix; Step 3.2: Construct an adaptive memory-aware event triggering mechanism, combine the current system state with the error of historical triggering times to construct triggering conditions, and determine the control signal update time; Step 3.3: Reconstruct the controller based on the memory trigger mechanism, and generate the current control input by weighted fusion of control signals from historical trigger moments.
[0027] In some exemplary embodiments, the adaptive memory-aware event triggering mechanism specifically includes:
[0028] in, Functions related to event triggering, The table indicates the next trigger time. For a fixed sampling period and satisfying , Indicates from the current trigger time Starting from the beginning, counting backwards... One sampling period; , , and ; This represents the floor function, which rounds its argument up to the nearest integer. express The upper bound, and satisfying , It is a non-zero real constant. , It is a given positive constant. It is a weight parameter, and satisfies , , This indicates the index of the previously transmitted data packet. It is an event triggering matrix, and .
[0029] In some exemplary embodiments, the memory-triggered controller reconstruction specifically includes:
[0030] According to a second aspect of the present invention, a storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the reinforcement learning-based multi-energy power system load frequency coordinated control method described in the first aspect.
[0031] According to a third aspect of the present invention, a computer program product is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, the load frequency coordinated control method for a multi-energy power system based on reinforcement learning described in the first aspect is implemented.
[0032] According to a fourth aspect of the present invention, an electronic device is provided, comprising: Processor; and Memory for storing the executable instructions of the processor; The processor is configured to implement the reinforcement learning-based multi-energy power system load frequency coordinated control method described in the first aspect by executing the executable instructions.
[0033] Compared with the prior art, the beneficial effects of the reinforcement learning-based multi-energy power system load frequency cooperative control method provided by the embodiments of the present invention are: This invention realizes coordinated load frequency control of a multi-energy power system under communication constraints based on reinforcement learning. By introducing various heterogeneous frequency regulation resources, such as wind power units, photovoltaic power units, and energy storage systems, into the traditional load frequency control model, it achieves unified load frequency control modeling for the multi-energy power system. This allows different frequency regulation resources to participate collaboratively in system frequency regulation, thereby improving the overall frequency regulation capability of the system and the flexibility of renewable energy participation in frequency regulation. By constructing a power allocation optimization mechanism based on reinforcement learning, and designing a comprehensive reward function that includes system frequency stability, available adjustable capacity constraints for wind and photovoltaic power, and the smoothness of energy storage power changes, the reinforcement learning algorithm can adaptively adjust the power allocation coefficients of each frequency regulation resource according to the system operating state. This improves the utilization rate of frequency regulation capabilities of renewable energy sources such as wind and photovoltaic power, suppresses drastic power fluctuations in the energy storage system, and enhances the coordinated control capability of multi-energy frequency regulation resources. Furthermore, a memory-aware triggering mechanism constructs trigger conditions by comprehensively utilizing current system output information and historical trigger moment output information, while introducing time-varying memory length to determine the number of historical states involved in the trigger conditions. This reduces communication resource consumption, controller update frequency, and system computational overhead, improving the control efficiency and operational reliability of the networked power system. By combining the reinforcement learning power allocation optimization mechanism with the control strategy with memory triggering mechanism, this invention can achieve efficient collaborative control of multi-energy frequency regulation resources under conditions of limited communication resources. While ensuring system frequency stability, it improves resource utilization efficiency and control performance, expands the application scope of load frequency control methods in communication-constrained power systems, and provides an effective control scheme for practical engineering applications of multi-energy power systems.
[0034] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description
[0035] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention. It is obvious that the drawings described below are merely some embodiments of the invention, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0036] Figure 1 This is the dynamic model of the i-th region in the load-frequency coordinated control scheme of a multi-energy power system. Figure 2 This is a diagram of the three-region power system structure; Figure 3 The diagram shows the collaborative optimization control framework based on the proximal policy optimization algorithm. Figure 4To reinforce the trend of reward function value changes during the learning and training process; Figure 5 A comparison of the frequency regulation power of the fan under different control schemes; Figure 6 A comparison of photovoltaic frequency regulation power under different control schemes; Figure 7 A comparison of energy storage frequency regulation power under different control schemes; Figure 8 A schematic diagram of load disturbances in the three-region power system; Figure 9 This is a schematic diagram of the frequency response of the three-region power system; Figure 10 This is a schematic diagram showing the triggering times of the three-region power system. Detailed Implementation
[0037] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the invention will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0038] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0039] To address the shortcomings and deficiencies of existing technologies, this exemplary embodiment provides a reinforcement learning-based method for coordinated load frequency control of multi-energy power systems under constrained communication networks. This includes: constructing a multi-energy power system load frequency control model comprising synchronous generator units, wind power generation units, photovoltaic power generation units, and energy storage systems, achieving unified modeling and coordinated frequency regulation of heterogeneous frequency regulation resources; designing a reinforcement learning-based power allocation optimization mechanism, which, by constructing a comprehensive reward function incorporating system frequency stability, available adjustable capacity constraints for wind and photovoltaic power, and the smoothness of energy storage power changes, enables the reinforcement learning algorithm to adaptively adjust the power allocation coefficients of each frequency regulation resource according to the system operating state, thereby improving the utilization rate of renewable energy frequency regulation and reducing energy storage power fluctuations; and designing a load frequency controller with a memory trigger mechanism in a constrained communication network environment, constructing trigger conditions by comprehensively utilizing current system output information and historical trigger moment output information, achieving on-demand updates of control signals, thereby reducing communication resource consumption and control computation overhead. This invention achieves adaptive optimization of power allocation and on-demand updates of control signals in load frequency control of multi-energy power systems under constrained communication networks, improving the system's coordinated frequency regulation capability and reducing communication and control overhead.
[0040] refer to Figure 1 As shown, the specific steps may include: Step 1: Establish a load frequency control model for a multi-energy power system that includes various heterogeneous frequency regulation resources such as synchronous generator sets, wind power generation units, photovoltaic power generation units, and energy storage units.
[0041] The multi-region interconnected power system is divided into multiple interconnected control regions, and the first... i State vector, output vector, and system disturbance for each control region: , .
[0042] Let be the system state vector. Indicates the output vector. For system disturbance, To control the input.
[0043] Establish the first i State-space equations of a regional load frequency control model:
[0044] in,
[0045] , .
[0046] in, and They represent the first Region and the Frequency deviation in the region. and These represent the turbine output power deviation and the governor output power deviation, respectively. as well as These represent the incremental adjustment power provided by the wind turbine generator, the photovoltaic power generation unit, and the energy storage system, respectively. and These represent tie line power deviation and load disturbance, respectively. This indicates the area control error. and These represent the time constants of the steam turbine and the governor, respectively. This is the frequency offset coefficient. This is the speed adjustment coefficient. Indicates the load damping coefficient. Represents the inertia constant. Parameter Characterizing the power exchange characteristics of the interconnected areas, while and These represent the dynamic time constants of the wind turbine generator, photovoltaic power generation unit, and energy storage system, respectively. and Let represent the power allocation coefficients for different frequency modulation resources, and satisfy . .
[0047] Step 2: Design a power allocation optimization mechanism based on reinforcement learning. Based on the dynamic changes in the available adjustment capacity of different frequency modulation resources and the differences in frequency adjustment characteristics, optimize the power allocation coefficient of each frequency modulation resource using reinforcement learning methods.
[0048] A reinforcement learning-based power allocation optimization framework is constructed, using a multi-energy power system load frequency control model as the reinforcement learning environment. System frequency deviation, output power of each frequency regulation resource, and actual available power of each frequency regulation resource are used as reinforcement learning state variables. The power allocation coefficient of each frequency regulation resource is defined as the action space of the reinforcement learning agent. A reinforcement learning reward function is designed, with system frequency stability as the basic objective of power allocation optimization. Specifically, the power system frequency deviation is used as an important indicator for evaluating the system's operating state. By introducing a frequency deviation penalty term into the reward function, the absolute value of the system frequency deviation is cumulatively penalized to reduce frequency fluctuations and improve system stability. Simultaneously, when the system frequency deviation remains within a preset allowable threshold range, a positive incentive term is set in the reward function to encourage the control strategy to maintain the system frequency within the stable operating range, thereby improving load frequency control performance. The reward function is specifically expressed as follows:
[0049] in, and These are the weighting coefficients. This indicates the frequency deviation threshold.
[0050] To improve the frequency regulation efficiency of wind power and photovoltaic (PV) power generation units and reduce wind and solar curtailment, relevant reward functions are constructed for wind power and PV power generation units, linking their frequency regulation power to the available regulation capacity of different power generation units. The reward functions can be expressed as follows:
[0051]
[0052] in, This represents the available power output of wind turbines. Weighting coefficient. and These respectively determine the penalty intensity when wind power utilization is insufficient and when the power of wind power participating in frequency regulation exceeds the available power. Similarly, This indicates the available photovoltaic output power. and It is the weighting coefficient.
[0053] To avoid excessive power fluctuations during frequency regulation, which could accelerate the aging of energy storage devices and reduce their lifespan, a reward function for the energy storage system is constructed to constrain the frequency regulation power variation of the energy storage units. The reward function of the energy storage system can be expressed as: , in, This indicates the current regulation power of the energy storage system. This indicates the most recently sent power command. Parameters It is a weighting coefficient.
[0054] The overall reward function is used to evaluate the comprehensive performance of the current control strategy in terms of system frequency stability, renewable energy utilization rate, and energy storage operation stability. Its expression is: .
[0055] Step 3: Design a load frequency controller with a memory trigger mechanism. Construct trigger conditions by combining the current system output information and the historical trigger time output information, determine whether the control signal is updated, and generate control input based on the trigger result to act on the load frequency control model.
[0056] Select The signal is used as a control signal, and the expression for the controller is:
[0057] in, This is the gain matrix.
[0058] To improve control efficiency and accuracy, the triggering conditions for the adaptive memory sensing event triggering mechanism are as follows:
[0059] in, Functions related to event triggering, The table indicates the next trigger time. For a fixed sampling period and satisfying , Indicates from the current trigger time Starting from the beginning, counting backwards... One sampling period; , , and . This represents the floor function, which rounds its argument up to the nearest integer. express The upper bound, and satisfying . It is a non-zero real constant. It is a given positive constant. It is a weight parameter, and satisfies . This indicates the index of the previously transmitted data packet. It is an event triggering matrix, and .
[0060] Considering the memory-triggered mechanism, the controller is refactored as follows:
[0061] Compared with the prior art, the innovation of this invention is reflected in the following aspects: 1. Solved the problem of load frequency coordinated control in multi-energy power systems under communication constraints; 2. A reinforcement learning reward function design method is proposed, taking into account the frequency regulation characteristics of wind power, photovoltaic power, and energy storage. This method ensures system frequency stability while improving the utilization rate of renewable energy frequency regulation and reducing energy storage power fluctuations, thereby achieving a rational and coordinated allocation of multi-energy frequency regulation resources. 3. A frequency regulation power adaptive allocation method based on reinforcement learning is proposed. This method achieves dynamic optimization of power allocation coefficients for various heterogeneous frequency regulation resources, improving the coordinated frequency regulation capability and resource utilization efficiency of multi-energy power systems. 4. An adaptive memory-sensing triggering mechanism is proposed, which comprehensively utilizes current system output information and historical trigger moment output information. While ensuring system frequency regulation performance, this effectively reduces the number of communication data transmissions, lowers communication resource consumption, and improves the operating efficiency of the control system.
[0062] Implementation Case: The performance of the designed controller is verified through a three-region interconnected power system, and the feasibility of the algorithm is verified through Matlab simulation. The system parameter settings are shown in Table 1: Table 1 Parameter Table
[0063] To verify the effectiveness of the control method proposed in this invention, a three-region multi-energy power system simulation model was constructed for verification. A schematic diagram of the three-region verification model is shown below. Figure 2 As shown. Figure 3 A cooperative optimization control framework based on a proximal policy optimization algorithm is presented to optimize power allocation among various heterogeneous frequency modulation resources. To further verify the advantages of the proposed reinforcement learning-based power allocation optimization method, the frequency modulation performance under different control strategies is compared and analyzed. Figure 4 The changes in the reward function during reinforcement learning training are presented. After about 150 training rounds, the average reward value basically stabilizes. Figure 5 , Figure 6 and Figure 7 The frequency regulation power output of the wind power generation unit, photovoltaic power generation unit, and energy storage system under different control strategies are presented. The frequency regulation power output of each power generation unit can be dynamically adjusted according to changes in operating status, and remains within its available regulation capacity throughout the simulation. Furthermore, the load disturbance situation of the three-region power system is as follows: Figure 8 As shown, load fluctuations during the actual operation of a power system are simulated by setting load changes over different time periods. Figure 9 The frequency response curve of the system under load disturbance is given, and it can be seen that the system frequency can be restored to a steady state under the proposed control strategy. Figure 10 The event triggering times of the three-region power system are shown, and the proposed event triggering mechanism can dynamically adjust the triggering interval according to changes in system state.
[0064] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0065] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the claims.
[0066] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is defined only by the appended claims.
Claims
1. A load frequency coordinated control method for multi-energy power systems based on reinforcement learning, characterized in that, The method includes: Step 1: Establish a load frequency control model for a multi-regional interconnected power system, dividing the multi-regional interconnected power system into multiple interconnected control regions; wherein each control region includes synchronous generator sets, wind power generation units, photovoltaic power generation units, and energy storage units, which serve as heterogeneous frequency regulation resources participating in frequency regulation; Step 2: Design a power allocation optimization mechanism based on reinforcement learning. Based on the dynamic changes in the available adjustment capacity of different frequency modulation resources and the differences in frequency adjustment characteristics, optimize the power allocation coefficient of each frequency modulation resource using reinforcement learning methods. Step 3: Design a load frequency controller with a memory trigger mechanism. Construct trigger conditions by combining the output information of the current control area with the output information of the historical trigger time of the control area, determine whether the control signal is updated, and generate control inputs that act on the load frequency control model based on the trigger result.
2. The multi-energy power system load frequency coordinated control method based on reinforcement learning according to claim 1, characterized in that, The specific content of step 1 includes: Step 1.1: Define the first i State vector, output vector, and system disturbance for each control region: in, Let be the system state vector. Indicates the output vector. For system disturbance; Step 1.2: Establish the first i State-space equations of a regional load frequency control model: , in, and They represent the first Region and the Frequency deviation in the region and These represent the turbine output power deviation and the governor output power deviation, respectively. as well as These represent the incremental regulation power provided by the wind turbine generator, the photovoltaic power generation unit, and the energy storage system, respectively. and These represent tie line power deviation and load disturbance, respectively. Indicates the area control error. and These represent the time constants of the steam turbine and the governor, respectively. This is the frequency offset coefficient. This is the speed adjustment coefficient. Indicates the load damping coefficient. Represents the inertia constant, parameter Characterizes the power exchange characteristics of the interconnected areas. and These represent the dynamic time constants of the wind turbine generator, the photovoltaic power generation unit, and the energy storage system, respectively. and Let represent the power allocation coefficients for different frequency modulation resources, and satisfy . .
3. The multi-energy power system load frequency coordinated control method based on reinforcement learning according to claim 1, characterized in that, Step 2 specifically includes: Step 2.1: Construct a reinforcement learning power allocation optimization framework. Use the multi-energy power system load frequency control model as the reinforcement learning environment. Use the system frequency deviation, the output power of each frequency regulation resource, and the actual available power of each frequency regulation resource as reinforcement learning state variables. Use the power allocation coefficient of each frequency regulation resource as the action space of the reinforcement learning agent. Step 2.2: Construct a comprehensive reward function that integrates frequency stability reward, wind power frequency regulation utilization reward, photovoltaic frequency regulation utilization reward, and energy storage power smoothing reward. Step 2.3: Train the power allocation strategy using the near-end strategy optimization algorithm, calculate the advantage function of the state transition samples, construct the truncated surrogate objective function to update the actor network parameters, update the critic network parameters according to the temporal difference error, and iteratively output the optimal power allocation coefficient.
4. The load frequency coordinated control method for multi-energy power systems based on reinforcement learning according to claim 3, characterized in that, The frequency stability reward item is specifically as follows: in, and These are the weighting coefficients. This indicates the frequency deviation threshold.
5. The multi-energy power system load frequency coordinated control method based on reinforcement learning according to claim 3, characterized in that, The wind power frequency regulation utilization incentive items are as follows: in, , This indicates the available power of the wind turbine. and These are the weighting coefficients.
6. The multi-energy power system load frequency coordinated control method based on reinforcement learning according to claim 3, characterized in that, The photovoltaic frequency regulation utilization incentive items are as follows: in, , This indicates the available power of the wind turbine. and These are the weighting coefficients.
7. The multi-energy power system load frequency coordinated control method based on reinforcement learning according to claim 3, characterized in that, The energy storage power smoothing reward item is specifically as follows: , in, This indicates the current regulation power of the energy storage system. This indicates the most recently sent power command, parameters. It is a weighting coefficient.
8. The multi-energy power system load frequency coordinated control method based on reinforcement learning according to claim 1, characterized in that, Step 3 specifically includes: Step 3.1, Select the area control error Design a PI-type load frequency controller for the control signal, with the following expression: in, This is the gain matrix; Step 3.2: Construct an adaptive memory-aware event triggering mechanism, combine the current system state with the error of historical triggering times to construct triggering conditions, and determine the control signal update time; Step 3.3: Reconstruct the controller based on the memory trigger mechanism, and generate the current control input by weighted fusion of control signals from historical trigger moments.
9. The multi-energy power system load frequency coordinated control method based on reinforcement learning according to claim 8, characterized in that, The adaptive memory-aware event triggering mechanism is as follows: in, Functions related to event triggering, The table indicates the next trigger time. For a fixed sampling period and satisfying , Indicates from the current trigger time Starting from the beginning, counting backwards... One sampling period; , , and ; This represents the floor function, which rounds its argument up to the nearest integer. express The upper bound, and satisfying , It is a non-zero real constant. , It is a given positive constant. It is a weight parameter, and satisfies , , This indicates the index of the previously transmitted data packet. It is an event triggering matrix, and .
10. The multi-energy power system load frequency coordinated control method based on reinforcement learning according to claim 8, characterized in that, The controller reconfiguration based on the memory triggering mechanism is specifically as follows: 。