Reinforcement learning-based network-source collaborative primary frequency optimization method and system
By constructing a hierarchical reinforcement learning framework, using a high-level policy network to generate dynamic frequency regulation weights and a low-level policy network to adjust output in real time, the frequency fluctuation problem caused by large-scale renewable energy access is solved, and efficient grid frequency regulation and frequency regulation accuracy are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GD POWER DEVELOPMENT CO LTD
- Filing Date
- 2025-09-27
- Publication Date
- 2026-05-12
AI Technical Summary
Existing frequency regulation methods cannot effectively cope with frequency fluctuations caused by the large-scale integration of renewable energy, resulting in slow response speed and low frequency regulation accuracy.
A hierarchical reinforcement learning framework is constructed, which includes a high-level policy network and a low-level policy network. The high-level policy network generates dynamic frequency modulation weights, which are combined with the low-level policy network to adjust the output in real time. A multi-level verification mechanism is used to ensure the security of the frequency modulation command.
It achieves efficient grid frequency regulation under renewable energy access conditions, improves primary frequency regulation efficiency and accuracy, and ensures grid stability and security.
Smart Images

Figure CN121192745B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of power dispatching and power generation technology, specifically to a grid-source collaborative primary frequency regulation optimization method and system based on reinforcement learning. Background Technology
[0002] With the rapid development of renewable energy, the integration of intermittent energy sources such as wind and solar power has significantly increased the complexity and challenges of power grid frequency regulation. Traditional frequency regulation methods typically rely on the regulation of large thermal power generating units, which have slow response times and high operating costs, making them difficult to effectively cope with frequently fluctuating power grid frequencies. Furthermore, the frequency regulation requirements of modern power systems are increasingly complex, involving the coordinated operation of multiple generating units. Traditional reinforcement learning methods often lack targeted strategies for power grid frequency regulation, making it difficult to simultaneously consider the real-time performance, economic efficiency, and unit health of frequency regulation. Summary of the Invention
[0003] This application provides a network-source collaborative primary frequency regulation optimization method and system based on reinforcement learning, which is used to solve the technical problem that existing frequency regulation methods cannot effectively cope with frequency fluctuations caused by large-scale renewable energy access, resulting in slow response speed and low frequency regulation accuracy.
[0004] The first aspect of this application provides a network-source collaborative primary frequency regulation optimization method based on reinforcement learning. The method includes: constructing a hierarchical reinforcement learning framework comprising a high-level policy network and a low-level policy network; based on the high-level policy network, outputting a dynamic frequency regulation weight vector as input to system frequency deviation, frequency regulation demand prediction, and unit operating status; based on the low-level policy network, outputting output adjustment commands for each frequency regulation unit as input to the dynamic frequency regulation weight vector and real-time frequency deviation, wherein each unit corresponds to a heterogeneous sub-policy network; performing real-time verification of the output adjustment commands through preset safety constraints; if the verification passes, executing frequency regulation control for each frequency regulation unit; if the verification fails, triggering a policy replanning mechanism, whereby the high-level policy network regenerates the dynamic frequency regulation weight vector and iteratively calculates new commands, which are then distributed to the governor via a network-source collaborative primary frequency regulation optimization device.
[0005] A second aspect of this application provides a network-source collaborative primary frequency regulation optimization system based on reinforcement learning. The system includes: a learning framework construction module for constructing a hierarchical reinforcement learning framework, which comprises a high-level policy network and a low-level policy network; a global frequency regulation analysis module for outputting a dynamic frequency regulation weight vector based on the high-level policy network, taking system frequency deviation, frequency regulation demand prediction, and unit operating status as inputs; a detailed output analysis module for outputting output adjustment commands for each frequency regulation unit based on the low-level policy network, taking the dynamic frequency regulation weight vector and real-time frequency deviation as inputs, wherein each unit corresponds to a heterogeneous sub-policy network; a real-time safety verification module for real-time verification of the output adjustment commands through preset safety constraints; if the verification passes, frequency regulation control of each frequency regulation unit is executed; and a policy replanning module for triggering a policy replanning mechanism if the verification fails, whereby the high-level policy network regenerates the dynamic frequency regulation weight vector and iteratively calculates new commands, which are then distributed to the governor via the network-source collaborative primary frequency regulation optimization device.
[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0007] The reinforcement learning-based grid-source collaborative primary frequency regulation optimization method and system provided in this application pertains to the fields of power dispatching and power generation technology. By constructing a hierarchical reinforcement learning framework including a high-level policy network and a low-level policy network, the high-level policy network generates dynamic frequency regulation weights based on global information, while the low-level policy network adjusts the output of each generating unit according to real-time data. A multi-level verification mechanism ensures the security of frequency regulation commands. This solves the technical problem that existing frequency regulation methods cannot effectively cope with frequency fluctuations caused by large-scale renewable energy access, resulting in slow response speed and low frequency regulation accuracy. It achieves efficient grid frequency regulation under renewable energy access through hierarchical reinforcement learning optimization, thereby improving the efficiency and accuracy of primary frequency regulation. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 A schematic flowchart of a network-source collaborative primary frequency modulation optimization method based on reinforcement learning provided in this application embodiment;
[0010] Figure 2 A schematic diagram of the network-source collaborative primary frequency modulation optimization system structure provided in this application embodiment.
[0011] Figure labeling: Learning framework construction module 11, global frequency modulation analysis module 12, detailed output analysis module 13, real-time safety verification module 14, strategy replanning module 15. Detailed Implementation
[0012] This application provides a network-source collaborative primary frequency regulation optimization method and system based on reinforcement learning, which is used to solve the technical problem that existing frequency regulation methods cannot effectively cope with frequency fluctuations caused by large-scale renewable energy access, resulting in slow response speed and low frequency regulation accuracy.
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0014] It should be noted that the terms "first," "second," etc., in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices.
[0015] Example 1, as Figure 1 As shown, this application provides a network-source cooperative primary frequency modulation optimization method based on reinforcement learning, which includes:
[0016] P10: Construct a hierarchical reinforcement learning framework, which includes a high-level policy network and a low-level policy network.
[0017] Furthermore, by combining offline training with online fine-tuning, the hierarchical strategy is first trained in the digital twin system and then transferred to actual power grid operation. Step P10 in this embodiment further includes:
[0018] P11: Construct a training environment based on the digital twin system, and train the high-level policy network and the low-level policy network in the training environment; P12: Migrate the trained high-level policy network and the low-level policy network to the actual power grid system, and fine-tune the parameters of the online primary frequency regulation optimization device of the high-level policy network and the low-level policy network based on the actual operating data.
[0019] The high-level policy network uses an attention mechanism to dynamically capture the state characteristics of key units; the low-level policy network adopts a heterogeneous network architecture, which includes heterogeneous sub-policy networks for different types of units, with each heterogeneous sub-policy network corresponding to a frequency-modulated unit; a cross-layer information fusion channel is configured between the high-level policy network and the low-level policy network, and bidirectional transmission of weight vectors and state information is performed based on the cross-layer information fusion channel.
[0020] It should be understood that this application optimizes the frequency regulation of the power grid by constructing a hierarchical reinforcement learning framework, and combines digital twin technology with an online fine-tuning mechanism to achieve efficient and accurate frequency regulation optimization in actual power grids.
[0021] Specifically, the process begins by constructing a complete training environment within the digital twin system. This environment is then used to train both the high-level and low-level policy networks. The trained networks are then migrated to the actual power grid system, and online fine-tuning of the primary frequency regulation optimization device parameters is performed based on real-world operational data. The digital twin system is a virtual digital model that closely mirrors the actual physical system, accurately simulating the power grid's operating state and dynamic characteristics. By constructing a training environment within the digital twin system, both the high-level and low-level policy networks can be trained under safe and controllable conditions.
[0022] In this training environment, the high-level policy network is primarily responsible for processing global information, such as system frequency deviation and load demand forecasting, to generate a frequency regulation weight vector. This information represents the regulation responsibility that each generating unit should undertake during frequency regulation. Furthermore, the high-level policy network employs an attention mechanism to dynamically capture key generating unit state characteristics. The attention mechanism is a neural network structure that simulates human attention selection, automatically identifying and focusing on the generating unit state characteristics most critical to frequency regulation control. For example, when the grid frequency deviation is large, the attention mechanism can prioritize the states of generating units with rapid response capabilities, thus providing more accurate input information to the high-level policy network and outputting a more reasonable dynamic frequency regulation weight vector. Through the analysis of the overall grid operating state, the high-level policy network can dynamically adjust the weights of each generating unit, enabling the system to effectively allocate frequency regulation tasks according to the current operating state.
[0023] The underlying strategy network focuses primarily on specific unit regulation operations, employing a heterogeneous network architecture that includes heterogeneous sub-strategy networks tailored to different types of units. This heterogeneous architecture can address the differences in frequency regulation capabilities, response speeds, and operating characteristics among different types of units (such as thermal power units, hydropower units, and renewable energy units). Each heterogeneous sub-strategy network corresponds to a frequency-regulating unit and can independently calculate output adjustment commands based on its respective unit's characteristics. For example, for hydropower units, the heterogeneous sub-strategy network might place greater emphasis on the unit's water flow inertia, mechanical inertia, and output adjustment range; while for renewable energy units, the heterogeneous sub-strategy network would focus more on the fluctuation and adjustability of their output. Based on real-time frequency deviation and unit status information, the underlying strategy network generates output adjustment commands for each unit according to the frequency regulation weight vector transmitted from the higher-level network. Furthermore, the underlying network can not only regulate based on guidance provided by the higher-level network but also independently optimize its regulation strategy based on changes in real-time data, thereby ensuring accurate unit responses in frequency control.
[0024] To achieve effective coordination between the higher-level and lower-level policy networks, a cross-layer information fusion channel is configured between them. This channel enables bidirectional information transmission between the two networks. Through this channel, weight vectors and state information can be exchanged between the two networks in a timely and accurate manner. For example, when the lower-level policy network detects a deviation between the actual output adjustment capability of a unit and the expectations of the higher-level policy network, it can feed this information back to the higher-level policy network through the cross-layer information fusion channel. The higher-level policy network can then readjust the dynamic frequency regulation weight vector based on this feedback, thereby optimizing the entire frequency regulation control process.
[0025] After training, the high-level and low-level policy networks are migrated to the actual power grid system. During this migration, due to the time-varying and complex nature of the actual power grid's operational data, online fine-tuning of the trained model is necessary. By collecting real-time operational data from the actual power grid, the parameters in the model are adjusted to better adapt to the actual power grid operating conditions. This online fine-tuning process continuously adjusts parameters based on real-time changes in the power grid and frequency regulation requirements, ensuring the model's effectiveness and stability in practical applications.
[0026] P20: Based on the aforementioned high-level strategy network, with system frequency deviation, frequency regulation demand prediction, and unit operating status as inputs, a dynamic frequency regulation weight vector is output.
[0027] Furthermore, step P20 in this embodiment of the application also includes:
[0028] P21: Real-time acquisition of power grid frequency measurement data and calculation of system frequency deviation; P22: Acquisition of frequency regulation demand forecast information from the dispatch center, and simultaneous acquisition of operating status parameters of each frequency regulation unit; P23: Inputting the system frequency deviation, frequency regulation demand forecast, and unit operating status into the high-level strategy network, calculating and outputting a dynamic frequency regulation weight vector, which includes three dimensions: frequency recovery weight, economic weight, and unit lifespan weight.
[0029] Optionally, the dynamic frequency regulation weight vector can be calculated based on the high-level strategy network. That is, the weight allocation of each frequency regulation unit can be dynamically adjusted according to the current power grid operating status and forecast information to optimize frequency regulation control.
[0030] First, real-time frequency measurement data of the power grid is acquired, and the system frequency deviation is calculated. The system frequency deviation refers to the difference between the actual frequency of the power grid and the target frequency, and is a key indicator for determining whether the power grid is operating stably. Frequency data of the power grid can be acquired in real time using measuring devices installed in the power grid, such as PMUs (Phasor Measurement Units). Then, the acquired frequency data is compared with the rated frequency of the power grid to calculate the system frequency deviation.
[0031] Simultaneously, the system acquires frequency regulation demand forecast information provided by the dispatch center and collects real-time operating status parameters of each frequency regulation unit. Frequency regulation demand forecast information can be calculated comprehensively based on multiple factors, including grid load forecasts, power generation plans, and forecasts of renewable energy generation, providing targets and direction for frequency regulation control. Meanwhile, the unit operating status parameters, such as output power and unit health status, provide detailed information about the current state of each unit, thus offering more comprehensive data support for frequency regulation decisions.
[0032] Next, frequency deviation, frequency regulation demand forecast information, and operating status parameters of each generating unit are input into the high-level strategy network for processing. The high-level strategy network, combined with reinforcement learning algorithms, analyzes and calculates these input data, outputting a dynamic frequency regulation weight vector. This dynamic frequency regulation weight vector includes three dimensions: frequency recovery weight, economic weight, and unit lifespan weight. The frequency recovery weight primarily focuses on how to quickly restore the grid frequency to its rated value to ensure grid stability; the economic weight considers the cost and benefits of frequency regulation, striving to achieve economical operation while meeting frequency regulation requirements; and the unit lifespan weight focuses on the long-term health of the units, avoiding excessive wear and tear caused by frequent frequency regulation and extending the units' service life. These weight vectors will guide the output adjustment of each frequency-regulating generating unit, ensuring that the grid achieves frequency restoration while balancing economic efficiency and equipment safety.
[0033] Furthermore, step P23 in this embodiment of the application also includes:
[0034] P23-1: Extract key features of the unit's operating status through the attention mechanism to generate feature attention weights; P23-2: Perform spatiotemporal correlation analysis on the system frequency deviation and frequency regulation demand prediction to generate a demand-deviation coupling coefficient; P23-3: Integrate the feature attention weights and the demand-deviation coupling coefficient to output a three-dimensional dynamic frequency regulation weight vector, wherein the weight values of each dimension are normalized by the Softmax function.
[0035] In one possible embodiment of this application, the calculation process of the dynamic frequency regulation weight vector can be further refined to ensure a more accurate reflection of the actual operating needs of the power grid and the status of the generating units.
[0036] First, key features of unit operating status are extracted using the attention mechanism in the high-level policy network. The attention mechanism is an advanced neural network technique that automatically weights important components of the input unit operating status parameters. By extracting features from the operating status parameters of each unit (including current output, ramp rate, and remaining frequency regulation capacity), a feature attention weight matrix is generated. This matrix quantifies the contribution of different units to frequency regulation capability under the current system state; for example, fast-response units receive higher attention weights. Based on these feature attention weights, the operating status of each unit can be focused on, improving the accuracy and effectiveness of frequency regulation control.
[0037] Next, a spatiotemporal correlation analysis is performed on the system frequency deviation and frequency modulation demand forecast to generate a demand-deviation coupling coefficient. Spatiotemporal correlation analysis is an analytical method that comprehensively considers time and space factors. It can reveal the relationship between frequency deviation and frequency modulation demand, and model these relationships based on temporal and spatial changes to generate the demand-deviation coupling coefficient. This coupling coefficient reflects the urgency of frequency modulation demand at different time scales, that is, the strength of the interaction between frequency modulation demand and frequency deviation under different time and space conditions.
[0038] Next, the feature attention weight matrix extracted through the attention mechanism is fused with the demand-bias coupling coefficient. These two key parameters are then fed into the feature fusion layer. The feature attention weight matrix undergoes a learnable linear transformation and is then multiplied element-wise with the demand-bias coupling coefficient using a Hadamard product. This is followed by a nonlinear mapping through a three-layer fully connected network containing a Tanh activation function, forming the final three-dimensional dynamic frequency regulation weight vector. This three-dimensional vector not only includes frequency recovery weights, economic weights, and unit lifespan weights, but also incorporates the results of the attention mechanism and spatiotemporal correlation analysis. This approach comprehensively considers different factors and calculates weights from multiple dimensions, ensuring comprehensive optimization during frequency regulation. Simultaneously, to ensure the balance and comparability of the weight vector, the weight values for each dimension are Softmax normalized, ensuring that all weight values are between 0 and 1 and sum to 1. This normalization process allows for comparison of weight values on the same scale, thus guaranteeing a reasonable weight allocation among different factors.
[0039] P30: Based on the underlying strategy network, with the dynamic frequency regulation weight vector and real-time frequency deviation as input, output the output adjustment command of each frequency regulation unit, wherein each unit corresponds to a heterogeneous sub-strategy network.
[0040] Furthermore, step P30 in this embodiment of the application also includes:
[0041] P31: Distribute the dynamic frequency modulation weight vector to each heterogeneous sub-strategy network; P32: Each heterogeneous sub-strategy network receives real-time frequency deviation data; P33: Calculate the benchmark output value of the corresponding unit based on the dynamic frequency modulation weight vector, and generate the final output adjustment command in combination with the real-time frequency deviation.
[0042] Specifically, the goal of the underlying policy network is to generate output adjustment commands for each frequency regulation unit based on the dynamic frequency regulation weight vector and real-time frequency deviation information output by the higher-level policy network, thereby achieving precise frequency regulation.
[0043] First, the dynamic frequency regulation weight vector obtained from the high-level policy network is distributed to each heterogeneous sub-policy network. Each frequency regulation unit corresponds to an independent heterogeneous sub-policy network, and these heterogeneous sub-policy networks make customized frequency regulation decisions for different types of units. Since the characteristics of each unit (such as maximum power, response time, etc.) are different, the use of heterogeneous sub-policy networks allows each unit to perform independent frequency regulation according to its own characteristics, ensuring a more flexible and efficient frequency regulation response.
[0044] Next, each heterogeneous sub-strategy network receives real-time frequency deviation information and uses this data as input, combined with information from various dimensions of the dynamic frequency regulation weight vector (such as frequency recovery weight, economic weight, and unit lifespan weight), to make regulation decisions. Real-time frequency deviation data reflects the difference between the current grid frequency and the rated frequency, and is a direct triggering factor for frequency regulation control. Each heterogeneous sub-strategy network receives real-time frequency deviation data through a connection to the grid monitoring system.
[0045] Next, each heterogeneous sub-strategy network calculates the baseline output value of the corresponding unit based on the dynamic frequency regulation weight vector. The baseline output value is a preliminary adjustment value calculated based on the unit's current operating status and the frequency regulation weight vector, reflecting the unit's ideal output level in the current frequency regulation task. However, frequency fluctuations in the actual power grid change in real time. Therefore, it is necessary to further combine real-time frequency deviation information to correct the baseline output value and generate the final output adjustment command. That is, based on the baseline output value and the magnitude and direction of the frequency deviation, the required increase or decrease in unit output is calculated, and the output adjustment command is generated.
[0046] Through this process, the output adjustment command of each frequency regulation unit is precisely calculated according to its independent sub-strategy network, ensuring that each unit performs at its best during dynamic frequency regulation.
[0047] Furthermore, step P33 in this embodiment of the application also includes:
[0048] P33-1: Invoke the corresponding heterogeneous sub-strategy network according to the unit type, input the frequency recovery weight of the dynamic frequency regulation weight vector, and calculate the initial power adjustment amount; P33-2: Based on the initial power adjustment amount, perform dynamic compensation by superimposing the differential term of the real-time frequency deviation to generate a benchmark output value; P33-3: Combine the economic weight and unit life weight of the dynamic frequency regulation weight vector to perform multi-objective optimization correction on the benchmark output value, and output the final output adjustment command that meets the unit ramp rate constraint.
[0049] Optionally, the generation process of output adjustment commands can be further refined to ensure that the output adjustment commands of each frequency regulation unit can not only respond quickly to frequency deviations, but also take into account economy and unit life.
[0050] First, based on the type of each generating unit, the corresponding heterogeneous sub-strategy network is invoked, and the frequency recovery weight from the dynamic frequency regulation weight vector output by the higher-level strategy network is input into this heterogeneous sub-strategy network to calculate the initial power adjustment. The frequency recovery weight reflects the system's priority for frequency recovery in the current frequency regulation task. Based on this weight value, the power adjustment model of the corresponding heterogeneous sub-strategy network can be fine-tuned, and the initial power adjustment for each generating unit can be calculated based on the adjusted power adjustment model. For example, for hydropower units with rapid response capabilities, their heterogeneous sub-strategy network may calculate a larger initial power adjustment to quickly respond to frequency deviations; while for thermal power units with slower response speeds, the initial power adjustment will be relatively smaller to avoid over-adjustment.
[0051] Subsequently, the initial power adjustment is dynamically compensated based on the rate of change of the real-time frequency deviation. The differential term of the frequency deviation refers to the rate of change of the frequency deviation over time, revealing the trend and speed of frequency fluctuations. Adding this differential term to the initial power adjustment allows for a rapid response to dynamic changes in frequency fluctuations, thereby generating a baseline output value. Specifically, the differential term can be calculated by differentiating the real-time frequency deviation over time, for example, by using a difference method to calculate the rate of change of the frequency deviation at adjacent time points. Adding this differential term to the initial power adjustment yields the baseline output value. This baseline value represents a preliminary power adjustment that reflects the system's initial demand for frequency recovery, but it does not yet take into account other constraints such as economic efficiency and unit lifespan.
[0052] After generating the baseline output value, a multi-objective optimization correction is further performed by combining the economic weight and the unit life weight in the dynamic frequency regulation weight vector. The economic weight reflects the cost and benefit considerations during frequency regulation, while the unit life weight focuses on the long-term health of the unit. The multi-objective optimization correction process is a comprehensive trade-off process, aiming to achieve a balance between frequency restoration, economy, and unit life. Specifically, this process can be achieved by establishing a multi-objective optimization model. The objective function of this model can be expressed as:
[0053] ;in, These are the frequency recovery weight, the economic weight, and the unit lifespan weight, respectively. Their values are dynamically adjusted by the high-level strategy network based on the current grid status and frequency regulation needs. f(frequency recovery), e(economic efficiency), and l(unit lifespan) are the optimization objective functions for frequency recovery, economic efficiency, and unit lifespan, respectively. For example, the frequency recovery objective function can be a quantitative indicator of the unit output adjustment response speed and recovery effect to frequency deviation; the economic objective function can be a quantitative indicator of fuel consumption, power generation costs, etc. during frequency regulation; and the unit lifespan objective function can be a quantitative indicator of unit equipment wear and maintenance costs, etc.
[0054] In multi-objective optimization models, trade-offs between different objectives can be achieved by adjusting weight values. For example, when the frequency deviation is large and rapid recovery is required, the frequency recovery weight can be increased, making the generating unit prioritize frequency recovery; when the grid operation is relatively stable and long-term economic efficiency needs to be considered, the economic efficiency weight can be increased, making the generating unit more cost-effective during frequency regulation; when it is necessary to extend the service life of the generating unit, the service life weight can be increased, preventing excessive wear and tear on the generating unit during frequency regulation.
[0055] Simultaneously, when performing multi-objective optimization corrections, the unit's ramp rate constraint must also be considered. The ramp rate constraint means that the unit's output adjustment rate cannot exceed its designed maximum ramp rate. This constraint ensures the unit can operate safely and stably during frequency regulation. For example, if the calculated output adjustment command exceeds the unit's maximum ramp rate, the command needs to be adjusted to comply with the ramp rate constraint. For instance, the output adjustment command can be executed in segments, with the output adjustment amount within each segment not exceeding the unit's maximum ramp rate.
[0056] Ultimately, after considering multi-objective optimization and ramp rate constraints, the generated output adjustment command will be able to maximize economy and protect the long-term health of the unit while meeting the needs of grid frequency restoration.
[0057] P40: The output adjustment command is verified in real time by a preset safety constraint. If the verification is successful, the frequency regulation control of each frequency regulation unit is executed.
[0058] Furthermore, step P40 in this embodiment of the application also includes:
[0059] P41: The preset safety constraints include multi-level verification channels, which include a pre-verification channel during instruction generation and a final verification channel before issuance; P42: Based on the real-time task stage, the corresponding verification channel is activated to perform real-time verification of the output adjustment instruction.
[0060] It should be understood that real-time verification of output adjustment commands is necessary to ensure the safety and reliability of frequency regulation control. This process includes multiple verification stages to ensure that the commands meet the system's safety requirements during generation, issuance, and execution.
[0061] First, pre-defined safety constraints are implemented through multi-level verification channels, including a pre-verification channel during command generation and a final verification channel before issuance. The pre-verification channel performs initial verification immediately after the output adjustment command is generated, ensuring that the command's basic parameters meet safety requirements. For example, it checks whether the output adjustment command exceeds the unit's output range and whether it complies with the unit's ramp rate constraints. The final verification channel performs final verification before the command is issued, ensuring that the command fully complies with grid operation safety standards before actual execution. The final verification channel considers more real-time operating conditions, such as the grid's current frequency and the operating status of other units, to ensure that the execution of the command does not pose a threat to grid security.
[0062] Based on different stages of the real-time task, corresponding verification channels are activated to perform real-time verification of the output adjustment commands. Real-time task stages refer to different points in time and states during the frequency regulation control process, such as after command generation and before command issuance. Activating corresponding verification channels according to different task stages ensures that the output adjustment commands are rigorously checked at each critical point. For example, after command generation, the pre-verification channel is activated immediately to verify the basic parameters of the command; before command issuance, the final verification channel is activated to perform a comprehensive final verification of the command. This phased verification mechanism can promptly identify and correct potential problems, ensuring that the output adjustment commands fully comply with safety requirements before execution.
[0063] The verification process not only checks whether the commands meet the basic requirements of the equipment's operational capabilities, but also takes into account the real-time operating status of the power grid, such as the severity of load fluctuations, frequency fluctuations, and other potential risk factors. Only when the verification passes will the system allow the issuance and execution of commands. This rigorous safety verification mechanism not only improves the reliability of frequency regulation control but also ensures the safety and stability of power grid operation.
[0064] P50: If the verification fails, the policy replanning mechanism is triggered. The higher-level policy network regenerates the dynamic frequency modulation weight vector and iteratively calculates the new instruction, which is then sent to the speed controller through the network-source collaborative frequency modulation optimization device.
[0065] Furthermore, step P50 in this embodiment of the application also includes:
[0066] P51: Record the type and degree of security constraint violation of the verification failure instruction and generate a replanning trigger signal; P52: Feed the replanning trigger signal back to the higher-level policy network through the cross-layer information fusion channel and adjust the weight allocation strategy of the attention mechanism; P53: Based on the corrected unit status characteristics and the updated frequency regulation demand prediction, regenerate the dynamic frequency regulation weight vector and start the iterative calculation of the lower-level policy network until the security verification pass instruction is output.
[0067] Optionally, when the output adjustment command fails during the safety verification process, the system will trigger a strategy replanning mechanism to regenerate the dynamic frequency regulation weight vector and iteratively calculate the new frequency regulation command, ensuring that the frequency regulation command of the power grid always complies with safety constraints and can effectively cope with various abnormal situations in real-time operation.
[0068] First, when an output adjustment command is deemed unsuccessful during the verification process, the system records the type and severity of the safety constraint violation. For example, if the command violates the unit's output range constraint, the specific value exceeding the range needs to be recorded; if it violates the ramp rate constraint, the rate of exceedance is recorded, etc. Based on these records, a replanning trigger signal is generated, which contains specific information about the verification failure.
[0069] Next, the replanning trigger signal is fed back to the higher-level policy network via a cross-layer information fusion channel. This channel transmits crucial information between the higher and lower-level policy networks to ensure synchronized adjustments. Upon receiving the replanning trigger signal, the higher-level policy network adjusts the weight allocation strategy of the attention mechanism based on the type and severity of the violation contained in the signal. For example, if the verification failure is due to a violation of the unit's output range constraint, the attention mechanism may increase the weight given to this feature to more accurately consider this constraint in subsequent policy generation.
[0070] Furthermore, based on the revised unit state characteristics and the updated frequency regulation demand forecast, the dynamic frequency regulation weight vector is recalculated. The revised unit state characteristics are the result adjusted based on the information from failed verifications, reflecting more accurate state information of the units under current operating conditions. The updated frequency regulation demand forecast considers real-time grid operation and potential trends. The higher-level policy network, based on the adjusted attention mechanism's weight allocation strategy and incorporating this revised and updated information, recalculates the dynamic frequency regulation weight vector to generate a frequency regulation strategy that better conforms to safety constraints. This ensures that the new frequency regulation weight vector more closely matches the actual operating environment, thereby improving the efficiency and accuracy of frequency regulation.
[0071] After the new dynamic frequency regulation weight vector is generated, iterative calculations of the underlying strategy network are initiated. Based on the new dynamic frequency regulation weight vector, the underlying strategy network recalculates the output adjustment commands for each frequency regulation unit and performs a safety check again. This process continues until the new frequency regulation command passes the safety check. Then, the new frequency regulation command is sent to the governor through the grid-source collaborative primary frequency regulation optimization device for grid frequency regulation control.
[0072] The introduction of a strategy replanning mechanism can ensure that when the system encounters instructions that do not meet safety constraints, it can adjust its strategy in a timely manner and regenerate instructions that meet safety requirements. This improves the adaptability and robustness of the power grid frequency regulation system, enabling it to respond to various emergencies in real time and ensure the stable operation of the power grid.
[0073] In summary, the embodiments of this application have at least the following technical effects:
[0074] This application employs a hierarchical reinforcement learning framework, where a high-level policy network dynamically generates frequency regulation weight vectors, and a low-level policy network rapidly calculates output adjustment commands based on real-time frequency deviations, significantly improving the response speed and accuracy of frequency regulation control. By combining offline training with a digital twin system and online fine-tuning with the actual power grid, the frequency regulation strategy can flexibly adapt to changes in the power grid's operating state, ensuring effectiveness in different scenarios. Multi-level verification channels are set up to perform real-time verification of output adjustment commands, triggering strategy replanning when safety constraints are not met, ensuring the safety and reliability of frequency regulation control. The dynamic frequency regulation weight vector covers three dimensions: frequency recovery, economy, and unit lifespan. Through multi-objective optimization and correction, it balances the speed, economy, and equipment lifespan of frequency regulation. It can effectively cope with frequency fluctuations caused by large-scale renewable energy integration, improve the power grid's frequency regulation capability, and ensure system stability and security.
[0075] This achieves the technical effect of optimizing the grid frequency under renewable energy access through hierarchical reinforcement learning, thereby improving the efficiency and accuracy of primary frequency regulation.
[0076] Example 2, based on the same inventive concept as the reinforcement learning-based network-source collaborative primary frequency modulation optimization method in the previous examples, such as... Figure 2 As shown, this application provides a network-source collaborative primary frequency modulation optimization system based on reinforcement learning. The system and method embodiments in this application are based on the same inventive concept. The system includes:
[0077] The learning framework construction module 11 is used to construct a hierarchical reinforcement learning framework, which includes a high-level policy network and a low-level policy network.
[0078] The global frequency regulation analysis module 12 is used to output a dynamic frequency regulation weight vector based on the high-level strategy network, taking system frequency deviation, frequency regulation demand prediction and unit operating status as inputs.
[0079] The detailed output analysis module 13 is used to output output adjustment commands for each frequency-regulating unit based on the underlying strategy network, with the dynamic frequency regulation weight vector and real-time frequency deviation as inputs. Each unit corresponds to a heterogeneous sub-strategy network.
[0080] The real-time safety verification module 14 is used to verify the output adjustment command in real time through preset safety constraints. If the verification is successful, the frequency regulation control of each frequency regulation unit is executed.
[0081] The strategy replanning module 15 is used to trigger the strategy replanning mechanism if the verification fails. The higher-level strategy network regenerates the dynamic frequency modulation weight vector and iteratively calculates the new instruction, which is then sent to the speed controller through the network-source collaborative frequency modulation optimization device.
[0082] Furthermore, the learning framework construction module 11 is also used to perform the following steps:
[0083] The method combines offline training with online fine-tuning. First, the hierarchical strategy is trained in the digital twin system and then migrated to the actual power grid operation. This includes: building a training environment based on the digital twin system, training the high-level strategy network and the low-level strategy network in the training environment; migrating the trained high-level strategy network and the low-level strategy network to the actual power grid system, and fine-tuning the parameters of the high-level strategy network and the low-level strategy network online based on actual operating data.
[0084] Furthermore, in the learning framework construction module 11:
[0085] The high-level policy network uses an attention mechanism to dynamically capture the state characteristics of key units; the low-level policy network adopts a heterogeneous network architecture, which includes heterogeneous sub-policy networks for different types of units, with each heterogeneous sub-policy network corresponding to a frequency-modulated unit; a cross-layer information fusion channel is configured between the high-level policy network and the low-level policy network, and bidirectional transmission of weight vectors and state information is performed based on the cross-layer information fusion channel.
[0086] Furthermore, the global frequency modulation analysis module 12 is also used to perform the following steps:
[0087] Real-time acquisition of power grid frequency measurement data, calculation of system frequency deviation; acquisition of frequency regulation demand forecast information from the dispatch center, and simultaneous acquisition of operating status parameters of each frequency regulation unit; input of the system frequency deviation, frequency regulation demand forecast, and unit operating status into the high-level strategy network, and calculation and output of dynamic frequency regulation weight vector, which includes three dimensions: frequency recovery weight, economic weight, and unit lifespan weight.
[0088] Furthermore, the global frequency modulation analysis module 12 is also used to perform the following steps:
[0089] The key features of the unit's operating status are extracted through the attention mechanism to generate feature attention weights; the system frequency deviation and frequency regulation demand prediction are analyzed in a spatiotemporal manner to generate a demand-deviation coupling coefficient; the feature attention weights and the demand-deviation coupling coefficient are fused to output a three-dimensional dynamic frequency regulation weight vector, wherein the weight values of each dimension are normalized by the Softmax function.
[0090] Furthermore, the detailed output analysis module 13 is also used to perform the following steps:
[0091] The dynamic frequency modulation weight vector is distributed to each heterogeneous sub-strategy network; each heterogeneous sub-strategy network receives real-time frequency deviation data; the reference output value of the corresponding unit is calculated based on the dynamic frequency modulation weight vector, and the final output adjustment command is generated in combination with the real-time frequency deviation.
[0092] Furthermore, the detailed output analysis module 13 is also used to perform the following steps:
[0093] Based on the unit type, the corresponding heterogeneous sub-strategy network is invoked, the frequency recovery weight of the dynamic frequency regulation weight vector is input, and the initial power adjustment is calculated. Based on the initial power adjustment, the differential term of the real-time frequency deviation is superimposed for dynamic compensation to generate a benchmark output value. Combining the economic weight and the unit life weight of the dynamic frequency regulation weight vector, the benchmark output value is optimized and corrected in multiple objectives, and the final output adjustment command that meets the unit ramp rate constraint is output.
[0094] Furthermore, the real-time security verification module 14 is also used to perform the following steps:
[0095] The preset safety constraints include multi-level verification channels, which include a pre-verification channel during instruction generation and a final verification channel before issuance. Based on the real-time task stage, the corresponding verification channel is activated to perform real-time verification of the output adjustment instruction.
[0096] Furthermore, the strategy replanning module 15 is also used to perform the following steps:
[0097] Record the type and degree of security constraint violation of the verification failure instruction, and generate a replanning trigger signal; feed the replanning trigger signal back to the high-level policy network through the cross-layer information fusion channel to adjust the weight allocation strategy of the attention mechanism; based on the corrected unit status characteristics and the updated frequency regulation demand prediction, regenerate the dynamic frequency regulation weight vector and start the iterative calculation of the bottom-level policy network until the security verification pass instruction is output.
[0098] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0099] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0100] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application intends to include such modifications and variations.
Claims
1. A network-source collaborative primary frequency modulation optimization method based on reinforcement learning, characterized in that, The method includes: A hierarchical reinforcement learning framework is constructed, which includes a high-level policy network and a low-level policy network. Based on the high-level strategy network, with system frequency deviation, frequency regulation demand forecast and unit operating status as inputs, a dynamic frequency regulation weight vector is output. Based on the underlying strategy network, the output adjustment command of each frequency-regulating unit is output with the dynamic frequency regulation weight vector and real-time frequency deviation as input, wherein each unit corresponds to a heterogeneous sub-strategy network. The output adjustment command is verified in real time by setting a safety constraint. If the verification passes, the frequency regulation control of each frequency regulation unit is executed. If the verification fails, the policy replanning mechanism is triggered. The higher-level policy network regenerates the dynamic frequency modulation weight vector and iteratively calculates the new instruction, which is then sent to the speed controller through network-source collaborative frequency modulation optimization device. A hierarchical reinforcement learning framework is constructed, comprising a high-level policy network and a low-level policy network, including: The high-level strategy network uses an attention mechanism to dynamically capture key unit status characteristics. The underlying policy network adopts a heterogeneous network architecture, which includes heterogeneous sub-policy networks for different types of units, with each heterogeneous sub-policy network corresponding to a frequency regulation unit. A cross-layer information fusion channel is configured between the high-level policy network and the low-level policy network, and bidirectional transmission of weight vectors and state information is performed based on the cross-layer information fusion channel. Based on the aforementioned high-level strategy network, taking system frequency deviation, frequency regulation demand forecast, and unit operating status as inputs, the output is a dynamic frequency regulation weight vector, including: Real-time acquisition of power grid frequency measurement data; calculation of system frequency deviation. Obtain frequency regulation demand forecast information from the dispatch center, and simultaneously collect operating status parameters of each frequency regulation unit; The system frequency deviation, frequency regulation demand prediction, and unit operating status are input into the high-level strategy network to calculate and output a dynamic frequency regulation weight vector, which includes three dimensions: frequency recovery weight, economic weight, and unit life weight.
2. The network-source collaborative primary frequency modulation optimization method based on reinforcement learning as described in claim 1, characterized in that, A combination of offline training and online fine-tuning was adopted, first training the hierarchical strategy in the digital twin system, and then migrating it to the actual power grid operation, including; A training environment is constructed based on a digital twin system, and the high-level policy network and the low-level policy network are trained in the training environment. The trained high-level and low-level policy networks are migrated to the actual power grid system, and the parameters of the online primary frequency regulation optimization device are fine-tuned based on the actual operating data.
3. The network-source collaborative primary frequency modulation optimization method based on reinforcement learning as described in claim 1, characterized in that, The system frequency deviation, frequency regulation demand forecast, and unit operating status are input into the high-level strategy network to calculate and output a dynamic frequency regulation weight vector, including: The key features of the unit's operating status are extracted using the attention mechanism, and feature attention weights are generated. The system frequency deviation and frequency modulation demand prediction are subjected to spatiotemporal correlation analysis to generate a demand-deviation coupling coefficient. By fusing the feature attention weights and the demand-bias coupling coefficient, a three-dimensional dynamic frequency tuning weight vector is output, wherein the weight values of each dimension are normalized by the Softmax function.
4. The network-source collaborative primary frequency modulation optimization method based on reinforcement learning as described in claim 3, characterized in that, Based on the underlying strategy network, and taking the dynamic frequency regulation weight vector and real-time frequency deviation as input, the output adjustment commands for each frequency regulation unit are output, including: The dynamic frequency-adjusted weight vector is distributed to each heterogeneous sub-policy network; Each heterogeneous sub-policy network receives real-time frequency deviation data; The reference output value of the corresponding unit is calculated based on the dynamic frequency adjustment weight vector, and the final output adjustment command is generated by combining the real-time frequency deviation.
5. The network-source collaborative primary frequency modulation optimization method based on reinforcement learning as described in claim 4, characterized in that, The reference output value of the corresponding unit is calculated based on the dynamic frequency adjustment weight vector, and the final output adjustment command is generated by combining the real-time frequency deviation, including: Based on the unit type, the corresponding heterogeneous sub-strategy network is invoked, the frequency recovery weight of the dynamic frequency modulation weight vector is input, and the initial power adjustment is calculated. Based on the initial power adjustment, a dynamic compensation is performed by superimposing the differential term of the real-time frequency deviation to generate a reference output value. Combining the economic weight and unit life weight of the dynamic frequency regulation weight vector, the benchmark output value is optimized and corrected in multiple objectives, and the final output adjustment command that meets the unit ramp rate constraint is output.
6. The network-source collaborative primary frequency modulation optimization method based on reinforcement learning as described in claim 1, characterized in that, The output adjustment command is verified in real time by means of preset safety constraints, including: The preset security constraints include multi-level verification channels, which include a pre-verification channel during instruction generation and a final verification channel before issuance. Based on the real-time task phase, the corresponding verification channel is activated to perform real-time verification of the output adjustment command.
7. The network-source collaborative primary frequency modulation optimization method based on reinforcement learning as described in claim 6, characterized in that, If the verification fails, a policy replanning mechanism is triggered, in which the higher-level policy network regenerates the dynamic frequency modulation weight vector and iteratively calculates the new instructions, including: Record the type and extent of security constraint violations in verification failure instructions, and generate a replanning trigger signal; The replanning trigger signal is fed back to the high-level policy network through a cross-layer information fusion channel to adjust the weight allocation strategy of the attention mechanism. Based on the corrected unit status characteristics and the updated frequency regulation demand forecast, the dynamic frequency regulation weight vector is regenerated and the iterative calculation of the underlying policy network is initiated until the safety verification command is output.
8. A network-source collaborative primary frequency modulation optimization system based on reinforcement learning, characterized in that, The system includes: A learning framework building module is used to build a hierarchical reinforcement learning framework, which includes a high-level policy network and a low-level policy network. The global frequency regulation analysis module is used to output a dynamic frequency regulation weight vector based on the high-level strategy network, taking system frequency deviation, frequency regulation demand forecast and unit operating status as inputs. The detailed output analysis module is used to output output adjustment commands for each frequency-regulating unit based on the underlying strategy network, with the dynamic frequency regulation weight vector and real-time frequency deviation as inputs. Each unit corresponds to a heterogeneous sub-strategy network. The real-time safety verification module is used to verify the output adjustment command in real time through preset safety constraints. If the verification is successful, the frequency regulation control of each frequency regulation unit is executed. The strategy replanning module is used to trigger the strategy replanning mechanism if the verification fails. The higher-level strategy network regenerates the dynamic frequency modulation weight vector and iteratively calculates the new instruction, which is then sent to the speed controller through network-source collaborative frequency modulation optimization device. Furthermore, in the learning framework construction module: The high-level policy network uses an attention mechanism to dynamically capture the state characteristics of key units; the low-level policy network adopts a heterogeneous network architecture, which includes heterogeneous sub-policy networks for different types of units, with each heterogeneous sub-policy network corresponding to a frequency-modulated unit; a cross-layer information fusion channel is configured between the high-level policy network and the low-level policy network, and bidirectional transmission of weight vectors and state information is performed based on the cross-layer information fusion channel. The global frequency modulation analysis module is also used to perform the following steps: Real-time acquisition of power grid frequency measurement data, calculation of system frequency deviation; acquisition of frequency regulation demand forecast information from the dispatch center, and simultaneous acquisition of operating status parameters of each frequency regulation unit; input of the system frequency deviation, frequency regulation demand forecast, and unit operating status into the high-level strategy network, and calculation and output of dynamic frequency regulation weight vector, which includes three dimensions: frequency recovery weight, economic weight, and unit lifespan weight.