Resource adjustment method and system based on power grid frequency deviation

By integrating reinforcement learning agents and model prediction control in the AGC controller and dynamically adjusting the control strategy, the problem of insufficient adaptability of traditional AGC controllers in the new energy environment is solved, and stable and efficient adjustment of the power grid frequency is achieved.

CN120497971AActive Publication Date: 2025-08-15HEFEI POWER SUPPLY COMPANY OF STATE GRID ANHUI ELECTRIC POWER +1

Patent Information

Application Number
CN202510990395.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-08-15
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

Traditional AGC controllers are difficult to adapt to the randomness and uncertainty of new energy output, especially when the penetration rate of new energy is high, it is difficult to adjust the grid frequency in a timely and precise manner, and lack adaptability, resulting in a lag in frequency adjustment.

Method used

Through reinforcement learning, the trained control agent is integrated into the AGC controller, combining soft switching and policy interpolation mechanisms, dynamically switch control strategies, and correct the unit setting values ​​through model prediction control to achieve adaptive selection and precise control.

Benefits of technology

It improves the adaptability and control accuracy of the AGC controller, enhances the robustness to load disturbances and new energy fluctuations, realizes the stability and sensitivity of frequency, avoids the regulation oscillation caused by sudden changes in the control signal, and improves the economic and stability of resource scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120497971A_ABST
    Figure CN120497971A_ABST
Patent Text Reader

Abstract

The invention discloses a resource adjustment method and system based on power grid frequency deviation, and relates to the technical field of AGC control, and the method comprises the following steps: obtaining environment state data, and constructing an environment state vector; initializing a control agent according to a preset control mode, taking the environment state vector as input, training the control agent through a reinforcement learning reward function, and outputting a strategy action; dynamically switching the strategy action into a control strategy of the AGC controller based on a soft switching and strategy interpolation mechanism; a model prediction control method is adopted, and the power set value of the AGC control unit is optimized in a rolling mode; according to the method, the trained control agent is integrated into the AGC controller through reinforcement learning, so that the AGC controller can dynamically switch the control strategy of the AGC controller, self-adaptive selection of control actions is realized, and the power grid frequency is controlled to be stable by correcting the set value of the unit in the AGC controller and improving the control precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of AGC control technology, and more particularly, to a resource regulation method and system based on power grid frequency deviation. Background Art

[0002] With the rapid development of renewable energy generation and energy storage technologies, grid frequency regulation faces unprecedented challenges. Traditional power systems rely primarily on large-scale thermal and hydropower plants with high inertia for frequency regulation. However, renewable energy generation is intermittent and volatile, and the controllability of energy storage systems is affected by their capacity and charge-discharge characteristics, making grid frequency regulation even more complex.

[0003] Automatic Generation Control (AGC) complementary coordinated control is a technology that optimizes the balance between generator scheduling and load demand in power grids, particularly when a large number of renewable energy sources and energy storage systems are connected to the grid. The primary purpose of this control technology is to ensure grid stability and power quality by coordinating the output of traditional generators, energy storage equipment, and renewable energy generation systems.

[0004] Specifically, AGC complementary coordinated control combines the control capabilities of different types of regulating resources and coordinates their output to balance the generation and load demand in the power grid, thereby maintaining the frequency stability and reliability of the power grid.

[0005] For example, the invention patent announcement with announcement number: CN103439962B discloses a closed-loop detection and verification method for automatic power generation control of a power grid, including: the power grid AGC master station calculates and generates an AGC strategy based on the power grid section, and sends the strategy to the unit AGC substation. After the unit AGC substation receives the strategy and corrects it, it starts to simulate the output change of the unit. The flow simulation module performs continuous time series simulation with a period of 1s based on the output change of the AGC unit, and generates a new power grid section for the power grid AGC master station of the next period to generate a new control strategy. The power grid AGC detection and verification method provided by the present invention has a simulation period of only 1s, which fully meets and exceeds the requirements of real-time data acquisition, simulates the changes in the power grid operating status during the AGC adjustment process very well, and can intuitively reflect the closed-loop control effect of the power grid AGC; it is a "plug and play" detection and verification method that is simple, easy to use, and effective.

[0006] For example, the invention patent publication number CN108845492A discloses an intelligent predictive control method for an AGC controller based on the CPS evaluation standard, including the following steps: 1. Establishing a unit coordinated control model through a neural network; 2. Using the neural network prediction model obtained by field data identification, the system predicts the output at the next moment based on the control inputs at the current moment and the next moment, as well as the actual output at the current moment; 3. Correcting the predicted output at the next moment using the deviation between the historically predicted output at the current moment and the actual output at the current moment; 4. Using the current control input, actual output, and set value as inputs to the neural network, the neural network predicts the control input at the next moment; 5. Repeating steps 2 to 4, the AGC controller selects a PID controller on the grid dispatching side to achieve system frequency control. Compared with the existing technology, the control effect of the present invention meets the AGC performance adjustment requirements and further improves the control performance.

[0007] The above disclosed technical solutions have at least the following technical problems: Traditional AGC controllers primarily rely on fixed rules and optimized scheduling to maintain grid frequency stability. These control strategies often rely on historical experience and are difficult to adapt to the randomness and uncertainty of renewable energy output. Especially with high renewable energy penetration, traditional rules struggle to adjust grid frequency in a timely and accurate manner. Furthermore, traditional AGC controllers have lags in regulation and lack adaptive capabilities, making them difficult to handle complex dynamic environments such as severe fluctuations in renewable energy output and rapid load changes.

[0008] In view of the above problems, the present invention proposes a solution. Summary of the Invention

[0009] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a resource regulation method and system based on frequency deviation, which integrates the trained control agent into the AGC controller through reinforcement learning, so that the AGC controller can dynamically switch its control strategy and realize adaptive selection of control actions. By correcting the set values of the units in the AGC controller, the control accuracy is improved, thereby controlling the grid frequency to achieve stability.

[0010] To achieve the above object, the present invention provides the following technical solutions: The resource regulation method based on grid frequency deviation includes the following steps: obtaining environmental state data and constructing an environmental state vector; initializing a control agent according to a preset control mode, using the environmental state vector as input, training the control agent through a reinforcement learning reward function, and outputting a strategic action; dynamically switching the control strategy of the AGC controller based on the soft switching and strategy interpolation mechanism; and using a model predictive control method to continuously optimize the power setpoint of the AGC control unit.

[0011] In a preferred embodiment, the environmental status data includes frequency deviation, historical load data, and new energy power, and the new energy power includes photovoltaic power and wind power; the frequency deviation is classified based on a dynamic threshold to obtain a frequency deviation category, and the frequency deviation is calculated using frequency data, and the frequency data includes a rated frequency and a real-time frequency measurement value; based on an ARIMA time series algorithm, the load fluctuation is obtained by monitoring the load change rate of historical load data; based on the ARIMA time series algorithm, a load demand change model is established through historical load data, and the load fluctuation is monitored in real time by calculating the load change rate; a new energy power prediction model is established based on machine learning through photovoltaic power and wind power, and the new energy output is obtained through the new energy power prediction model; and the environmental state vector is obtained through the frequency deviation category, the load fluctuation, and the new energy output.

[0012] In a preferred embodiment, a method for constructing a reinforcement learning reward function includes: setting the environment state vector as the state input of the intelligent agent; defining the control action of the intelligent agent as a preset control mode set, wherein the control mode set includes a short-cycle control mode, a medium-cycle control mode, and a long-cycle control mode; constructing a reinforcement learning reward function and constraint conditions for the control action of the intelligent agent, wherein the constraint conditions include control cost constraint, control capability constraint, new energy utilization constraint, and control direction constraint.

[0013] In a preferred embodiment, a control agent is trained by a reinforcement learning reward function to output a strategic action, including: obtaining an independent state vector for each agent in the environment; each agent performs different control actions based on its state vector; calculating the immediate reward value of each control action based on a preset control action reinforcement learning reward function; training each agent through a policy optimization algorithm with the goal of maximizing the cumulative reward; obtaining a trained control agent, in which each state is mapped to a unique optimal control action, forming a deterministic strategic action.

[0014] In a preferred embodiment, based on the soft switching and strategy interpolation mechanism, the strategy action is dynamically switched to the control strategy of the AGC controller, specifically: the current environment state vector is matched with the trained control agent, and the corresponding optimal strategy action, i.e., the target strategy, is output; the similarity between the historical strategy action and the target strategy is calculated, and when the similarity is less than a preset threshold, the exponential interpolation method is used to generate a smooth control action, and when the similarity is greater than the threshold, the target strategy is directly output; the smooth control action is output to the AGC controller to execute the corresponding control strategy, and the current strategy action is updated to the smooth control action output this time.

[0015] In a preferred embodiment, an exponential interpolation method is used to generate a smooth control action, specifically: mapping the historical policy action and the target policy action into a policy embedding space, wherein the policy embedding space is constructed by a pre-trained variational autoencoder; performing embedding space interpolation through the encoder to obtain a policy embedding vector; mapping the policy embedding vector through the decoder to obtain a candidate control action; inputting the candidate control action into a preset action safety verification module to determine whether it belongs to the safe and feasible action domain, and if so, outputting it as the final smooth control action; if not, performing a minimum distance constraint projection on the candidate control action to correct it and obtain the final smooth control action.

[0016] In a preferred embodiment, frequency deviations are classified based on dynamic thresholds, specifically by: obtaining historical frequency deviation data; using an adaptive Gaussian mixture model to train the historical data to separate two independent Gaussian distributions of positive fluctuations and negative fluctuations; extracting the mean and standard deviation of positive fluctuations, and the mean and standard deviation of negative fluctuations; calculating the upper threshold of frequency deviation and the lower threshold of frequency deviation; and classifying the frequency deviation according to the upper threshold and the lower threshold.

[0017] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. By constructing a multidimensional environmental state vector and combining dynamic threshold determination, adaptive Gaussian mixture models, ARIMA time series analysis, and machine learning prediction methods, the control agent's ability to perceive and respond to the dynamics of power grid operation is comprehensively enhanced. On this basis, the agent leverages a reinforcement learning mechanism to optimize the regulation strategy based on a tuned reinforcement learning reward function, achieving adaptive decision-making and precise control of control actions. This effectively enhances the regulation's feedforward nature, stability, and multi-scenario adaptability, improving the overall control efficiency of the AGC controller's stable control.

[0018] 2. By integrating the trained control agent into the AGC controller through reinforcement learning, strategic, dynamic, and intelligent regulation of grid frequency deviations can be achieved. Unlike traditional AGC control logic based on fixed rules, this solution drives the selection of control actions through the environmental state vector and has the adaptive capability of state-action mapping. With the help of soft switching and strategy interpolation mechanisms, smooth transitions between multiple control modes can be achieved, avoiding regulation oscillations caused by sudden changes in control signals; at the same time, model predictive control (MPC) is introduced to constrain and correct the control actions generated by the agent, thereby improving the real-time and accuracy of frequency regulation while ensuring control safety. Overall, this fusion mechanism significantly enhances the robustness, sensitivity, and economy of the AGC system to load disturbances and fluctuations in renewable energy, effectively realizing the coordinated optimization of resource scheduling and frequency stability control. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A flow chart of a resource regulation method based on grid frequency deviation provided in an embodiment of the present application.

[0020] Figure 2 A schematic diagram of the structure of a resource regulation system based on grid frequency deviation provided in an embodiment of the present application.

[0021] Figure 3 This is a frequency deviation curve diagram provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0023] Example 1, Figure 1 The schematic flow chart of the resource adjustment method based on grid frequency deviation provided in the embodiment of the present application includes the following steps: S1, obtains environmental state data, constructs the environmental state vector, and initializes the control agent based on the preset control mode.

[0024] In this embodiment, the environmental state vector refers to the mathematical modeling and feature extraction of the current operating state of the power grid, enabling intelligent algorithms to understand the dynamic state of the power grid. It converts various operational data of the power grid into a structured, multi-dimensional mathematical description, enabling control strategies to make optimal decisions based on this data. The construction of the environmental state vector enables deep reinforcement learning algorithms to understand the operating state of the power grid: by constructing the environmental state vector, the reinforcement learning agent can "observe" the real-time state of the power grid and select the optimal resource control strategy accordingly.

[0025] Obtain the environment state data to construct the environment state vector, specifically: Real-time collection of frequency deviation, historical load data, photovoltaic power and wind power; Classifying the frequency deviation based on a dynamic threshold to obtain a frequency deviation category, and calculating the frequency deviation using frequency data, wherein the frequency data includes a rated frequency and a real-time frequency measurement value; Based on the ARIMA time series algorithm, the load fluctuation is obtained by monitoring the load change rate of historical load data. Based on the ARIMA time series algorithm, a load demand change model is established through historical load data, and the load change rate is calculated to monitor the load fluctuation in real time; Establish a new energy power prediction model based on machine learning through photovoltaic power and wind power, and obtain new energy output through the new energy power prediction model; The frequency deviation category, load fluctuation value and renewable energy output are encoded as an environmental state vector.

[0026] like Figure 3 As shown, the frequency deviation is classified based on the dynamic threshold to obtain the frequency deviation category, specifically: Obtain historical frequency deviation data, use the historical frequency deviation data as parameters to train an adaptive Gaussian mixture model, and obtain the mean and standard deviation of positive fluctuations and the mean and standard deviation of negative fluctuations through the trained Gaussian mixture model; The mean of positive fluctuations plus a preset multiple of the standard deviation is used as the upper threshold of the frequency deviation; The mean of negative fluctuations minus the preset multiple of standard deviation is used as the lower threshold of frequency deviation; If the frequency deviation is less than the lower threshold, it is a mild deviation; If the absolute value of the frequency deviation is greater than the lower threshold and less than the upper threshold, it is a moderate deviation; If the absolute value of the frequency deviation is greater than the upper threshold, it is a serious deviation.

[0027] It should be noted that frequency deviation data is often affected by multiple factors and may contain multiple fluctuation patterns. The adaptive Gaussian mixture model is a probability-based model that automatically identifies multiple fluctuation patterns in the data and dynamically adjusts model parameters based on the actual data distribution. The mean and standard deviation of positive and negative fluctuations trained by the adaptive Gaussian mixture model provide a more precise basis for threshold setting. This dynamic threshold setting method accurately reflects the normal fluctuation range and abnormal fluctuation boundaries of the power grid frequency based on the actual data distribution, thereby improving the accuracy of anomaly detection. Figure 3 The two horizontal lines represent the upper threshold and the lower threshold respectively.

[0028] The specific calculation formula for the frequency deviation is as follows:

[0029] In the formula, is the frequency deviation, is the real-time measurement value of frequency, is the rated frequency.

[0030] The specific calculation formula of the load demand change model is as follows:

[0031] The specific calculation formula of the load change rate is as follows:

[0032] In the formula, For the current moment The load forecast value, is the autoregressive order, is the moving average order, is the autoregressive coefficient, For the past moment The load value, is the moving average coefficient, is the error term, is the number of historical load data, is the number of error terms, is the load change rate, For the previous moment load forecast value.

[0033] It should be noted that a positive load change rate indicates an increase in load, while a negative load change rate indicates a decrease in load. The autoregressive order indicates that the model uses several past load data points to predict the current load; the moving average order indicates that the model uses several past error terms to correct the predicted value; the autoregressive coefficient measures the impact of past load data on the current load; the moving average coefficient measures the impact of past errors on the current prediction; and the error term represents the deviation between the past predicted value and the true value.

[0034] The specific calculation formula of the new energy power prediction model is as follows:

[0035] In the formula, is the new energy power prediction model, is the weight of photovoltaic power, is the photovoltaic power, is the weight of wind power, is wind power.

[0036] S2, based on the reinforcement learning framework, initializes the control agent according to the preset control mode, takes the environment state vector as input, trains the control agent through the reinforcement learning reward function, and outputs the policy action.

[0037] The method for constructing the reinforcement learning reward function includes: Setting the environment state vector as the state input of the agent, wherein the state of the agent includes frequency deviation category, load fluctuation and new energy output; The control action of the intelligent agent is defined as a preset control mode set, which includes a short-cycle control mode, a medium-cycle control mode, and a long-cycle control mode. The short-cycle control mode controls the charge and discharge power of the energy storage unit, the medium-cycle control mode controls the output of the thermal power unit, and the long-cycle control mode controls the output of the new energy unit and the demand response side unit. Construct a reinforcement learning reward function and constraints for the control action of the intelligent agent. The constraints include control cost constraints, control ability constraints, new energy utilization rate constraints, and control direction constraints. The reinforcement learning reward function includes new energy utilization, frequency deviation, control cost, and control ability.

[0038] It should be noted that by incorporating frequency deviation categories into state variables, the control agent can accurately classify disturbance scenarios, avoid a unified control strategy for all disturbances, and quickly switch control response modes, such as switching from primary frequency regulation to a secondary control strategy. By using load fluctuations and renewable energy output as state variables, the control agent can assess the availability of system control resources and the system's dynamic characteristics in real time. This allows for adaptive decision-making on the timing of control strategy switching, dynamic reconstruction of control weights, and robust control responses under abnormal conditions, resulting in greater robustness to sudden output reductions or large-scale load disturbances.

[0039] The control action reinforcement learning reward function is calculated as follows:

[0040] In the formula, is the reward value, is the utilization rate of new energy, is the frequency deviation, To control costs, For control ability, is the weight of new energy utilization rate, is the frequency deviation weight, To control the cost weight, To control the ability weight.

[0041] The control action reinforcement learning reward function constraints are as follows: The constraints on the utilization rate of new energy are as follows:

[0042] In the formula, is the utilization rate of new energy, It is the maximum value of new energy utilization rate.

[0043] The constraints for controlling costs are:

[0044] In the formula, To control costs, To minimize the cost of control.

[0045] The constraints for controlling the direction are:

[0046] In the formula, To control the direction change, For energy storage systems in time The power of the moment, The energy storage system at the last moment power.

[0047] The constraints on control capability are as follows:

[0048] In the formula, To control the capacity constraint, To adjust the power, For maximum available control capability.

[0049] It should be noted that the reinforcement learning reward function for control actions is set to comprehensively minimize the combined loss between frequency deviation, adjustment cost, and action execution. This integration can achieve the following control effects: reducing frequent AGC command switching and resource fatigue, and improving the AGC control's tolerance to the uncertainty of new energy sources.

[0050] The intelligent agent is trained to automatically select the optimal control action for each state by solving the control problem of the control mode through state, action, and reinforcement learning reward functions, thus achieving fine control. By maximizing the reward value, the intelligent agent tends to choose a control mode that is both fast and stable during training.

[0051] The control agent is trained through reinforcement learning reward function and the policy action is output, specifically: For each agent in the environment, obtain its independent state vector; Each agent performs different control actions based on its state vector; Based on the preset control action reinforcement learning reward function, the immediate reward value of each control action is calculated; With the goal of maximizing cumulative rewards, each agent is trained through a strategy optimization algorithm; Obtain a trained control agent where each state is mapped to a unique optimal control action, forming a deterministic policy action.

[0052] It should be noted that the state of the intelligent agent is obtained based on the environmental state vector. It can be understood that through the classification of frequency deviation, the state of the intelligent agent can have a slight frequency deviation, a moderate deviation, and a severe deviation. Through the load demand change model, the state of the intelligent agent can have a load increase and a load decrease. Through the new energy power prediction model, the state of the intelligent agent can have a new energy predicted output.

[0053] S3, according to the trained control agent, switches the control strategy of the AGC controller based on the soft switching and strategy interpolation mechanism, and corrects the set value of the AGC control unit based on the model predictive control method.

[0054] In this embodiment, the traditional power grid automatic generation control (AGC) usually relies on fixed rules or optimized scheduling. The intelligent agent dynamically selects the control mode, which can quickly respond to frequency deviations. The intelligent agent can work closely with the AGC controller to make the control logic of the entire power grid more reasonable and realize true intelligent scheduling. The control agent switches the control strategy of the AGC controller based on the soft switching and strategy interpolation mechanism, which can significantly improve the continuity, stability and response accuracy of the control process. Compared with the traditional hard switching method, this mechanism achieves a smooth transition of control actions through weighted interpolation of similar states, avoids system disturbances and execution shocks caused by sudden changes in control strategies, and effectively suppresses frequency oscillations and regulation lags. In addition, the mechanism can also flexibly integrate multi-strategy characteristics according to the real-time environmental status, realize coordination and optimization between different control modes, and improve the dynamic adaptability and control flexibility of the AGC controller under frequency fluctuation conditions, thereby supporting the intelligent, flexible and high-reliability control goals of power grid operation. According to the trained control agent, the control strategy of the AGC controller is switched based on the soft switching and strategy interpolation mechanism. Specifically: Match the current environment state vector with the trained control agent and output the corresponding optimal policy action, i.e. the target policy action; The similarity between historical policy actions and target policy actions is calculated based on weighted Euclidean distance. When the similarity is less than a preset threshold, an exponential interpolation method is used to generate a smooth control action. When the similarity is greater than the threshold, the target policy action is directly output to the AGC controller. The smooth control action is output to the AGC controller to execute the corresponding control strategy, and the current strategy action is updated to the smooth control action output this time.

[0055] Exponential interpolation is used to generate smooth control actions, specifically: Set historical policy actions Action with target policy Mapping into a policy embedding space, where the policy embedding space is constructed by a pre-trained variational autoencoder. The encoder and decoder jointly learn the low-dimensional embedding structure of the policy space to maintain the continuity of the policy manifold and the validity of the interpolation. Based on the adaptive interpolation factor through the encoder Perform embedding space interpolation to obtain the strategy embedding vector ; Among them, the adaptive interpolation factor represents the weight ratio between the historical policy action and the target policy action at the current moment in the policy embedding space; Embedding the policy into a vector Through the decoder Mapping to obtain candidate control actions ; The candidate control action Input to the preset motion safety verification module , judge whether it belongs to the safe and feasible action domain, if it satisfies , that is, it belongs to , then it is output as the final smooth control action. The action safety verification module is a neural network classifier. The training data comes from the control action record of the system under stable operation, which is used to determine whether the interpolation action can ensure frequency stability and power generation regulation constraints; like , that is, it does not belong to, then the minimum distance constraint projection is performed on the candidate action Make corrections, including The final safe action at the current moment, For safe action collection, The smooth control action that is finally output to the AGC controller at the current moment is the set of all control actions that satisfy the physical and operational constraints. By minimizing the distance between them, a legal control action that is closest to it is found in the known safe action set as the final output control action, so that the final output control action falls into the known safe control action set of the system, and the final smooth control action is obtained.

[0056] The specific calculation formula of the similarity is as follows:

[0057] The specific calculation formula of the strategy embedding vector is as follows:

[0058] The specific calculation formula for the candidate control action is as follows:

[0059] In the formula, is the similarity, is the target policy action, is the historical strategic action, The target policy action is The value on the dimension, For the historical strategy action The value on the dimension, For the The weight of the dimension, is the total feature dimension, is the policy embedding vector, is the adaptive interpolation factor, For the encoder, is a candidate control action, For the decoder.

[0060] Furthermore, the calculated smooth control action will be mapped to the AGC controller. Since the control action is a smooth control signal that has been interpolated and fused, it can effectively avoid frequency oscillation or system instability caused by hard switching when the control strategy is switched. The introduction of the soft switching mechanism makes the transition of the control strategy more continuous, especially when the system state is close to the switching boundary, it can avoid sudden changes in the control strategy. In this way, the system can smoothly transition between different control modes, ensuring the stability and continuity of the control process, thereby improving the overall control performance of the AGC controller. Soft switching means that when the system state changes, instead of directly performing a "hard" switch (i.e., switching to another control strategy immediately), it makes the switching process smoother through gradual adjustments. Strategy interpolation refers to the calculation of a transition control strategy between two or more control strategies through weighted averaging and interpolation function mathematical methods.

[0061] Furthermore, the AGC controller combines the decisions of the agents and issues the final control instructions. In the AGC controller, each agent has independent decision-making capabilities, and the system uses deep reinforcement learning algorithms to achieve coordinated control among the agents. The agents weigh factors such as their own control capabilities and control costs and report the optimal decision to the AGC control center. After collecting the agent's decisions, the AGC control center calculates the final control mode switching instructions, assigns them to the corresponding control unit for execution, and reports the execution status to the AGC controller.

[0062] Furthermore, the set value of the AGC control unit is corrected based on the model predictive control method.

[0063] In this embodiment, based on a control agent-assisted AGC control strategy, a model predictive control (MPC) approach is incorporated to modify the setpoints of the control units. These modified setpoints are then fed back into the AGC control strategy for precise control of the control units, enabling proactive, optimized, and constraint-consistent control of the unit's regulation behavior. MPC constructs a state-space model and utilizes a rolling-horizon optimization mechanism to predict frequency response and unit output changes over multiple future time steps. It then generates an optimal setpoint modification scheme, taking into account multiple constraints such as ramp rate limits, output constraints, and overall output balance. This modification mechanism effectively improves regulation accuracy and system stability under AGC control, avoids energy waste and equipment stress caused by overshoot and frequent adjustments, and enables dynamic and precise control of the control units. Control units refer to conventional thermal power units, natural gas turbine units, energy storage systems, and other generators or energy storage devices with frequency regulation capabilities that can dynamically adjust output according to control commands.

[0064] The model predictive control method is used to correct the set value of the AGC control unit, specifically: Establishing an MPC state space model, wherein the MPC state space model is constructed by state variables and control variables, wherein the state variables are the environment state vector and the output of each control unit, and the control variables are the set values of the control units; Based on the preset prediction time domain length, the optimization objective function and constraints of the MPC controller are set. The optimization objective function includes the system frequency deviation in the prediction time domain, the adjustment amount of the control unit, and the output deviation of the control unit from the target set value. The constraints include the control unit output constraint, the ramp rate constraint, and the total output balance constraint. Obtain the initial set value and environmental state vector of each control unit in the AGC controller, perform rolling horizon MPC optimization based on the MPC state space model and MPC controller according to the preset prediction horizon length, and obtain the revised set value of each control unit; The corrected set value is synchronously sent to the control unit to achieve the stabilization of the power grid frequency.

[0065] The specific calculation formula of the MPC state space model is as follows:

[0066]

[0067] In the formula, The state of the next moment, For the current moment The state variables, For the moment The control input variables, For the current moment The load forecast value, is the output of the system, is the state transition matrix, is the control input matrix, is the load mapping matrix, is the output mapping matrix.

[0068] The specific calculation formula of the optimization objective function is as follows:

[0069] In the formula, To optimize the objective function, To predict the time domain length, is the frequency deviation, For the The regulation amount of each control unit, To control the output deviation of the unit from the target setting value, 、 、 is the weight parameter, For time.

[0070] Control unit output constraints, ramp rate constraints, and total output balance constraints, specifically: Control unit output constraints:

[0071] Ramp rate constraint:

[0072] Total output balance constraint:

[0073] In the formula, For the The regulation amount of each control unit, is the minimum output limit of the unit, is the maximum output limit of the unit, is the time interval, is the minimum ramp rate limit of the unit, The maximum ramp rate limit of the unit. For the current moment load forecast value.

[0074] It should be noted that the state transfer matrix defines the dynamic evolution characteristics of the system under uncontrolled conditions, typically reflecting the system's internal inertia, damping, frequency recovery characteristics, and the inherent coupling relationships between various states. The control input matrix maps the impact of control input variables on the system state. The load mapping matrix indicates the intensity of the impact of load changes on the system state. The output mapping matrix maps the system state to the actual system output. The MPC state-space model can dynamically characterize the response evolution of the system's controlled variables to load disturbances and control actions.

[0075] Example 2, Figure 2 This is a schematic diagram of the structure of a resource regulation system based on grid frequency deviation provided in an embodiment of the present application, including a state setting module, an agent training module, an AGC control switching module, and an AGC controller adjustment module, with connections between the modules: The state setting module is used to obtain the environment state data and construct the environment state vector; The agent training module is used to initialize the control agent according to the preset control mode, take the environment state vector as input, train the control agent through the reinforcement learning reward function, and output the policy action; AGC control switching module is used to dynamically switch the control strategy of the AGC controller based on soft switching and strategy interpolation mechanism; The AGC controller adjustment module is used to adopt the model predictive control method to continuously optimize the power setting value of the AGC control unit.

[0076] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.

[0077] The above embodiments may be implemented in whole or in part through software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product.

[0078] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0079] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0080] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0081] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A resource regulation method based on power grid frequency deviation, characterized in that: The steps include: Obtain environmental state data and construct an environmental state vector; Based on the reinforcement learning framework, the control agent is initialized according to the preset control mode, the environment state vector is used as input, the control agent is trained through the reinforcement learning reward function, and the policy action is output; Based on the soft switching and strategy interpolation mechanism, the strategy action is dynamically switched to the control strategy of the AGC controller; The model predictive control method is used to continuously optimize the power setting value of the AGC control unit.

2. The resource adjustment method based on power grid frequency deviation according to claim 1, characterized in that: The acquisition of environmental state data to construct an environmental state vector is specifically as follows: Real-time collection of frequency deviation, historical load data, photovoltaic power and wind power; Classifying the frequency deviation based on the dynamic threshold to obtain the frequency deviation category; Use ARIMA time series model to analyze historical load data and output load fluctuation value; The power output of renewable energy can be predicted by integrating photovoltaic power and wind power through machine learning models; The frequency deviation category, load fluctuation value and renewable energy output are encoded as an environmental state vector.

3. The resource adjustment method based on power grid frequency deviation according to claim 1, characterized in that: The method for constructing the reinforcement learning reward function includes: Set the environment state vector as the agent's state input; The control action of the intelligent agent is defined as a preset control mode set, wherein the control mode set includes a short-cycle control mode, a medium-cycle control mode, and a long-cycle control mode; Construct a control action reinforcement learning reward function and constraints for the intelligent agent, wherein the constraints include control cost constraints, control capability constraints, new energy utilization rate constraints, and control direction constraints.

4. The resource adjustment method based on power grid frequency deviation according to claim 1, characterized in that: The control agent is trained by reinforcement learning reward function to output strategy actions, including: For each agent in the environment, obtain its independent state vector; Each agent performs different control actions based on its state vector; Based on the preset control action reinforcement learning reward function, the immediate reward value of each control action is calculated; With the goal of maximizing cumulative rewards, each agent is trained through a strategy optimization algorithm; Obtain a trained control agent where each state is mapped to a unique optimal control action, forming a deterministic policy action.

5. The resource adjustment method based on power grid frequency deviation according to claim 1, characterized in that: The control strategy of the AGC controller is dynamically switched based on the soft switching and strategy interpolation mechanism, specifically: Match the current environment state vector with the trained control agent and output the corresponding optimal policy action, i.e. the target policy; Calculate the similarity between the historical policy action and the target policy. When the similarity is less than the preset threshold, use the exponential interpolation method to generate a smooth control action. When the similarity is greater than the threshold, directly output the target policy. The smooth control action is output to the AGC controller to execute the corresponding control strategy, and the current strategy action is updated to the smooth control action output this time.

6. The resource adjustment method based on power grid frequency deviation according to claim 5, characterized in that: The exponential interpolation method is used to generate a smooth control action, specifically: Mapping historical policy actions and target policy actions into a policy embedding space constructed by a pre-trained variational autoencoder; Perform embedding space interpolation through the encoder to obtain the policy embedding vector; Map the policy embedding vector through the decoder to obtain candidate control actions; The candidate control action is input into the preset action safety verification module to determine whether it belongs to the safe and feasible action domain. If it does, it is output as the final smooth control action; If it does not belong to, the candidate control action is corrected by performing the minimum distance constraint projection to obtain the final smooth control action.

7. The resource adjustment method based on power grid frequency deviation according to claim 1, characterized in that: The model predictive control method is used to rollingly optimize the power setting value of the AGC control unit, specifically: Establish an MPC state space model; Set up a multi-objective optimization function within the prediction time domain, and simultaneously optimize the system frequency deviation, unit regulation, and the deviation between unit output and the target set value; Define triple constraints, including upper and lower limits on unit output, unit ramp rate limit, and real-time power balance constraints for the entire network; Obtain the real-time environmental state vector and the initial setting value, and modify the setting value sequence through rolling optimization calculation; The first correction setting value is synchronously sent to the control unit for execution.

8. The resource adjustment method based on power grid frequency deviation according to claim 2, characterized in that: The frequency deviation is classified based on the dynamic threshold, specifically: Get historical frequency deviation data; An adaptive Gaussian mixture model is used to train historical data to separate two independent Gaussian distributions of positive fluctuations and negative fluctuations; Extract the mean and standard deviation of positive fluctuations, and the mean and standard deviation of negative fluctuations; Calculate the frequency deviation upper threshold and the frequency deviation lower threshold; Frequency deviations are classified based on upper and lower thresholds.

9. A system using the resource regulation method based on power grid frequency deviation according to any one of claims 1 to 8, characterized in that: It includes the state setting module, the agent training module, the AGC control switching module, and the AGC controller adjustment module. There are connections between the modules: The state setting module is used to obtain the environment state data and construct the environment state vector; The agent training module is used to initialize the control agent according to the preset control mode, take the environment state vector as input, train the control agent through the reinforcement learning reward function, and output the policy action; AGC control switching module is used to dynamically switch the control strategy of the AGC controller based on soft switching and strategy interpolation mechanism; The AGC controller adjustment module is used to adopt the model predictive control method to continuously optimize the power setting value of the AGC control unit.

Citation Information

Patent Citations

  • AGC dynamic optimization control method based on multiple agents

    CN117767422A

  • Intelligent power grid voltage control method based on quantum multi-agent reinforcement learning

    CN118868113A

  • Energy storage auxiliary thermal power generating unit deep reinforcement learning load frequency control method

    CN119051070A

  • Energy storage and multi-resource optimization scheduling method and system, electronic equipment and storage medium

    CN119398389A

  • Power grid frequency cooperative control method and system

    CN119518833A

Cited By

  • Low-voltage distributed photovoltaic group scheduling and group control method and system based on multi-source information fusion

    CN120749912A

  • Hydropower station AGC collaborative optimization method and device based on deep learning load prediction and model prediction control

    CN121367187A

  • Resonant sensor reading adaptive optimization method and system based on reinforcement learning

    CN121388395A

  • Power system load frequency adaptive regulation and control system and method

    CN121642981A