New energy AGC intelligent agent decision rule optimization method

By optimizing the decision-making rules of the new energy AGC intelligent body through deep neural networks and reinforcement learning, the problem of inaccurate response of the new energy AGC system to complex power grid faults is solved, and the stability of the power grid and the adaptability of new energy power generation are improved.

CN120742856AActive Publication Date: 2025-10-03YUNNAN POWER GRID CO LTD +1

Patent Information

Application Number
CN202511220796.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-10-03
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

Existing new energy AGC systems have difficulty responding accurately and comprehensively when simulating complex power grid faults, and are unable to effectively adapt to the uncertainty of power grid operating conditions and the intermittent nature of new energy power generation, affecting the stability and reliability of the power grid.

Method used

Deep neural networks and reinforcement learning methods are used to optimize the decision-making rules of the new energy AGC intelligent agent. By collecting historical power grid fault data and new energy AGC control data, the decision-making model of the intelligent agent is trained using supervised learning and deep Q-learning algorithms to optimize its decision-making strategy in complex fault scenarios.

Benefits of technology

It improves the accuracy of control effect evaluation of the new energy AGC system in complex fault scenarios, enhances the grid's adaptability to fluctuations in new energy power generation, and ensures the stable operation of the grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120742856A_ABST
    Figure CN120742856A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a new energy AGC intelligent agent decision rule optimization method, and relates to the technical field of power system testing, and the method comprises the steps: collecting historical power grid fault data, new energy AGC control data and simulation test data; preprocessing the collected data to obtain preprocessed data; based on the preprocessed data, a supervised learning method is adopted to carry out optimization training on a decision rule model of an intelligent agent, a deep neural network is used to predict the optimal AGC unit power adjustment amount, a mean square error is used as a loss function to carry out model optimization, and an optimized intelligent agent decision rule model is obtained; according to the method, the prediction precision of the model and the operation efficiency of the power grid are improved by optimizing the decision rule of the intelligent agent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power system testing, and in particular to a new energy AGC intelligent body decision rule optimization method. Background Art

[0002] As global demand for clean energy continues to grow, the proportion of renewable energy in the power system is increasing. Automatic Generation Control (AGC) is a key technology for ensuring the reliable integration of renewable energy into the grid and maintaining stable power system operation, playing an increasingly important role. In practical applications, AGC is widely used in various renewable energy sites, such as wind farms and photovoltaic power stations. Through the AGC system, renewable energy sites can adjust power generation in real time based on grid demand, maintaining power balance and frequency stability in the power system.

[0003] However, current simulations of renewable energy AGC systems' responses to complex grid faults are limited, making it difficult to accurately and comprehensively simulate a wide range of potential grid faults and operating conditions. Therefore, a new approach is needed to improve the accuracy and efficiency of simulation testing of renewable energy AGC systems while optimizing the decision-making rules of intelligent agents to better adapt to the uncertainties of grid operation and the intermittent nature of renewable energy generation. This will help improve the grid's adaptability to fluctuations in renewable energy generation, enhancing its stability and reliability. Summary of the Invention

[0004] The main purpose of the present invention is to provide a new energy AGC intelligent agent decision rule optimization method, which can improve the accuracy of the control effect evaluation of the new energy AGC in complex fault scenarios.

[0005] To achieve the above objectives, the present application provides, in a first aspect, a method for optimizing decision rules of a new energy AGC agent, the method comprising: Collect historical power grid fault data, new energy AGC control data and simulation test data; Preprocessing the collected data to obtain preprocessed data; Based on the preprocessed data, a supervised learning method is used to optimize and train the decision rule model of the intelligent agent, wherein a deep neural network is used to predict the optimal AGC unit power adjustment amount, and the mean square error is used as the loss function to optimize the model to obtain an optimized intelligent agent decision rule model.

[0006] Optionally, the method further includes: Reinforcement learning is used to further optimize the adaptive decision-making ability of the intelligent agent. By setting the state space, action space and reward function, the deep Q learning algorithm is used to enable the intelligent agent to iteratively optimize the decision-making strategy through Q value.

[0007] Optionally, the state space is set to:

[0008] in, is the grid frequency deviation, is the excess power, is the renewable energy power variation, V bus is the key node voltage; The action space is set as:

[0009] in, is the active power regulation of the AGC unit, is the reactive power regulation quantity; The reward function is set as:

[0010] in: α, β, γ is the weight coefficient; is the rated voltage.

[0011] Optionally, the decision strategy is optimized through Q-value iteration using the following formula:

[0012] in: m is the learning rate; q is the discount factor; Q ( s,a ) is the Q value of the current state-action pair; Q t+1 ( s , a ) is the new Q value corrected after iteration; r For immediate rewards.

[0013] Optionally, the historical power grid fault data includes power flow exceeding limit at the transmission section, power grid frequency deviation, and voltage fluctuation; The new energy AGC control data includes wind power, photovoltaic output data, AGC unit response data, and load changes; The simulation data includes data related to the interactive behaviors of the new energy intelligent agent, the AGC unit intelligent agent, and the transmission section intelligent agent under different fault scenarios.

[0014] Optionally, the preprocessing includes: Normalization, outlier removal, and time series resampling.

[0015] Optionally, the method further includes: Using the optimized agent decision rule model as the decision rule of the agent in the simulation model, performing a new energy AGC active power control joint simulation test based on the simulation model to obtain simulation results; The simulation results are compared and verified with actual power grid monitoring data or theoretical analysis results; if there is a deviation, the simulation model and algorithm are calibrated.

[0016] A second aspect of the present application provides a new energy AGC intelligent agent decision rule optimization device, comprising: Data acquisition module, used to collect historical power grid fault data, new energy AGC control data and simulation test data; A preprocessing module is used to preprocess the collected data to obtain preprocessed data; A training optimization module is used to optimize and train the decision rule model of the intelligent agent based on the preprocessed data using a supervised learning method, wherein a deep neural network is used to predict the optimal AGC unit power adjustment amount, and the mean square error is used as the loss function to optimize the model to obtain an optimized intelligent agent decision rule model.

[0017] A third aspect of the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the first aspect and any possible implementation thereof.

[0018] The present application provides a new energy AGC intelligent agent decision rule optimization method, which collects historical power grid fault data, new energy AGC control data and simulation test data; preprocesses the collected data to obtain preprocessed data; based on the preprocessed data, uses a supervised learning method to optimize and train the intelligent agent's decision rule model, wherein a deep neural network is used to predict the optimal AGC unit power adjustment amount, and the mean square error is used as the loss function to optimize the model to obtain an optimized intelligent agent decision rule model; this method optimizes the decision rule of the intelligent agent to make the simulation results closer to the actual situation, thereby improving the accuracy of the control effect evaluation of the new energy AGC in complex fault scenarios, which not only improves the prediction accuracy of the model and the operating efficiency of the power grid, but also enhances the power grid's adaptability to fluctuations in new energy power generation, thereby providing strong technical support for the stable operation of the power grid. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] in: Figure 1 A flowchart of a new energy AGC agent decision rule optimization method provided in an embodiment of the present application; Figure 2 A schematic diagram of the structure of a new energy AGC intelligent agent decision rule optimization device provided in an embodiment of the present application; Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0022] The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0023] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0024] Some of the terms and concepts involved in the embodiments of this application or related background are explained as follows: Automatic Generation Control (AGC): Within a defined area, when the power system frequency or tie-line power changes, the active power of generator sets is adjusted through automatic control programs to maintain system frequency or ensure predetermined power exchange between regions. Its technical equipment system primarily includes grid operation control systems at all levels of dispatching organizations, telecontrol transmission channels, remote terminal equipment or computer monitoring systems at power plants, generator set coordination and control systems, generator sets and their active power regulation devices, and application software that implements AGC functions. For regulation and control involving renewable energy AGC, under normal circumstances, renewable energy stations generate power at maximum capacity. When the grid experiences cross-section limits, peak-shaving difficulties, insufficient reserve, or frequency fluctuations, the renewable energy AGC automatically adjusts the output power of renewable energy generation equipment (such as photovoltaic inverters) to meet the grid's security and stability requirements.

[0025] Power Conditioning Controller (PLC): Based on the instructions of the AGC system, the PLC can accurately control the output power of renewable energy power generation equipment. It calculates the difference between the total power delivered and the sum of the current power generation of each device as the power to be adjusted. Then, it prioritizes each device based on its adjustable power ratio and allocates the power to be adjusted to each device. This process ensures that renewable energy power generation equipment can achieve efficient and stable operation while meeting grid demand. In the new energy AGC system, the PLC can also integrate optimization algorithms and strategies, such as proportional allocation and priority sorting, to achieve more accurate and efficient power regulation. These algorithms and strategies can fully consider the characteristics of renewable energy power generation equipment and grid requirements, ensuring the system can maintain stable operation under complex and changing operating conditions.

[0026] Principle of PLC equivalence for local dispatching: Provincial and local coordinated control should support the equivalence of a single local dispatching system into multiple virtual machines according to group objects, grid topology nodes, and section partitions, so as to achieve refined control of photovoltaic / wind farms and related sections on the central dispatching side. The principles for PLC equivalence of local dispatching new energy farms in the central dispatching master station are as follows: 1) The geostationary stations whose output has an impact on the same provincial and local coordinated section can be equivalent to one PLC, and other new energy stations that are not affected by the provincial and local coordinated section can be equivalent to one PLC; 2) A new energy station connected to the same 220kV node can be equivalent to a PLC; 3) Each prefecture-level new energy station is equivalent to only one PLC.

[0027] 4. Provincial-Local Coordination: An effective coordination mechanism should be established between provincial and local grid dispatchers. Generally, provincial dispatchers formulate overall control strategies and targets based on the overall power supply and demand situation and renewable energy generation forecasts, and issue them to local dispatchers. Local dispatchers then formulate specific control plans based on the provincial dispatchers' control strategies and targets, combined with local renewable energy generation conditions and grid structure. This coordinated provincial-local control strategy optimizes the dispatch and control of renewable energy generation. This not only improves the security and stability of the power grid, but also enhances the utilization rate and power generation efficiency of renewable energy.

[0028] The following mentioned in the embodiments of this application describes the embodiments of this application in conjunction with the drawings in the embodiments of this application.

[0029] Figure 1 A schematic diagram of a new energy AGC agent decision rule optimization method provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the method includes: 101. Collect historical power grid fault data, new energy AGC control data and simulation test data.

[0030] The execution subject of the method in the embodiment of the present application can be a new energy AGC intelligent body decision rule optimization device, which can be implemented on an electronic device in practical applications.

[0031] In a multi-agent modeling approach, the decision-making rules of each agent need to be adaptive to cope with the uncertainty of complex grid failures and fluctuations in renewable energy output. Therefore, in the embodiments of the present application, a machine learning algorithm can be used to optimize and train the agent, enabling it to learn the optimal decision-making strategy based on historical grid failure data and renewable energy AGC control data. The training goal is to enable the agent to dynamically adjust its strategy based on existing experience and real-time data when faced with different fault conditions, thereby improving the response speed, regulation accuracy, and grid stability of the renewable energy AGC system.

[0032] Specifically, the optimization training of the intelligent agent includes data collection and model training.

[0033] In this data collection step, the data obtained can be used as sample data for training, mainly including historical power grid fault data, new energy AGC control data and simulation test data.

[0034] Optionally, the above-mentioned historical power grid fault data includes power flow exceeding limit of transmission section, power grid frequency deviation, and voltage fluctuation; The above-mentioned new energy AGC control data includes wind power, photovoltaic output data, AGC unit response data, and load changes; The above simulation data includes data related to the interactive behavior of the new energy intelligent agent, the AGC unit intelligent agent, and the transmission section intelligent agent under different fault scenarios.

[0035] Specifically, historical power grid fault data can mainly include transmission section power flow exceeding limits (fault time, section number, active power flow, reactive power flow, and exceeding limits), power grid frequency deviation (frequency changes before and after the fault), and voltage fluctuations (voltage changes at key nodes). New energy AGC control data, which mainly includes wind power and photovoltaic output data (real-time power, wind speed, and light intensity), AGC unit response data (regulation rate, regulation amplitude, and control delay), and load changes (system load forecast and actual load data). The simulation data may mainly include data related to the interactive behaviors of the new energy intelligent agent, the AGC unit intelligent agent, and the transmission section intelligent agent under different fault scenarios.

[0036] 102. Preprocess the collected data to obtain preprocessed data.

[0037] Optionally, the above preprocessing includes: Normalization, outlier removal, and time series resampling.

[0038] The preprocessing mentioned in the embodiments of this application refers to a series of processing steps performed on the collected data before using a machine learning algorithm to train the decision rule model of the intelligent agent. The purpose of preprocessing is to improve the quality of the data, making it more suitable for model training and improving the performance and accuracy of the model. The appropriate data preprocessing method can be selected and adjusted as needed.

[0039] The main normalization method can be Min-Max normalization:

[0040] in: X is the original data; 、 are the minimum and maximum values ​​of the data respectively; X′ The data are normalized.

[0041] 103. Based on the above preprocessed data, the supervised learning method is used to optimize the training of the decision rule model of the intelligent agent. In particular, a deep neural network is used to predict the optimal AGC unit power adjustment amount, and the mean square error is used as the loss function to optimize the model to obtain the optimized intelligent agent decision rule model.

[0042] Supervised learning, as mentioned in the examples of this application, is a method in machine learning that involves learning a model from labeled training data in order to predict outputs for unseen data. In supervised learning, each training example consists of input data (features) and a corresponding output label (target value). The goal of the model is to learn the relationship between input features and output labels so that it can make accurate predictions for new, unseen data.

[0043] For some tasks of new energy AGC control (such as power adjustment decisions of AGC units), supervised learning methods are used for training so that it can learn to make appropriate power adjustment strategies under different fault conditions.

[0044] The deep neural network (DNN) mentioned in the embodiments of this application is a type of artificial neural network with multiple hidden layers that can learn complex patterns and relationships in data. By increasing the depth of the network (i.e., the number of hidden layers), deep neural networks can capture more abstract data features, thereby achieving excellent performance in many tasks.

[0045] Specifically, a deep neural network (DNN) may be used in the embodiments of the present application, the goal of which is to predict the optimal AGC unit power adjustment:

[0046] Where: Input X includes grid frequency deviation , Changes in new energy output , current exceeding limit etc.; output It is the power adjustment of the AGC unit.

[0047] The loss function of the neural network can be the mean square error (MSE):

[0048] in: is the target output in the training data; is the adjusted power predicted by the model; N is the total number of samples.

[0049] In an optional embodiment, the above method further includes: Reinforcement learning is used to further optimize the adaptive decision-making ability of the intelligent agent. By setting the state space, action space and reward function, the deep Q learning algorithm is used to enable the intelligent agent to iteratively optimize the decision-making strategy through Q value.

[0050] Reinforcement learning (RL), mentioned in the examples of this application, is a method in machine learning that learns how to make decisions through interaction with an environment. In RL, an agent takes actions (actions) in a given environment and receives feedback (rewards) based on these actions, with the goal of maximizing cumulative rewards. This method is particularly useful in tasks that require sequential decision-making, such as gaming, robotic control, and resource management.

[0051] When dealing with complex environments such as renewable energy fluctuations and sudden failures, intelligent agents need to possess adaptive decision-making capabilities. Reinforcement learning can be used for optimization, enabling agents to learn optimal decision-making strategies through trial and error. The core of reinforcement learning is that agents execute actions in an environment and adjust their strategies based on reward feedback.

[0052] In one embodiment, the states, actions, and rewards may be set as follows: State space S:

[0053] in, is the grid frequency deviation, is the excess power, Renewable energy power change refers to the change in power generation from renewable energy sources such as wind and solar energy; V bus is the key node voltage.

[0054] Action space A:

[0055] in, is the active power regulation of the AGC unit, It is the reactive power regulation quantity.

[0056] Reward function R:

[0057] in: α, β, γ is the weight coefficient; is the rated voltage.

[0058] Deep Q Learning (DQN) mentioned in the embodiments of this application is a technology that combines deep learning and reinforcement learning. It selects the optimal action by estimating the value of each action. During the training process, the agent obtains reward signals by interacting with the environment and updates its estimate of the value of the action based on these signals. By repeating this process continuously, the agent can gradually learn how to make the best decisions in different situations. The core idea of ​​deep Q learning is to construct a deep neural network as the decision-making model of the agent. The input of this neural network is the agent's observation of the environment, and the output is the value estimate of each action. During the training process, we use a reinforcement learning algorithm to update the parameters of the neural network so that the agent can better estimate the value of the action.

[0059] Specifically, deep Q learning can be used to train the agent so that it can iteratively optimize its decision-making strategy through Q-values:

[0060] in: m is the learning rate; q is the discount factor; Q ( s,a ) is the Q value of the current state-action pair; Q t+1 ( s , a ) is the new Q value corrected after iteration; r For immediate rewards.

[0061] In the above formula s Represents the "state space" of the intelligent agent. In the power grid simulation scenario of this application, s Can include grid frequency deviation f 2. Transmission section current exceeds the limit P , Changes in new energy output P N , key node voltage V bus and other key information, which comprehensively reflects the current operating status of the power grid; a Represents the "action space" that the intelligent agent can perform, corresponding to the power grid scenario, a Can include active power regulation of AGC units P adjust , reactive power regulation Q adjust , that is, the specific actions that the agent can take to deal with the current state.

[0062] Q ( s , a) is the “state-action value function”, which represents the agent’s current state s Next, execute the action a The expectation of the long-term cumulative rewards that can be obtained after the investment. Q ( s , a ) value is higher, it means that under the current power grid state, the action is taken a , the more beneficial it is to maintain grid stability and optimize AGC control effects in the future.

[0063] Q ( s' , a' ) indicates that when the agent enters a new state s' Then, for all possible actions a' For example, when the power grid status changes due to frequency deviation or section flow changes s' After that, the agent determines the action to take a' That is, the benefit that can be obtained after AGC performs power regulation.

[0064] Q ( s' , a' ) states and actions are "potential", s' Is to perform an action a The state that will appear later, a' is s 'All possible actions that can be selected in the state. max Q ( s' , a' ) is in the state s' The maximum Q value among all possible actions.

[0065] During the training process of deep Q learning, the agent continuously interacts with the simulation environment and updates the Q ( s , a ), where r Is to perform an action a If the frequency deviation decreases after the action, a positive reward will be obtained; if the cross-section exceeds the limit, a negative reward will be obtained. s' Is to perform an action a The new state entered later, max Q ( s' , a' ) is in the new state s' The maximum Q value among all possible actions, m is the learning rate (controls the magnitude of each update), q is the discount factor (weighing the importance of immediate rewards and future rewards). Through such iterative updates, Q( s , a ) can gradually approach the optimal value, allowing the intelligent agent to continuously learn the optimal decision-making strategy in complex power grid fault scenarios. For example, when multiple faults are superimposed, it can accurately select the adjustment amount of the AGC unit to achieve stable control of the power grid.

[0066] The reference values ​​of each weight in the above reward function R can be determined according to the process of theoretical derivation, testing and verification, and continuous iteration. Specifically: 1. Theoretical derivation (1) Determine the "influence weight" of each indicator on the power grid According to the goal of "stabilizing frequency, controlling power flow, and maintaining voltage" of the power grid, key indicators can be initially prioritized. For example, frequency collapse will directly cause large-scale power outages and need to be controlled first, so a larger weight can be set. Power flow exceeding the limit may cause equipment overload and line tripping, so it can be set as a lower priority. Based on the deviation limit of each indicator, the theoretical weight can be initially set as α : β : γ =0.5:0.2:0.1=5:2:1.

[0067] (2) Combined with "control sensitivity" correction Use small disturbance stability analysis to calculate the sensitivity of AGC control to various indicators. For example, if the sensitivity of AGC to frequency is S f =0.8Hz / MW, meaning that increasing the output by 1MW results in a 0.8Hz increase in frequency. Through calculation and testing, the adjustment sensitivity of each indicator is determined. The higher the sensitivity, the easier it is for the AGC to control that indicator, and the weight can be appropriately reduced. This method is used to adjust the weight.

[0068] 2. Test and Verification After initially obtaining the theoretically calculated weights, different weight combinations can be set in the simulation test to observe the AGC control effect, so as to judge the error in the weight setting and make timely adjustments. α If it is too small, the frequency limit violation in the simulation cannot be suppressed in time, and the fault loss is large, so the corresponding weight should be appropriately increased.

[0069] Through multi-scenario testing, the algorithm is used to optimize the weights to minimize the expected failure loss and ultimately determine the available engineering α 、 β 、 γ .

[0070] 3. Continuous Iteration Regular recalibration is required as the grid topology and the proportion of renewable energy access change, such as annual updates or re-optimization when new large-scale renewable energy sites are added.

[0071] Further optionally, after the above step 103, the method further includes: The optimized agent decision rule model is used as the decision rule of the agent in the simulation model. Based on the simulation model, a joint simulation test of new energy AGC active power control is carried out to obtain simulation results. The above simulation results are compared and verified with actual power grid monitoring data or theoretical analysis results; if there are deviations, the above simulation models and algorithms are calibrated.

[0072] Specifically, after obtaining the optimized intelligent agent decision rule model, it can be used in a simulation model to conduct a joint simulation test of the new energy AGC active control. Among them, the simulation model can be obtained based on a multi-agent modeling method, and the transmission section exceeding the limit, the new energy power generation equipment affected by changes in natural conditions, and the new energy AGC system are regarded as independent intelligent agents. Each intelligent agent is given a clear function, initial state and decision rule, so that it can make real-time decisions and adjust its behavior based on its own state and other intelligent agent information, accurately simulate the dynamic interaction of various factors in complex scenarios, and thus improve the test accuracy. The intelligent agent decision rule used in the simulation can be implemented by the optimization model in the embodiment of the present application. By optimizing the intelligent agent decision rule, it can make more realistic decisions when facing different situations, further improving the accuracy and efficiency of the test.

[0073] In the process of setting and simulating complex fault scenarios, it is necessary to consider a variety of grid equipment and their interactions, including different factors such as transmission sections, wind farms, photovoltaic power stations and other modules, while introducing extreme changes in natural conditions and severe fluctuations in grid frequency and voltage. Through simulation, it is possible to dynamically simulate these fault scenarios, perform fault injection and optimize scheduling decisions. When setting specific fault scenarios, the first thing to consider is the transmission section exceeding the limit. To simulate this fault, by increasing or decreasing the load or adjusting the power output at one end of the transmission section, the fluctuation of the system load is simulated, thereby triggering the transmission section exceeding the limit. In this case, the specific method of increasing or decreasing the load can be defined by the following algorithm:

[0074] in, is the load change, is the load adjustment factor, which indicates the ratio of load increase or decrease. This adjustment simulates various load fluctuations and simulates power overload or power shortage scenarios by adjusting power output. This process requires real-time monitoring of the status of each transmission section of the grid, and the use of an agent-based model to determine the range of load fluctuations to ensure dynamic system response and stability.

[0075] In the embodiments of the present application, after simulating a complex fault scenario, the simulation results can be compared and verified with monitoring data of similar faults in the actual power grid or authoritative theoretical analysis results. If there is a large deviation, the simulation results can be made closer to the actual situation by adjusting the interaction coefficients of various factors in the model and correcting the parameters of the simulation algorithm. At the same time, a sensitivity analysis method can be used to determine the factors and parameters that have a greater impact on the simulation results. These key factors and parameters can be adjusted and calibrated first to improve calibration efficiency, thereby improving the accuracy of the control effect evaluation of the new energy AGC in complex fault scenarios.

[0076] Optionally, after adjusting and calibrating the parameters, the simulation test may be re-performed and the above steps may be repeated until the deviation between the simulation result and the actual data is within an acceptable range or a predetermined accuracy requirement is met.

[0077] In the embodiment of the present application, the simulation results are compared with the monitoring data of the actual power grid or the theoretical analysis results through the data acquisition step to evaluate the control effect of the new energy AGC system in complex fault scenarios. Through comparative analysis, the differences between the simulation model and the actual system can be revealed, thereby providing a basis for subsequent model adjustment and algorithm optimization. The main contents of the comparison include the active control response characteristics of the new energy AGC system, the regulation effect of the grid frequency and voltage, and the control of the transmission section flow. The verification of these factors can ensure the accuracy of the simulation model and make adjustments when necessary to improve its consistency with the actual power grid operation. The above comparative analysis may specifically include: During the comparison, if there is a large deviation between the simulation results and the actual or theoretical analysis data, the model needs to be calibrated. The goal of model calibration is to make the simulation results more consistent with the actual situation, thereby improving the accuracy of the evaluation of the control effect of the renewable energy AGC. During the calibration process, the interaction coefficients of the various factors in the model can be adjusted first, such as changing the influence weights between different fault agents and the renewable energy AGC agent. By adjusting these weights, the interaction between the agents can be changed, thereby optimizing the accuracy of the simulation results. For example, if the extreme wind speed changes in the wind farm have too great an impact on the grid frequency, the regulation effect can be optimized by adjusting the weight of the wind speed change on the response of the renewable energy AGC system.

[0078] Furthermore, parameters within the simulation algorithm can be modified, such as adjusting the threshold and step size within the agent's decision-making model. These parameters determine the agent's decision sensitivity and adjustment speed during the simulation. For example, if the power regulation deadband set in the agent's regulation rules is too large, the new energy AGC system may not be sufficiently responsive to grid frequency fluctuations. In this case, reducing the deadband can enhance sensitivity to frequency fluctuations, thereby improving control accuracy.

[0079] Through these adjustments, the deviation between the simulation results and the actual grid operation data can be effectively reduced, and the accuracy of the model can be improved. For example, during the simulation process, if the power regulation of the new energy AGC system under extreme wind speed conditions is not timely enough, adjusting the step size and regulation rate in the intelligent agent decision rule can make the regulation faster and avoid excessive fluctuations in the grid frequency. In this way, not only can the control effect be improved, but also more accurate decision support can be provided for fault response in actual operation of the system. Ultimately, the calibrated simulation model can better simulate the response of the power grid in complex fault scenarios, thereby optimizing the performance of the new energy AGC system and ensuring the stable operation of the power grid.

[0080] The method in the embodiment of the present application collects historical power grid fault data, new energy AGC control data and simulation test data; preprocesses the collected data to obtain preprocessed data; based on the preprocessed data, a supervised learning method is used to optimize and train the decision rule model of the intelligent agent, wherein a deep neural network is used to predict the optimal AGC unit power adjustment amount, and the mean square error is used as the loss function to optimize the model to obtain an optimized intelligent agent decision rule model; this method optimizes the decision rule of the intelligent agent to make the simulation results closer to the actual situation, thereby improving the accuracy of the control effect evaluation of the new energy AGC in complex fault scenarios, which not only improves the prediction accuracy of the model and the operating efficiency of the power grid, but also enhances the adaptability of the power grid to fluctuations in new energy power generation, thereby providing strong technical support for the stable operation of the power grid.

[0081] Based on the description of the aforementioned method embodiment, the embodiment of the present application also provides a new energy AGC intelligent body decision rule optimization device.

[0082] Figure 2 This is a schematic diagram of the structure of a new energy AGC intelligent body decision rule optimization device provided in an embodiment of the present application. Figure 2 As shown, the new energy AGC intelligent agent decision rule optimization device 200 includes: Data acquisition module 210, used to collect historical power grid fault data, new energy AGC control data and simulation test data; The preprocessing module 220 is used to preprocess the collected data to obtain preprocessed data; The training optimization module 230 is used to optimize the training of the decision rule model of the intelligent agent based on the preprocessed data using a supervised learning method, wherein a deep neural network is used to predict the optimal AGC unit power adjustment amount, and the mean square error is used as the loss function to optimize the model to obtain an optimized intelligent agent decision rule model.

[0083] Optionally, the above-mentioned training optimization module 230 is also used to: use reinforcement learning to further optimize the adaptive decision-making ability of the intelligent agent, wherein, by setting the state space, action space and reward function, the deep Q learning algorithm is used to enable the intelligent agent to iteratively optimize the decision-making strategy through Q value.

[0084] Optionally, the above state space is set to:

[0085] in, is the grid frequency deviation, is the excess power, is the renewable energy power variation, V bus is the key node voltage; The above action space is set as:

[0086] in, is the active power regulation of the AGC unit, is the reactive power regulation quantity; The above reward function is set as:

[0087] in: α, β, γ is the weight coefficient; is the rated voltage.

[0088] Optionally, the above-mentioned Q-value iterative optimization decision strategy adopts the following formula:

[0089] in: m is the learning rate; q is the discount factor; Q ( s,a ) is the Q value of the current state-action pair; Q t+1 ( s , a ) is the new Q value corrected after iteration; r For immediate rewards.

[0090] Optionally, the above-mentioned historical power grid fault data includes power flow exceeding limit of transmission section, power grid frequency deviation, and voltage fluctuation; The above-mentioned new energy AGC control data includes wind power, photovoltaic output data, AGC unit response data, and load changes; The above simulation data includes data related to the interactive behavior of the new energy intelligent agent, the AGC unit intelligent agent, and the transmission section intelligent agent under different fault scenarios.

[0091] Optionally, the above preprocessing includes: Normalization, outlier removal, and time series resampling.

[0092] Optionally, also include: A simulation execution module 240 is configured to use the optimized agent decision rule model as the decision rule of the agent in the simulation model, perform a new energy AGC active power control joint simulation test based on the simulation model, and obtain simulation results; The verification and calibration module 250 is used to compare and verify the above simulation results with actual power grid monitoring data or theoretical analysis results; if there is a deviation, the above simulation model and algorithm are calibrated.

[0093] Understandably, Figure 2 The relevant contents of each module in the above method embodiment have been described in detail, and the details can be referred to the contents of the method embodiment; Figure 2 The provided new energy AGC intelligent agent decision rule optimization device 200 can perform the following Figure 1 Any steps in the illustrated embodiment will not be described in detail here.

[0094] In one embodiment of the present application, an electronic device is also provided. Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 3 As shown, the electronic device 300 includes a processor 301 and a memory 302. The memory 302 stores a computer program. When the computer program is executed by the processor 301, the following operations are performed: Figure 1 The electronic device 300 may further include an input / output device, etc. In a specific embodiment, the electronic device may be a terminal device, etc.

[0095] In one embodiment, a computer-readable storage medium is further provided. The computer-readable storage medium stores a computer program. When the computer program is executed by the processor 301, the processor 301 executes any step in the above method embodiment.

[0096] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0097] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0098] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A new energy AGC agent decision rule optimization method, characterized in that: The method comprises: Collect historical power grid fault data, new energy AGC control data and simulation test data; Preprocessing the collected data to obtain preprocessed data; Based on the preprocessed data, a supervised learning method is used to optimize and train the decision rule model of the intelligent agent, wherein a deep neural network is used to predict the optimal AGC unit power adjustment amount, and the mean square error is used as the loss function to optimize the model to obtain an optimized intelligent agent decision rule model.

2. The new energy AGC agent decision rule optimization method according to claim 1 is characterized in that: The method further comprises: Reinforcement learning is used to further optimize the adaptive decision-making ability of the intelligent agent. By setting the state space, action space and reward function, the deep Q learning algorithm is used to enable the intelligent agent to iteratively optimize the decision-making strategy through Q value.

3. The new energy AGC agent decision rule optimization method according to claim 2 is characterized in that: The state space is set as: in, is the grid frequency deviation, is the excess power, is the renewable energy power variation, V bus is the key node voltage; The action space is set as: in, is the active power regulation of the AGC unit, is the reactive power regulation quantity; The reward function is set as: in: α, β, γ is the weight coefficient; is the rated voltage.

4. The new energy AGC agent decision rule optimization method according to claim 2 is characterized in that: The Q-value iterative optimization decision strategy adopts the following formula: in: m is the learning rate; q is the discount factor; Q t ( s,a ) is the Q value of the current state-action pair; Q t+1 ( s , a ) is the new Q value corrected after iteration; r is the immediate reward.

5. The new energy AGC agent decision rule optimization method according to claim 1 is characterized in that: The historical power grid fault data includes power transmission section power flow exceeding limit, power grid frequency deviation, and voltage fluctuation; The new energy AGC control data includes wind power, photovoltaic output data, AGC unit response data, and load changes; The simulation data includes data related to the interactive behaviors of the new energy intelligent agent, the AGC unit intelligent agent, and the transmission section intelligent agent under different fault scenarios.

6. The new energy AGC agent decision rule optimization method according to claim 1 is characterized in that: The pretreatment includes: Normalization, outlier removal, and time series resampling.

7. The new energy AGC agent decision rule optimization method according to claim 1 is characterized in that: The method further comprises: Using the optimized agent decision rule model as the decision rule of the agent in the simulation model, performing a new energy AGC active power control joint simulation test based on the simulation model to obtain simulation results; The simulation results are compared and verified with actual power grid monitoring data or theoretical analysis results; if there is a deviation, the simulation model and algorithm are calibrated.

8. A new energy AGC intelligent agent decision rule optimization device, characterized in that: include: Data acquisition module, used to collect historical power grid fault data, new energy AGC control data and simulation test data; A preprocessing module is used to preprocess the collected data to obtain preprocessed data; A training optimization module is used to optimize and train the decision rule model of the intelligent agent based on the preprocessed data using a supervised learning method, wherein a deep neural network is used to predict the optimal AGC unit power adjustment amount, and the mean square error is used as the loss function to optimize the model to obtain an optimized intelligent agent decision rule model.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that When the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • AGC unit dynamic optimization method based on deep reinforcement learning

    CN112186811A

  • Power grid multi-section power automatic control method based on distributed multi-agent reinforcement learning

    CN112615379A

  • Multi-mobile emergency power supply toughness optimization scheduling method based on data driving

    CN118449131A

  • AGC performance evaluation method and system in approximate follow-up control operation mode

    CN119596904A

  • Power grid real-time scheduling optimization method and system, computer device and storage medium

    US20250210996A1

Cited By

  • Microgrid demand side management method and system based on group enhancement strategy optimization

    CN121192715A