AGC adaptive cooperative control method based on multi-target driving

By introducing a multi-objective dynamic weight adaptive mechanism, time-varying network collaborative scheduling, and multi-source information fusion technology, combined with two-layer game-driven and online reinforcement learning, the collaborative control problem of the AGC system in complex power environments was solved, achieving adaptive optimization of frequency stability, economical operation, and equipment health.

CN121923162APending Publication Date: 2026-04-24THREE GORGES JINSHAJIANG CHUANYUN HYDROPOWER DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THREE GORGES JINSHAJIANG CHUANYUN HYDROPOWER DEV CO LTD
Filing Date
2025-12-15
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing AGC control methods are difficult to effectively cope with complex, dynamic and uncertain power system environments. They suffer from response lag, control rigidity and local optima, and cannot achieve dynamic weight allocation and coordinated scheduling among multiple objectives.

Method used

By employing a multi-objective dynamic weight adaptive determination mechanism, time-varying complex interconnected network collaborative scheduling, multi-source information fusion and deep identification technology, two-layer adaptive game-driven approach, and online reinforcement learning-assisted method, adaptive collaborative control that ensures frequency stability, economical operation, and equipment health is achieved.

Benefits of technology

It enhances the adaptive and collaborative capabilities of the AGC system in complex environments, improves the operating efficiency, reliability, and equipment health of the power system, and realizes dynamic weight allocation and rapid disturbance identification and compensation for multiple objectives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121923162A_ABST
    Figure CN121923162A_ABST
Patent Text Reader

Abstract

The invention discloses an AGC adaptive cooperative control method based on multi-target driving, and belongs to the technical field of automatic control and scheduling of a power system. Aiming at the problems of insufficient multi-target coordination, poor system sensitivity, disturbance response lag and the like existing in AGC control in large-scale new energy access and complex interconnection system environments, a comprehensive technical scheme based on multi-target dynamic weight adjustment, complex network collaborative scheduling, multi-source information fusion identification and reinforcement learning online optimization is provided. According to the method, the system state can be sensed in real time, multi-target adaptive collaborative optimization of stable frequency, economic operation, equipment health and the like is realized, various disturbances are quickly identified and accurately compensated, and the multi-region control elasticity and intelligent level are improved. Experimental results show that the dynamic response capability, the robustness and the economical efficiency of the AGC system are remarkably enhanced, and the method is suitable for intelligent regulation and control and high-quality operation requirements of a modern electric power system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power system automatic control and dispatching technology, and more specifically relates to an AGC adaptive cooperative control method based on multi-objective driving. Background Technology

[0002] Automatic Generation Control (AGC), as a core technology for secondary frequency regulation and coordinated operation of modern power systems, aims to achieve system frequency stability, economical power dispatch, and safe and healthy equipment operation. With the large-scale integration of new energy sources, increasingly complex grid structures, frequent load fluctuations, and higher requirements for equipment health management, traditional AGC control methods, relying on fixed parameter settings and empirical models, are ill-suited to effectively cope with complex, dynamic, and uncertain operating environments. Existing technologies often focus on optimizing a single objective of frequency or power, neglecting the coupling and trade-offs between multiple objectives. Furthermore, they generally suffer from problems such as response lag, control rigidity, and local optima in areas such as disturbance identification, heterogeneous data fusion, and regional coordinated dispatch.

[0003] Furthermore, the dynamic changes in network topology, insufficient real-time sensing capabilities for multi-source heterogeneous data in the power system, and the empirical nature of controller parameter adjustment further limit the adaptive and collaborative optimization capabilities of AGC systems under large-scale and complex conditions. In recent years, although some studies have attempted to introduce methods such as artificial intelligence, game theory, and deep reinforcement learning to improve control intelligence and autonomy, unified and systematic solutions are still lacking in areas such as multi-objective dynamic weight allocation, collaborative scheduling under complex time-varying networks, accurate disturbance identification and compensation, and full lifecycle adaptive optimization. Therefore, there is an urgent need to develop a novel AGC control method capable of addressing multiple objectives and possessing strong adaptive and collaborative capabilities to comprehensively improve the operational resilience, economic efficiency, and equipment health of the power grid, meeting the requirements for high-quality operation of modern power systems. Summary of the Invention

[0004] Addressing the prominent issues of slow response, rigid control, insufficient information utilization, and local optima in existing AGC technologies, particularly in areas such as multi-objective coordination, system state sensitivity assessment, collaborative scheduling in complex topologies, rapid disturbance identification, and adaptive parameter optimization, this invention aims to solve the key technical problem of achieving dynamic weight adaptive allocation for multiple objectives, including frequency stability, economical operation, and equipment health. This will enhance the collaborative optimization capabilities of AGC controllers in various regions of complex interconnected systems. By employing multi-source information fusion and deep identification technologies, the invention achieves rapid disturbance identification and source localization. Furthermore, throughout the entire lifecycle, intelligent algorithms such as reinforcement learning are used to continuously correct and optimize control parameters, ultimately realizing efficient, robust, and intelligent adaptive collaborative control of the AGC system in dynamically changing environments.

[0005] To achieve the above objectives, the present invention employs the following technical solution: the method comprises: A multi-objective dynamic weight adaptive determination mechanism is constructed, which introduces information entropy and system sensitivity analysis to explore the sensitivity of the current system state to the objectives of frequency stability, economy and equipment health, and dynamically generates real-time weights for different control objectives based on fuzzy Bayes inference. For time-varying complex interconnected network collaborative scheduling, an adaptive collaborative scheduling algorithm for multi-region generators based on time-varying topological complex network is proposed. By utilizing evolutionary network reconstruction theory, the collaborative graph is actively reconstructed according to the interaction influence between each generator unit at regular scheduling cycles. By introducing a topology evolution-driven spatiotemporal coupling controller, the elastic cooperation and response capability of AGC controllers in each region are improved. Adaptive disturbance identification and compensation based on multi-source information fusion introduces multi-modal sensor fusion technology, comprehensively utilizes multi-source heterogeneous data such as frequency fluctuations, active power snapshots, equipment status, meteorological and load forecasts of the power system, and achieves rapid disturbance identification, source location and quantitative analysis of disturbance magnitude through deep tensor decomposition and symbol dynamic identification algorithms, and adaptively adjusts AGC control output. The AGC instruction optimization driven by a two-level adaptive game theory has two levels. The first level focuses on each generator unit and introduces a mutually beneficial evolutionary game learning mechanism to adaptively adjust the control strategy based on local objectives and global feedback. The second level is guided by the overall network optimization and realizes multi-objective global coordination of AGC instructions of each unit. Online reinforcement learning-assisted autonomous performance correction utilizes real-time feedback from global performance metrics to autonomously explore and adjust parameters at each stage of AGC. By integrating a reward and penalty function system, it balances rapid response with robustness, driving the system to continuously optimize its performance under changing environments and loads, ultimately achieving full lifecycle adaptive collaborative control of the AGC system.

[0006] In one scheme, the construction of the multi-objective dynamic weight adaptive determination mechanism includes: a quantification evaluation system for the sensitivity of each control objective to the current system state; selecting a set of related feature indicators for each control objective as input; normalizing each indicator; constructing the associated feature vector of each objective based on historical and real-time data; applying the information entropy method to calculate the information entropy of each feature indicator; and obtaining the overall sensitivity of each control objective. This paper introduces fuzzy Bayesian inference, using the sensitivity and current operating state of each objective as input membership degrees. Based on a fuzzy rule base and a Bayesian conditional probability model, it outputs the comprehensive priority of each objective. Finally, based on the comprehensive priority of each objective, dynamic objective weights are dynamically allocated to achieve real-time balance and on-demand adjustment among multiple objectives.

[0007] In one scheme, the time-varying complex interconnected network collaborative scheduling includes: abstracting each power generation unit in the region as a network node, the edges between nodes representing the physical interconnection relationship or the coupling strength of the interaction, using an adjacency matrix to represent the dynamic interaction influence degree between each node, and the quantification of the influence degree combined with the weight allocation result, obtained through a multi-index weighting function. An evolutionary network reconstruction mechanism is adopted to re-evaluate and adjust the existence and weight of edges based on the influence between nodes, thereby achieving adaptive evolution of the topology.

[0008] In one scheme, the adaptive disturbance identification and compensation of multi-source information fusion includes: the adaptive disturbance identification and compensation of multi-source information fusion includes: constructing a high-dimensional tensor representation for multi-modal time series data of frequency, active power, equipment health status, meteorological characteristics, and load forecast, so as to realize the extraction of deep coupling relationship of multi-source data and dimensionality reduction and noise reduction; Tensor decomposition is used to obtain potential disturbance features, and a symbol dynamic identification algorithm is introduced. By discretizing symbol encoding and Markov probability transition matrix, the dynamic changes of the system before and after the disturbance are characterized. Combined with KL divergence or mutual information algorithm, the disturbance is quantitatively detected. The identification results are fed back to the AGC collaborative control module after Bayesian uncertainty analysis. Based on the disturbance type and characteristics, the multi-objective weights and control parameters are adaptively corrected, and a response compensation bias is generated to adjust the AGC control output in real time, thereby achieving accurate identification and adaptive compensation of multi-source large disturbances and unknown disturbances.

[0009] In one scheme, the dual-layer adaptive game-driven AGC instruction optimization includes: adopting a dual-layer game structure, with each generator unit as an independent game subject at the bottom layer, and adaptively optimizing and adjusting the AGC parameters of the unit based on its own strategy space and payoff function, combined with multi-objective weights, through evolutionary game learning, and sensing the system objective weights and disturbance feedback in real time. The high-level management treats the entire system as a game, gathers feedback from each unit, and, based on the global payoff function and the weights of the main control objectives, achieves dynamic weight allocation and multi-objective joint optimal allocation of AGC commands according to multiple indicators such as frequency, economy, and safety. The bottom-level disturbance identification and collaborative network state information are pushed to the top level, and the optimization instructions from the top level are fed back to the bottom level. This constructs a dynamic closed-loop, multi-objective, multi-agent game-coordinated adaptive control system to achieve optimal robust operation at the system level.

[0010] In one scheme, the online reinforcement learning-assisted autonomous performance correction includes: introducing an augmented reinforcement learning algorithm based on global feedback, modeling the AGC control process as a Markov decision process, and incorporating frequency deviation, control energy consumption, unit health status, target overshoot, response rate, robustness, and disturbance tolerance into a multi-objective weighted system by formulating an innovative reward and punishment function system, thereby achieving a dynamic comprehensive evaluation of system performance.

[0011] In one scheme, the time-varying complex interconnected network collaborative scheduling introduces a topology evolution-driven controller parameter update mechanism, in which the control parameters change in conjunction with the network structure and target weights, thereby achieving automatic and elastic cooperation between regions.

[0012] In one scheme, the time-varying complex interconnected network collaborative scheduling defines the output of the regional collaborative controller and adopts a dynamic collaborative mechanism. It combines local feedback information optimized by multi-objective weights and integrates dynamic adjustment of network topology weights to achieve information sharing and sensitivity adaptive adjustment between regions.

[0013] In one approach, the online reinforcement learning-assisted autonomous performance correction employs deep reinforcement learning or explicit policy iterative algorithms, autonomously exploring and optimizing AGC parameters based on real-time updates of policy parameters. By combining experience playback mechanisms and multi-granularity reward adjustment, the AGC system can adapt to environmental changes in the long term and continuously achieve efficient, robust, and autonomous intelligent full life cycle collaborative optimization control.

[0014] Beneficial effects of this invention: This invention constructs a multi-objective dynamic weight adaptive determination mechanism, which can automatically balance multiple control objectives such as frequency stability, economy, and equipment health according to the system's operating status, achieving a sensitive response to actual operating conditions. It employs a collaborative scheduling method based on a time-varying complex interconnected network, effectively improving the flexible cooperation and rapid response capabilities of AGC controllers in different regions under interconnected conditions on the power generation side. The introduction of multi-source heterogeneous data fusion and deep identification methods makes disturbance identification, source location, and quantitative analysis more accurate and reliable, achieving adaptive disturbance compensation and significantly enhancing the system's security and robustness. Through game-driven and hierarchical optimization mechanisms, dynamic coordination between local objectives and global performance is ensured, avoiding the drawbacks of previous local optima. Combined with online reinforcement learning-assisted autonomous performance correction, the AGC system can continuously learn and optimize itself during operation, significantly improving its adaptability and intelligence under environmental changes and load disturbances.

[0015] In summary, this invention significantly improves the adaptive and cooperative control capability of the AGC system, enhances the overall operating efficiency, reliability, and equipment health assurance level of the power system, and has significant application and promotion value. Attached Figure Description

[0016] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0017] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Typical embodiments of the invention are shown in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.

[0018] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. To facilitate understanding, the invention will now be described more fully with reference to the accompanying drawings. Typical embodiments of the invention are shown in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to make the disclosure of the invention more thorough and complete.

[0019] like Figure 1 As shown, the AGC adaptive cooperative control method based on multi-objective driving specifically includes: Step 1: Construction of a multi-objective dynamic weight adaptive determination mechanism The multi-objective weight adaptive allocation algorithm introduces information entropy and system sensitivity analysis to autonomously explore the sensitivity of the current system state to objectives such as frequency stability, economy, and equipment health. Based on a fuzzy Bayesian inference system, it dynamically generates real-time weights for different control objectives, enabling flexible adjustment of weights according to the system's operating state and ensuring that each objective is balanced as needed.

[0020] To implement a multi-objective dynamic weight adaptive determination mechanism, it is first necessary to construct an evaluation system that can quantify the sensitivity of each control objective (such as frequency stability, economy, equipment health, etc.) to the current system state. In the specific implementation process, the first step is to address each control objective... Select a set of related feature indicators Inputs include frequency deviation, tie-line power, load forecast deviation, or equipment health status factors. For each index xi, normalization is applied to obtain a dimensionless index. Based on historical and real-time data, correlation feature vectors for each target are constructed. Then, the information entropy method is used to calculate the information entropy of each feature indicator. :

[0021] in, Where m is the number of samples. The lower the information entropy, the more important the indicator. This method yields the overall sensitivity of each control objective Oj. For example, relevant feature indicators can be combined using a weighted method:

[0022] in, For each feature in the target The basic weights are determined based on the objectives. To further enhance the flexibility and intelligence of weight allocation, a fuzzy Bayesian inference system is introduced. This inference system uses each objective as a basis for... Sensitivity Using the current operating state as input membership, and based on a pre-designed fuzzy rule base and a Bayesian conditional probability model, the system outputs the comprehensive priority of each objective. Specifically, fuzzy inference sets rules such as "if the frequency is stable and the sensitivity is high and the health status is average, then the frequency weight is greater than the health weight," and updates the weight uncertainty according to the probability distribution.

[0023] Ultimately, dynamic target weights The calculation is as follows:

[0024] This dynamic allocation process has good adaptive capabilities: under different operating conditions (such as sudden changes in system frequency, load disturbances, or decline in equipment health), the weights automatically increase due to increased sensitivity, realizing real-time balance and on-demand adjustment among multiple objectives, and providing quantitative and flexible weight inputs for subsequent AGC optimization control.

[0025] Step 2: Cooperative Scheduling of Time-Varying Complex Interconnected Networks A multi-regional generator adaptive cooperative scheduling algorithm based on time-varying complex topological networks is proposed. Utilizing evolutionary network reconstruction theory, the cooperative graph is proactively reconstructed at regular scheduling cycles based on the interaction influence between each generator unit. By introducing a topology-evolution-driven spatiotemporal coupled controller, the elastic cooperation and responsiveness of the AGC controllers in each region are improved, effectively addressing complex operating conditions and sudden events.

[0026] In the implementation of collaborative scheduling in time-varying complex interconnected networks, it is first necessary to abstract each power generation unit in the region as a network node, and the edges between nodes represent their physical interconnection relationships or the coupling strength of their interactions. An adjacency matrix is ​​then defined. ,in This represents the dynamic interaction influence between generator units i and j at time t. The quantification of this influence can be achieved by combining the results of the weight allocation in step 1, referencing indicators such as power exchange, frequency coupling sensitivity, and geographical distance, and using the following weighting function:

[0027] in, For real-time active power interaction between the two generating units, For frequency dynamic sensitivity, To normalize geographical distance, The function represents the weighting coefficients for each part. This is used to normalize the influence degree. To adapt to the dynamic changes in system structure and operating conditions, an evolutionary network reconstruction mechanism is adopted. Every scheduling cycle T, the existence and weight of edges are reassessed and adjusted based on the influence degree between nodes, achieving adaptive evolution of the topology. Specifically, for each pair of nodes, if... If the threshold is exceeded, the association is maintained or strengthened in the collaborative scheduling network; otherwise, the association is weakened or even disconnected, thus realizing the time-varying reorganization of the network.

[0028] In this adaptive network structure, the output of the regional cooperative controller is defined as follows: Its dynamic coordination mechanism can draw on the concept of spatiotemporal coupling consistency control, and the following coordination control law can be constructed:

[0029] in, The state of node i (such as terminal frequency deviation or output index). This is the feedback information for local AGC based on multi-objective weight optimization (associated with step 1). The coefficients are adaptively adjusted. This control law, through dynamic adjustment of network topology weights and organic embedding of multi-layer target feedback, enables regions to share global frequency and power regulation information while also possessing the ability to flexibly adjust according to operational sensitivity.

[0030] Furthermore, by introducing a topology evolution-driven controller parameter update mechanism, where control parameters change in tandem with network structure and target weights, spatiotemporal coupling and flexible cooperation of the network are achieved. Specifically, local control parameters can be defined as:

[0031] in For the dynamic weights in step 1, This is the parameter scheduling mapping function. Therefore, the system can achieve flexible linkage and optimal cooperation between AGC controllers in different areas under changes such as normal operation and disturbances, significantly improving the system's robustness and responsiveness to complex operating conditions and emergencies.

[0032] Step 3: Adaptive Perturbation Identification and Compensation Based on Multi-Source Information Fusion A disturbance prediction and adaptive compensation algorithm based on multi-source information fusion is designed. This algorithm introduces multi-modal sensor fusion technology, comprehensively utilizing heterogeneous data from multiple sources such as power system frequency fluctuations, active power snapshots, equipment status, meteorological data, and load forecasts. Through deep tensor decomposition and symbolic dynamic identification algorithms, it achieves rapid disturbance identification, source location, and quantitative analysis of disturbance magnitude. Furthermore, it adaptively adjusts the AGC control output to improve adaptability to large and unknown disturbances.

[0033] The core of the multi-source information fusion adaptive disturbance identification and compensation algorithm lies in the efficient integration of various heterogeneous data from the power system, enabling rapid disturbance perception, type identification, magnitude judgment, and targeted compensation. Specifically, firstly, regarding system frequency... A high-dimensional tensor representation is constructed from multimodal time-series data including active power, equipment health status, meteorological characteristics, and load forecasting. Where N is the number of data sources, M is the feature dimension, and T is the number of time steps. To extract deep coupling relationships from multi-source information and to reduce dimensionality and noise, tensor decomposition methods such as CP decomposition are used to represent the observation tensor as a product of a set of weight vectors. ,in For component intensity weights, These are the feature vectors for each mode. Based on the decomposed latent factors, the inherent perturbation characteristics and co-evolutionary relationships of each source data can be revealed.

[0034] Building upon tensor decomposition, a Symbolic Dynamic Filtering (SDF) algorithm is further introduced. Its basic process involves identifying key sense quantities such as... Discretize the symbols to obtain the symbol sequence. By constructing the Markov probability transition matrix This can characterize the dynamic pattern transformation of the system before and after a disturbance. Combining the dynamic weights from step 1 and the network interaction states from step 2, key disturbance signs such as feature anomalies, mutation boundaries, and co-evolutionary mutations are extracted in the symbolic domain and tensor latent space. Quantitative detection is then performed using algorithms such as KL divergence or mutual information. For example:

[0035] when Once the adaptive threshold is exceeded, a disturbance can be identified, and the source and magnitude of the disturbance can be further located and quantified based on the modal weights of the tensor decomposition.

[0036] The identification results, after undergoing Bayesian uncertainty analysis, are fed back to the AGC collaborative control module to adaptively adjust the multi-objective weights. and control parameters. After a disturbance occurs, a response compensation bias is generated according to the type (such as sudden frequency drop, sudden load increase, equipment failure, etc.). Its expression is:

[0037] in, By integrating the disturbance symbol sequence, local tensor disturbance components, and the current target weight, a corresponding compensation strategy is generated, and the AGC control output is adjusted in real time. In this way, the system can accurately identify and locate multi-source large disturbances and unknown disturbances, and achieve adaptive and targeted responses from the controller, significantly enhancing its robustness and flexibility under complex operating conditions.

[0038] Step 4: Optimization of AGC instructions driven by two-layer adaptive game theory The first stage focuses on each generator unit and introduces a mutually beneficial evolutionary game learning mechanism to adaptively adjust the control strategy based on local objectives and global feedback. The second stage is guided by the overall network optimization and achieves multi-objective global coordination of AGC commands of each entity, maximizing the overall system operating efficiency and response robustness while ensuring frequency, economy and safety.

[0039] The dual-layer adaptive game-driven AGC command optimization mechanism significantly improves the intelligence of frequency regulation commands and the system's response efficiency by simultaneously advancing the coordination of generator unit interests and the global optimization of the system. Specifically, in the lower layer (first level), each generator unit is treated as an independent game player, with each player possessing its own strategy space. and payoff function The profit function is based on the multi-objective weights dynamically generated in step 1. This reflects the varying contributions of factors such as frequency stability, economic efficiency, and health status to local operation. Each agent adopts evolutionary game learning based on its local state and feedback from higher levels; policy updates can employ modified replication dynamics.

[0040] in, For learning rate, This represents the individual return under the current portfolio strategy. The average payoff for the entire strategy group is denoted as . This mechanism uses a mutually beneficial (cooperative-win) game model to perceive information about neighboring units and the system in real time, and adjusts the unit AGC parameters under the drive of target weights and disturbance feedback (steps 1 and 3) to achieve a locally optimal response.

[0041] Building upon this, the higher-level (second-level) system treats the entire system as a single game entity, aggregating feedback from each unit through a global payoff function \(U_{sys}\) to achieve multi-objective optimal allocation of cross-regional AGC commands. The global payoff is defined as the weighted sum of all major control objectives:

[0042] in , , The comprehensive indicators for each unit in terms of frequency, economy, and safety. The weights are dynamically allocated by the system's current adaptive weight mechanism (step 1). The global AGC instruction optimization model is a multi-objective joint optimization game problem, which can be solved using the Lagrange multiplier method or a multi-objective evolutionary algorithm.

[0043] Under constraints such as system frequency limitations and safe operating range of unit output, a Qatar point or Nash equilibrium solution is formed. During the optimization process, the higher management will adjust the weight allocation and command dispatch based on the optimal game parameters fed back from the lower level of each unit, so as to achieve flexible coordination of multiple objectives and global optimization, ensuring that frequency regulation is fast, economical, and safe to operate.

[0044] In actual operation, the underlying adaptive game recursion will continuously report the disturbance identification (step 3) and cooperative network status (step 2) information to the upper layer. The AGC instructions of the upper layer global optimization will act as feedback in real time on each subject, continuously correcting their behavior, thereby forming a dynamic closed-loop, multi-objective, multi-subject game coordination and adaptive control system, and finally achieving optimal robust operation at the system level.

[0045] Step 5: Correction of Autonomous Performance with Online Reinforcement Learning Assistance By leveraging real-time feedback from global performance metrics, the system autonomously explores and adjusts parameters at each stage of AGC. Integrating a reward and penalty function system, it balances rapid response with robustness, driving the system to continuously optimize its performance under changing environments and loads, ultimately achieving adaptive and collaborative control throughout the entire lifecycle of the AGC system.

[0046] In the autonomous performance correction stage assisted by online reinforcement learning, an augmented reinforcement learning algorithm based on global feedback is introduced, enabling the AGC system to automatically optimize its parameters based on real-time performance during actual operation. First, the AGC control process is modeled as a Markov decision process (MDP), assuming the system state at time t is... The actions taken (such as adjusting the AGC adjustment parameter vector) are as follows: The dynamics of environmental transfer are The system receives rewards based on environmental feedback. Its value is dynamically calculated by an innovative reward and penalty function system. The reward function not only incorporates frequency deviation... Controlling energy consumption Unit health status In addition to target overshoot and response rate, robustness and disturbance tolerance are also incorporated into a multi-objective weighted system, specifically expressed as follows:

[0047] in For real-time frequency deviation, Energy consumption deviation, Deviation of unit health parameters Penalties for overshoot and response latency, As a positive excitation for recovery rate or robustness to strong earthquakes, Each performance objective is assigned a weight. This innovative reward structure dynamically balances response speed and system robustness in the feedback process and can adaptively adjust according to operating conditions.

[0048] The learning algorithms employ either deep reinforcement learning-based algorithms (such as DDPG, PPO, etc.) or explicit policy iterative algorithms. Taking the policy gradient class as an example, the parameterized policy of the agent... Based on the current state, output AGC parameter adjustment instructions to maximize the cumulative expected return.

[0049] Continuously adjust strategy parameters Perform gradient updates:

[0050] in As a return discount factor, The learning rate is used. The system continuously acquires the global state, including the network evolution characteristics of the aforementioned four steps, multi-source disturbance assessment, and instruction game optimization results. This global data is input into the Agent, which continuously and autonomously explores the best parameter adjustment scheme, and quickly learns and corrects erroneous operations when faced with changes in operating conditions or load.

[0051] Furthermore, the system ensures that the parameters of each AGC stage continuously converge towards the optimal state throughout its lifecycle through online reward updates. Combined with an experience playback mechanism, historically optimal state-action pairs can be reused in a timely manner to accelerate convergence in new environments. Multi-level feedback information, including local disturbances and global responses, is integrated, and multi-granularity reward adjustment is employed to prevent getting trapped in local optima. Ultimately, the reinforcement learning-driven AGC adaptive adjustment system enables the control system to maintain high efficiency, robustness, and autonomous intelligence in collaborative scheduling under long-term operation and environmental changes, truly achieving adaptive collaborative optimization control throughout the entire lifecycle.

[0052] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0053] It should be understood that the above detailed description of the technical solutions of the present invention with reference to preferred embodiments is illustrative and not restrictive. Those skilled in the art can modify the technical solutions described in the embodiments or make equivalent substitutions for some of the technical features based on reading this specification; however, these modifications or substitutions do not cause the essence of the corresponding technical solutions to depart from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An AGC adaptive cooperative control method based on multi-objective drive, characterized in that: The method includes: A multi-objective dynamic weight adaptive determination mechanism is constructed, which introduces information entropy and system sensitivity analysis to explore the sensitivity of the current system state to the objectives of frequency stability, economy and equipment health, and dynamically generates real-time weights for different control objectives based on fuzzy Bayes inference. For time-varying complex interconnected network collaborative scheduling, an adaptive collaborative scheduling algorithm for multi-region generators based on time-varying topological complex network is proposed. By utilizing evolutionary network reconstruction theory, the collaborative graph is actively reconstructed according to the interaction influence between each generator unit at regular scheduling cycles. By introducing a topology evolution-driven spatiotemporal coupling controller, the elastic cooperation and response capability of AGC controllers in each region are improved. Adaptive disturbance identification and compensation based on multi-source information fusion introduces multi-modal sensor fusion technology, comprehensively utilizes multi-source heterogeneous data such as frequency fluctuations, active power snapshots, equipment status, meteorological and load forecasts of the power system, and achieves rapid disturbance identification, source location and quantitative analysis of disturbance magnitude through deep tensor decomposition and symbol dynamic identification algorithms, and adaptively adjusts AGC control output. The AGC instruction optimization driven by a two-level adaptive game theory has two levels. The first level focuses on each generator unit and introduces a mutually beneficial evolutionary game learning mechanism to adaptively adjust the control strategy based on local objectives and global feedback. The second level is guided by the overall network optimization and realizes multi-objective global coordination of AGC instructions of each unit. Online reinforcement learning-assisted autonomous performance correction utilizes real-time feedback from global performance metrics to autonomously explore and adjust parameters at each stage of AGC; it integrates a reward and penalty function system to balance rapid response and robustness, driving the system to continuously optimize its performance under environmental and load changes, ultimately achieving full lifecycle adaptive collaborative control of the AGC system.

2. The AGC adaptive cooperative control method based on multi-objective driving according to claim 1, characterized in that: The construction of the multi-objective dynamic weight adaptive determination mechanism includes: a quantitative evaluation system for the sensitivity of each control objective to the current system state; selecting a set of related feature indicators for each control objective as input; normalizing each indicator; constructing the associated feature vector of each objective based on historical and real-time data; applying the information entropy method to calculate the information entropy of each feature indicator; and obtaining the overall sensitivity of each control objective. Fuzzy Bayesian inference is introduced, using the sensitivity and current operating state of each objective as input membership. Based on the fuzzy rule base and Bayesian conditional probability model, the comprehensive priority of each objective is output. Finally, dynamic objective weights are dynamically allocated according to the comprehensive priority of each objective, so as to achieve real-time balance and on-demand adjustment among multiple objectives.

3. The AGC adaptive cooperative control method based on multi-objective driving according to claim 1, characterized in that: The aforementioned time-varying complex interconnected network collaborative scheduling includes: abstracting each power generation unit in the region as a network node, with the edges between nodes representing the physical interconnection relationship or the coupling strength of the interaction, using an adjacency matrix to represent the dynamic interaction influence degree between each node, and the quantification of the influence degree combined with the weight allocation result, obtained through a multi-index weighting function; An evolutionary network reconstruction mechanism is adopted to re-evaluate and adjust the existence and weight of edges based on the influence between nodes, thereby achieving adaptive evolution of the topology.

4. The AGC adaptive cooperative control method based on multi-objective driving according to claim 1, characterized in that: The adaptive disturbance identification and compensation of multi-source information fusion includes: constructing a high-dimensional tensor representation for multi-modal time series data of frequency, active power, equipment health status, meteorological characteristics, and load forecast, to realize the extraction of deep coupling relationships of multi-source data and dimensionality reduction and noise reduction; Tensor decomposition is used to obtain potential disturbance features, and a symbol dynamic identification algorithm is introduced. By discretizing symbol encoding and Markov probability transition matrix, the dynamic changes of the system before and after the disturbance are characterized. Combined with KL divergence or mutual information algorithm, the disturbance is quantitatively detected. The identification results are fed back to the AGC collaborative control module after Bayesian uncertainty analysis. Based on the disturbance type and characteristics, the multi-objective weights and control parameters are adaptively corrected, and a response compensation bias is generated to adjust the AGC control output in real time, thereby achieving accurate identification and adaptive compensation of multi-source large disturbances and unknown disturbances.

5. The AGC adaptive cooperative control method based on multi-objective drive according to claim 1, characterized in that: The aforementioned dual-layer adaptive game-driven AGC instruction optimization includes: adopting a dual-layer game structure, with each generator unit as an independent game subject at the bottom layer, based on its own strategy space and payoff function, combined with multi-objective weights, adaptively optimizing and adjusting the generator unit's AGC parameters through evolutionary game learning, and sensing the system's objective weights and disturbance feedback in real time; The top management treats the entire system as a game, gathers feedback from each unit, and, based on the global payoff function and the weights of the main control objectives, achieves dynamic weight allocation and multi-objective joint optimal allocation of AGC commands according to multiple indicators such as frequency, economy, and safety. The bottom-level disturbance identification and collaborative network state information are pushed to the top level, and the optimization instructions from the top level are fed back to the bottom level. This constructs a dynamic closed-loop, multi-objective, multi-agent game-coordinated adaptive control system to achieve optimal robust operation at the system level.

6. The AGC adaptive cooperative control method based on multi-objective driving according to claim 1, characterized in that: The aforementioned online reinforcement learning-assisted autonomous performance correction includes: introducing an enhanced reinforcement learning algorithm based on global feedback, modeling the AGC control process as a Markov decision process, and incorporating frequency deviation, control energy consumption, unit health status, target overshoot, response rate, robustness, and disturbance tolerance into a multi-objective weighted system by formulating an innovative reward and punishment function system, thereby achieving dynamic comprehensive evaluation of system performance.

7. The AGC adaptive cooperative control method based on multi-objective driving according to claim 3, characterized in that: The aforementioned time-varying complex interconnected network collaborative scheduling introduces a topology evolution-driven controller parameter update mechanism, in which the control parameters change in tandem with the network structure and target weights, thereby achieving automatic and elastic collaboration between regions.

8. The AGC adaptive cooperative control method based on multi-objective driving according to claim 3, characterized in that: The aforementioned time-varying complex interconnected network collaborative scheduling defines the output of the regional collaborative controller and adopts a dynamic collaborative mechanism. It combines local feedback information optimized by multi-objective weights and integrates dynamic adjustment of network topology weights to achieve information sharing and sensitivity adaptive adjustment between regions.

9. The AGC adaptive cooperative control method based on multi-objective driving according to claim 6, characterized in that: The aforementioned online reinforcement learning-assisted autonomous performance correction uses deep reinforcement learning or explicit policy iterative algorithms to autonomously explore and optimize AGC parameters based on real-time updates of policy parameters. By combining experience playback mechanisms and multi-granularity reward adjustment, the AGC system can adapt to environmental changes in the long term and continuously achieve efficient, robust, and autonomous intelligent full life cycle collaborative optimization control.