Coordinated control method for AGC and primary frequency modulation of thermal power unit

CN122532947APending Publication Date: 2026-08-07HUANENG JILIN POWER GENERATION CO LTD CHANGCHUN THERMAL POWER PLANT
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUANENG JILIN POWER GENERATION CO LTD CHANGCHUN THERMAL POWER PLANT
Filing Date
2026-04-29
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0002]随着新型电力系统建设的深入推进,电网的电源结构发生了深刻变革;在包含高比例波动性新能源、电化学储能、具有调节能力的水电以及传统火电机组的省级区域电网中,系统频率的稳定运行依赖于多种异质调频资源的协同支撑;传统上,火力发电机组凭借其可靠的出力和调节能力,一直是自动发电控制的主力调频资源;然而,在新型电力系统场景下,一方面,风电、光伏等新能源的大规模并网引入了显著且快速的功率波动,对调频资源的响应速度和调节精度提出了更高要求;另一方面,电化学储能等新型资源具有毫秒级响应和近乎线性的功率调节特性,其调频性能与经济性在某些场景下优于传统火电机组,水电则具备调节容量大但受限于水库调度计划与振动区约束;在此多能互补的复杂格局中,如何高效、经济地协调这些特性迥异的调频资源,特别是重新定位火电机组的角色并优化其与储能、水电的协同方式,以实现全网调频综合成本最优,已成为系统运行面临的核心挑战之一

Benefits of technology

[0054]本发明的有益效果是:通过融合精细化动态成本建模、激励相容的集中优化与具备协同感知能力的分布式决策,实现了多能互补系统调频经济性、安全性及设备友好性的统一提升;其核心价值在于将调频服务转化为可动态定价的商品,利用统一出清价格引导异质资源在追求个体收益的同时自发达成全局最优,并通过在线学习机制使得模型能够持续跟踪系统变化,自适应地优化成本预测与协同策略,从而在应对高比例新能源随机扰动时,依然可保障频率稳定并降低整体运行成本,形成具备持续进化能力的智能调频生态。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122532947A_ABST
    Figure CN122532947A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of based on thermal power unit AGC and primary frequency coordination control method, specifically relates to the field of thermal power unit coordinated control, by fusing fine dynamic cost modeling, incentive compatible centralized optimization and with collaborative perception ability distributed decision, realized the unity of multiple energy complementary system frequency economic efficiency, security and equipment friendliness It is promoted;Its core value lies in that frequency modulation service is converted into dynamically priced goods, using unified clearing price to guide heterogeneous resources to achieve global optimum while pursuing individual income, and through online learning mechanism, the model can continuously track system changes, adaptively optimize cost prediction and collaborative strategy, so that when responding to high proportion new energy random disturbance, frequency stability can still be guaranteed and overall operating cost is reduced, forming an intelligent frequency modulation ecology with continuous evolution ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of coordinated control of thermal power units, and more specifically, to a method for coordinated control of thermal power unit AGC and primary frequency regulation. Background Technology

[0002] With the deepening of the construction of new power systems, the power supply structure of the power grid has undergone profound changes. In provincial regional power grids that include a high proportion of volatile new energy sources, electrochemical energy storage, hydropower with regulation capabilities, and traditional thermal power units, the stable operation of the system frequency depends on the coordinated support of various heterogeneous frequency regulation resources. Traditionally, thermal power generating units, with their reliable output and regulation capabilities, have always been the main frequency regulation resource for automatic generation control. However, in the context of new power systems, on the one hand, the large-scale grid connection of new energy sources such as wind power and photovoltaics has introduced significant and rapid power fluctuations, impacting the frequency regulation resources. Higher demands are placed on speed and regulation accuracy. On the other hand, new resources such as electrochemical energy storage have millisecond-level response and near-linear power regulation characteristics. Their frequency regulation performance and economy are superior to traditional thermal power units in some scenarios. Hydropower has a large regulation capacity but is limited by reservoir scheduling plans and vibration zone constraints. In this complex pattern of multi-energy complementarity, how to efficiently and economically coordinate these frequency regulation resources with different characteristics, especially to redefine the role of thermal power units and optimize their coordination with energy storage and hydropower, so as to achieve the optimal overall frequency regulation cost of the entire network, has become one of the core challenges facing the system operation.

[0003] Currently, solutions to the aforementioned multi-resource frequency regulation coordination problem mainly rely on fixed priority ranking or simple proportional allocation strategies pre-set by the dispatch center. These methods fail to fully consider the differences in dynamic adjustment costs of different frequency regulation resources on a time scale of seconds to minutes, as well as the real-time state constraints of the equipment itself. For example, the frequency regulation cost of electrochemical energy storage is strongly correlated with its charge and discharge depth, current state of charge, and cycle life loss; the frequency regulation cost of thermal power units is implicit in their fuel consumption, equipment wear, and life loss due to frequent operations, and their regulation capability and response speed are significantly affected by the current operating conditions; the response of hydropower units needs to avoid vibration zones and consider hydrological conditions. Existing technologies lack a mechanism that can integrate the dynamic marginal costs of various resources in real time. The current joint optimization framework based on equipment health status often fails to achieve global economic optimization in command allocation. As a result, energy storage resources with high regulation costs may be used to handle low-cost regulation tasks that could be handled by hydropower, while thermal power units, which should focus on providing support for large-capacity, slow regulation, are forced to frequently perform inefficient high-frequency, small-amplitude regulation. This not only wastes and reduces the overall efficiency of frequency regulation resources but also accelerates the fatigue wear of key equipment in thermal power units, bringing enormous long-term economic operation and maintenance pressure. Therefore, there is an urgent need for a collaborative control method that can comprehensively consider the dynamic economic characteristics of heterogeneous resources and the physical constraints of equipment on a rapid time scale to achieve refined and intelligent allocation of frequency regulation commands driven by economic efficiency. Summary of the Invention

[0004] This invention addresses the technical problems existing in the prior art by providing a method for coordinated control of AGC and primary frequency regulation in thermal power units, thereby solving the problems mentioned in the background art.

[0005] The technical solution of this invention to solve the above-mentioned technical problems is as follows: specifically, it includes the following steps:

[0006] Step S1: The dispatch center collects real-time status data of various frequency regulation resources in the power grid, including thermal power units, energy storage devices and hydropower units, and calls the dynamic marginal cost calculation model to calculate the dynamic marginal cost curve of various frequency regulation resources based on the real-time status data. The dynamic marginal cost curve is used to characterize the economic cost of providing a unit of regulation power in the future period.

[0007] Step S2: Based on the dynamic marginal cost curves of various frequency regulation resources obtained in Step S1, the dispatch center constructs a real-time optimization model with the goal of minimizing the total frequency regulation cost of the entire network and solves it to obtain the power command pre-allocation scheme and unified clearing price for each frequency regulation resource within a future rolling time window.

[0008] Step S3: The local control unit configured for each frequency regulation resource receives the power command pre-allocation scheme and unified clearing price, and based on the pre-built multi-agent reinforcement learning decision model, with the goal of maximizing their respective long-term net benefits, performs distributed collaborative decision-making and dynamic fine-tuning on the received power command to generate the final execution command.

[0009] Step S4: The scheduling center collects the actual running data after the final execution instruction is executed, and uses the actual running data to update and optimize the dynamic marginal cost calculation model and the multi-agent reinforcement learning decision model online.

[0010] In a preferred embodiment, in step S1, the process of the dispatch center collecting real-time status data of various frequency regulation resources in the power grid and calculating the dynamic marginal cost curves of various frequency regulation resources using a dynamic marginal cost calculation model is as follows:

[0011] The real-time status data collected includes the real-time output power, regulation rate, main steam pressure, main steam temperature, reheat steam temperature and first-stage metal temperature of the high-pressure cylinder of the thermal power unit; the current state of charge, maximum charge and discharge power and current cycle life count of the energy storage device; and the current output, vibration zone boundary data and available regulating capacity of the reservoir of the hydropower unit.

[0012] The process of calling the dynamic marginal cost calculation model and calculating the dynamic marginal cost curves of various frequency regulation resources based on real-time status data is as follows: For thermal power units, their dynamic marginal cost consists of fuel cost component and life loss cost component; the fuel cost component is based on the real-time coal consumption characteristic curve of the thermal power unit, and the marginal fuel cost is calculated by differentiating the real-time output power of the current thermal power unit.

[0013] The calculation of the life loss cost component is achieved through a real-time life loss assessment model. This model takes the operating status sequence data of the thermal power unit within the recent time window as input. The operating status sequence data includes key metal temperature, main steam pressure, main steam temperature, reheat steam temperature, real-time output power, and regulation rate. The real-time life loss assessment model outputs a life loss coefficient. The life loss coefficient is multiplied by the unit capacity investment cost conversion factor to obtain the first product. The first product is then multiplied by the regulation severity additional factor to obtain the life loss cost component. The regulation severity additional factor is the sum of the product of a number 1, the regulation severity weighting coefficient, and the absolute value of the current reference power change rate. The current reference power change rate represents the severity of regulation.

[0014] In a preferred embodiment, the dynamic marginal cost of the thermal power unit is the sum of the marginal fuel cost, the lifetime depreciation cost component, and the operation and maintenance marginal cost constant.

[0015] For energy storage devices, their dynamic marginal cost mainly consists of the cycle life cost component and the state of charge deviation penalty cost component. The cycle life cost component is calculated by dividing the single full cycle cost by the difference between the maximum tolerable cycle life and the current accumulated equivalent cycle number, and then multiplying it by the ratio of the current absolute value of the charge / discharge power to the current charge / discharge efficiency function.

[0016] The penalty cost component for state of charge deviation is calculated by multiplying the penalty coefficient by the square of the difference between the current state of charge and the preset optimal state of charge.

[0017] For hydropower units, the calculation of their dynamic marginal cost must avoid the vibration operating zone. When the current output of the hydropower unit is within the range formed by the lower limit and upper limit of the vibration zone power, the dynamic marginal cost is set to a very large positive number. When the current output of the hydropower unit is outside the range formed by the lower limit and upper limit of the vibration zone power, the dynamic marginal cost is obtained by multiplying the water price coefficient, the current net head, and the water consumption rate per unit power.

[0018] In a preferred embodiment, the specific process of the dispatch center constructing a real-time optimization model with the objective of minimizing the total frequency regulation cost of the entire network in step S2 is as follows:

[0019] The scheduling center uses a future rolling time window as the optimization period. The rolling time window is discretized into multiple consecutive optimization periods. An objective function is constructed, which minimizes the sum of the costs of all time periods and all frequency regulation resource calls within the future rolling time window.

[0020] The cost of calling a single frequency regulation resource in a single time period is obtained by definite integration of the dynamic marginal cost curve of the frequency regulation resource in the time period obtained in step S1 from the power zero point to the planned regulation power value of the resource in the time period. The planned regulation power is the decision variable to be optimized.

[0021] When calculating the sum of call costs, the scheduling center calculates a regulation contribution quality factor for each frequency regulation resource in each time period, multiplies the call cost by the corresponding regulation contribution quality factor, and then sums them up.

[0022] The adjustment contribution quality factor is calculated based on the recent historical comprehensive frequency modulation performance index of the frequency modulation resource through an exponential function with the natural constant as the base. The exponent of this exponential function is the product of the negative sensitivity coefficient and the number minus the difference of the historical comprehensive frequency modulation performance index. The higher the historical comprehensive frequency modulation performance index, the less the cost of calling the resource is discounted in the optimization.

[0023] In a preferred embodiment, the constraints of the real-time optimization model include power balance constraints, operational constraints of each frequency regulation resource, and state coupling constraints of the energy storage device.

[0024] The power balance constraint requires that the total planned regulation power equals the sum of the total demand under automatic generation control and the predicted primary frequency regulation demand.

[0025] Operational constraints for each frequency regulation resource include upper and lower power limits based on real-time status data and ramping constraints based on regulation rate;

[0026] For energy storage devices, additional coupling constraints and state of charge limit constraints are added to characterize the continuous change of the state of charge.

[0027] The dispatch center solves the real-time optimization model that satisfies all constraints. The optimal solution obtained includes the power command pre-allocation scheme for each time period and each frequency regulation resource within the future rolling time window, as well as the unified clearing price corresponding to each time period. The unified clearing price is the Lagrange multiplier corresponding to the power balance constraint at the optimal solution of the optimization problem. Its numerical meaning is the minimum marginal cost that the system needs to pay for each additional unit of power regulation demand under optimal scheduling.

[0028] In a preferred embodiment, in step S3, the local control unit configured for each frequency modulation resource makes decisions as an independent intelligent agent, and the state observed by each intelligent agent at the decision time includes the local state and the neighborhood state.

[0029] The local state includes: the power instruction pre-allocation value for the current time period corresponding to this resource in the power instruction pre-allocation scheme from step S2, and the corresponding unified clearing price for the current time period;

[0030] The real-time status data from step S1 includes, for thermal power units, the first-stage metal temperature of the high-pressure cylinder, and for energy storage devices, the current state of charge; as well as the locally measured or received grid frequency deviation.

[0031] The neighborhood state is the state summary information of adjacent frequency modulation resource agents obtained through communication, including the changing trend of the lifetime loss coefficient of adjacent thermal power units and the current state of charge of adjacent energy storage devices; the local state and the neighborhood state together constitute the complete observation state of the agent;

[0032] Each agent, based on the complete observation state, uses its internal policy network to output a fine-tuning action for the pre-allocated power command value for the current time period. This fine-tuning action is a power offset. This power offset is added to the pre-allocated power command value for the current time period to obtain the final execution command for the frequency modulation resource controlled by the agent.

[0033] In a preferred embodiment, the policy network inside the intelligent body dynamically fuses local and neighboring information using a graph attention mechanism before generating fine-tuning actions; the specific process is as follows:

[0034] The agent first calculates the correlation weight between its own state and the state of each neighboring agent, which is called the graph attention weight. This graph attention weight represents the importance of the state information of each neighboring agent to the current decision.

[0035] The calculation process of graph attention weights is as follows: First, using a shared linear transformation weight matrix, linear transformations are performed on the agent's own local state vector and the state vectors of each neighboring agent. Then, a trainable attention vector is concatenated with the transformed state of the agent and the transformed states of each neighboring agent, and an inner product operation is performed on the inner product result. A leaky linear rectified unit activation function is applied to the activation result, and then an exponential operation with the natural constant base is performed on the activation result to obtain the unnormalized weights of each neighbor and the agent itself. Finally, all unnormalized weights are summed, and each unnormalized weight is divided by the sum to obtain the normalized graph attention weights.

[0036] Then, the agent performs a linear transformation on the state information of all neighboring agents and itself using a shared linear transformation weight matrix, and then performs a weighted summation according to the corresponding graph attention weights to obtain the aggregated collaborative perception information.

[0037] Ultimately, the policy network, based on this aggregated collaborative perception information and the agent's own local state, jointly decides to generate fine-tuning actions; enabling each agent to selectively and adaptively consider the operating state of other frequency modulation resources associated with it when making local decisions, thereby achieving distributed collaboration.

[0038] In a preferred embodiment, the reward function used to drive the agent policy network learning and updating consists of four parts:

[0039] The first part is the income incentive item, which is the product of the current clearing price and the value of the fine-tuning action performed by the agent;

[0040] The second part is the incremental cost term, which is the difference between the dynamic marginal cost generated by the agent executing the final execution instruction and the dynamic marginal cost generated by only executing the power instruction pre-allocated value. The specific calculation method of the incremental cost term is as follows: take the final execution instruction value as input, call the dynamic marginal cost calculation model to obtain a cost value; then take the power instruction pre-allocated value as input, call the same dynamic marginal cost calculation model to obtain another cost value; calculate the difference between the two cost values.

[0041] The third part is the quality adjustment penalty term, which is the product of a positive weighting coefficient and the square of the local frequency deviation;

[0042] The fourth part is the policy regularization term, which is used to constrain the deviation between the agent's current decision policy and a reference policy;

[0043] After each decision and execution, each agent shares its decision experience data to a central experience pool, which is then trained by a central trainer using a multi-agent reinforcement learning algorithm. The multi-agent reinforcement learning algorithm is a multi-agent proximal policy optimization algorithm that periodically and centrally updates the policy network parameters of all agents based on the reward function.

[0044] In a preferred embodiment, step S4 involves the scheduling center collecting actual operational data after the final execution instruction is executed, specifically including:

[0045] The dispatch center aligns all collected actual operational data with a unified timescale at the second level, and then processes it through sliding window filtering and outlier detection to form a time series sample set.

[0046] The dispatch center calls a newly constructed data value assessment network. The data value assessment network calculates a scalar value score for each sample in the time series sample set. The value score is the product of three factors: the first factor is the stable value factor, which is calculated by using the natural constant as the base and the value of the absolute value of the actual frequency deviation of the power grid corresponding to the negative sample divided by a frequency deviation normalization scaling factor as the index.

[0047] The second factor is the prediction error value factor, which is calculated as follows: add the absolute value of the difference between the actual operating cost of the sample and the predicted cost obtained after inputting the real-time status data of the same sample into the dynamic marginal cost calculation model in step S1, and then divide by the sum of the actual operating cost of the sample and a very small positive number.

[0048] The third factor is the exploratory value factor, which is calculated by adding the information entropy of the output fine-tuning action distribution after inputting the states of the same sample into the multi-agent reinforcement learning decision model in step S3. This information entropy represents the exploratory nature of the strategy. The scheduling center selects high-value data samples ranked in the top 20% to 30% of the value scores to form a high-value sample subset.

[0049] In a preferred embodiment, the process of online updating and feedback optimization of the dynamic marginal cost calculation model and the multi-agent reinforcement learning decision model using actual operating data is as follows:

[0050] First, the dynamic marginal cost calculation model is updated by adopting the elastic weight joint algorithm. The loss function of the elastic weight joint algorithm includes two terms: the first term is the mean squared error term, and the second term is the elastic constraint term.

[0051] Secondly, the multi-agent reinforcement learning decision-making model is updated. The scheduling center divides the samples into several batches from easy to difficult based on the value scores of the samples in the high-value sample subset. The course learning strategy is adopted. The model is first trained using simple batches of samples with higher value scores or smaller frequency deviations. After the model performance is stable, difficult batches of samples with lower value scores or larger frequency deviations are gradually introduced for training, so that the policy network learns complex collaborative behaviors progressively.

[0052] Meanwhile, based on the recent changes in the mean square value of the power grid frequency deviation relative to the target threshold, the weight coefficient of the adjustment quality penalty term in the reward function is dynamically adjusted to achieve the adaptive learning objective.

[0053] After each model update, the scheduling center will evaluate the performance of the updated dynamic marginal cost calculation model and the multi-agent reinforcement learning decision model in an offline simulation environment. If the evaluation results do not meet the preset performance improvement requirements, a safety rollback mechanism will be triggered to abandon the current model update and automatically adjust the hyperparameters in subsequent model updates.

[0054] The beneficial effects of this invention are as follows: by integrating refined dynamic cost modeling, incentive-compatible centralized optimization, and distributed decision-making with collaborative perception capabilities, it achieves a unified improvement in the economy, safety, and equipment friendliness of frequency regulation in multi-energy complementary systems. Its core value lies in transforming frequency regulation services into dynamically priced commodities, using a unified clearing price to guide heterogeneous resources to spontaneously achieve global optimization while pursuing individual benefits, and enabling the model to continuously track system changes and adaptively optimize cost prediction and collaborative strategies through an online learning mechanism. Thus, even when dealing with random disturbances from high proportions of new energy sources, it can still ensure frequency stability and reduce overall operating costs, forming an intelligent frequency regulation ecosystem with continuous evolution capabilities. Attached Figure Description

[0055] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0056] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0057] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0058] In the description of this application, the term "for example" is used to indicate that it is used as an example, illustration, or illustration. Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0059] Example 1

[0060] This embodiment provides, for example Figure 1 The method shown here is a coordinated control method based on AGC and primary frequency regulation of thermal power units, which specifically includes the following steps:

[0061] Step S1: The dispatch center collects real-time status data of various frequency regulation resources in the power grid. Frequency regulation resources include at least thermal power units, energy storage devices and hydropower units. It also calls the dynamic marginal cost calculation model to calculate the dynamic marginal cost curve of various frequency regulation resources based on the real-time status data. The dynamic marginal cost curve is used to characterize the economic cost of providing a unit of regulation power in the future period.

[0062] Step S2: Based on the dynamic marginal cost curves of various frequency regulation resources obtained in Step S1, the dispatch center constructs a real-time optimization model with the goal of minimizing the total frequency regulation cost of the entire network and solves it to obtain the power command pre-allocation scheme and unified clearing price for each frequency regulation resource within a future rolling time window.

[0063] Step S3: The local control unit configured for each frequency regulation resource receives the power command pre-allocation scheme and unified clearing price, and based on the pre-built multi-agent reinforcement learning decision model, with the goal of maximizing their respective long-term net benefits, performs distributed collaborative decision-making and dynamic fine-tuning on the received power command to generate the final execution command.

[0064] Step S4: The scheduling center collects the actual running data after the final execution instruction is executed, and uses the actual running data to update and optimize the dynamic marginal cost calculation model and the multi-agent reinforcement learning decision model online.

[0065] In this embodiment, it is specifically necessary to explain the process in step S1 where the dispatch center collects real-time status data of various frequency regulation resources in the power grid and calls the dynamic marginal cost calculation model to calculate the dynamic marginal cost curves of various frequency regulation resources:

[0066] The collected real-time status data includes the real-time output power, regulation rate, main steam pressure, main steam temperature, reheat steam temperature, and first-stage metal temperature of the high-pressure cylinder of the thermal power unit; the current state of charge, maximum charge / discharge power, and current cycle life count of the energy storage device; and the current output, vibration zone boundary data, and available regulating capacity of the reservoir of the hydropower unit. The real-time output power of the thermal power unit and the current output of the hydropower unit are values ​​between the minimum and maximum technical output of the corresponding unit. The current state of charge of the energy storage device is a value between 0 and 1. The vibration zone boundary data of the hydropower unit is defined by the lower limit and upper limit of the vibration zone power.

[0067] The process of using the dynamic marginal cost calculation model to calculate the dynamic marginal cost curves of various frequency regulation resources based on real-time status data is as follows: For thermal power units, their dynamic marginal cost consists of a fuel cost component and a life loss cost component; the fuel cost component is calculated based on the real-time coal consumption characteristic curve of the thermal power unit by differentiating the real-time output power of the current thermal power unit to obtain the marginal fuel cost; the coal consumption characteristic curve is obtained by fitting unit performance tests or historical operating data, and characterizes the functional relationship between output power and standard coal consumption;

[0068] The calculation of the lifespan loss cost component is achieved through a real-time lifespan loss assessment model. This model takes the operating status sequence data of the thermal power unit within the recent time window as input. This data includes key metal temperatures, main steam pressure, main steam temperature, reheat steam temperature, real-time output power, and regulation rate. The model outputs a lifespan loss coefficient, a non-negative real number typically much less than 1. The real-time lifespan loss assessment model is a trained recurrent neural network, such as a long short-term memory network. Its training data comes from the unit's historical operating data and the lifespan loss assessment results from equipment maintenance records after the corresponding time period. The lifespan loss coefficient is multiplied by the unit capacity investment cost conversion factor to obtain the first product. The capacity investment cost conversion factor is a positive real number used for monetization conversion. Its value can be obtained by dividing the total investment in unit construction by the design life and rated capacity and then converting it to the frequency regulation service period. Multiplying the first product by the regulation severity additional factor yields the life loss cost component. The regulation severity additional factor is the sum of the product of the number one, the regulation severity weighting coefficient, and the absolute value of the current reference power change rate. The regulation severity weighting coefficient is a positive real number determined according to the characteristics of the unit, and its value can range from 0.1 to 0.5. It is used to amplify the additional losses caused by rapid regulation. The longer the unit has been in operation or the worse the equipment condition, the larger the value of this coefficient should be. The current reference power change rate represents the degree of regulation severity and is calculated by dividing the difference between the power command at the current moment and the power command at the previous moment by the sampling period.

[0069] The dynamic marginal cost of a thermal power unit is the sum of the marginal fuel cost, the lifetime depreciation cost component, and the operation and maintenance marginal cost constant; the operation and maintenance marginal cost constant covers additional labor and material consumption costs related to power regulation.

[0070] For energy storage devices, the dynamic marginal cost mainly consists of the cycle life cost component and the state-of-charge deviation penalty cost component. The cycle life cost component is calculated by dividing the single full cycle cost by the difference between the maximum tolerable cycle life and the current accumulated equivalent cycle life, and then multiplying it by the ratio of the current absolute value of charge / discharge power to the current charge / discharge efficiency function. The single full cycle cost is a positive real number, calculated based on the initial investment cost and guaranteed cycle life of the energy storage battery. The maximum tolerable cycle life is the rated value, and the current accumulated equivalent cycle life is the historical cumulative value. It is obtained by converting the depth of charge and discharge of each cycle into the equivalent full cycle life.

[0071] The penalty cost component for state of charge deviation is calculated by multiplying the penalty coefficient by the square of the difference between the current state of charge and the preset optimal state of charge. The preset optimal state of charge is a constant between zero and one; it is usually set to 0.5 to maintain the regulation margin of the energy storage device.

[0072] For hydropower units, the calculation of their dynamic marginal cost must avoid the vibration operating zone. When the current output of the hydropower unit is within the interval formed by the lower limit and upper limit of the vibration zone power, the dynamic marginal cost is set to a very large positive number, such as 10^6 yuan / MWh, to ensure that this interval is automatically avoided during optimization. When the current output of the hydropower unit is outside the interval formed by the lower limit and upper limit of the vibration zone power, the dynamic marginal cost is obtained by multiplying the water price coefficient, the current net head, and the unit power water consumption rate function. The unit power water consumption rate function is determined according to the turbine characteristic curve and characterizes the slight increase in water consumption rate under different outputs.

[0073] In this embodiment, it is particularly important to explain the specific process in step S2 where the scheduling center constructs a real-time optimization model with the goal of minimizing the total frequency regulation cost of the entire network:

[0074] The dispatch center uses a future rolling time window as the optimization period. The rolling time window is discretized into multiple consecutive optimization periods. An objective function is constructed, which minimizes the sum of the costs of all time periods and all frequency regulation resource calls within the future rolling time window. The length of the rolling time window is 5 to 15 minutes, and it is discretized into multiple optimization periods. The duration of a single optimization period is 15 to 60 seconds.

[0075] The cost of calling a single frequency regulation resource in a single time period is obtained by definite integration of the dynamic marginal cost curve of the frequency regulation resource in the time period obtained in step S1 from the power zero point to the planned regulation power value of the resource in the time period. The planned regulation power is the decision variable to be optimized.

[0076] When calculating the sum of call costs, the scheduling center calculates a regulation contribution quality factor for each frequency regulation resource in each time period, multiplies the call cost by the corresponding regulation contribution quality factor, and then sums them up.

[0077] The adjustment contribution quality factor is calculated based on the recent historical comprehensive frequency modulation performance index of the frequency modulation resource using an exponential function with the natural constant e as the base. The exponent of this exponential function is the product of a negative sensitivity coefficient and a number minus the difference in the historical comprehensive frequency modulation performance index. The historical comprehensive frequency modulation performance index is calculated by the dispatch center using a weighted average of the frequency modulation resource's response speed to frequency deviation commands, the root mean square value of the difference between actual output and target output, and the proportion of online available time over the past few minutes to one hour. The historical comprehensive frequency modulation performance index is evaluated based on the recent adjustment rate, adjustment accuracy, and availability of the resource, and its value ranges from zero to one. The sensitivity coefficient is a constant greater than zero, with a value range between 1 and 5. It is used to adjust the degree of influence of performance differences on cost discounts. The larger the value, the more significant the impact of performance differences on the adjustment contribution quality factor. The higher the historical comprehensive frequency modulation performance index, the closer the calculated adjustment contribution quality factor is to one, resulting in less discounting of the resource's call cost during optimization.

[0078] The constraints of the real-time optimization model include power balance constraints, operational constraints of each frequency regulation resource, and state coupling constraints of the energy storage device.

[0079] The power balance constraint requires that the sum of the planned regulation power of all frequency regulation resources in each time period be equal to the sum of the total demand command for automatic generation control from the upper-level dispatching agency in that time period and the primary frequency regulation demand power obtained based on the frequency deviation prediction in that time period.

[0080] The automatic generation control total demand instruction is the regional regulation total demand issued by the superior dispatching agency to this dispatching center. The primary frequency regulation demand power obtained based on frequency deviation prediction is predicted by this dispatching center based on the real-time frequency deviation of the power grid and the characteristics of load disturbance. The prediction of the primary frequency regulation demand power uses linear extrapolation or Kalman filtering algorithm to estimate the frequency change trend in the next few seconds to tens of seconds, and calculates the required compensation power in combination with the system inertia constant.

[0081] The operational constraints for each frequency regulation resource require that the planned regulation power of each frequency regulation resource in each time period must not exceed the upper and lower limits of the power of that resource in that time period determined based on the real-time status data collected in step S1. Furthermore, the change in planned regulation power between adjacent optimization time periods must not exceed the product of the allowable regulation rate of that resource and the duration of a single optimization time period. For thermal power units, the regulation rate is the regulation rate collected in step S1; for energy storage devices and hydropower units, it is the maximum power change rate allowed by their technical specifications. The unit of regulation rate is megawatts per minute.

[0082] For thermal power units and hydropower units, the upper and lower power limits are their minimum and maximum technical output, respectively. For energy storage devices, they are their maximum charging power and maximum discharging power. For energy storage devices, an additional energy storage state coupling constraint is added. This constraint requires that the current state of charge of the energy storage device in each time period is equal to the current state of charge of the previous time period minus the product of the planned adjustment power of the current time period and the length of a single optimization time period, divided by the rated capacity of the energy storage device. It also ensures that the current state of charge of the energy storage device in each time period is between its allowable minimum and maximum values.

[0083] The rated capacity of energy storage is the nominal total energy capacity of the energy storage device. The minimum and maximum allowable values ​​for the current state of charge are set according to the operating procedures of the energy storage device; usually, they are zero and one. To prevent overcharging and discharging, the minimum allowable value can be set to 0.1 and the maximum allowable value can be set to 0.9 in actual operation.

[0084] The dispatch center solves the real-time optimization model that satisfies all constraints. The optimal solution obtained includes the power command pre-allocation scheme for each time period and each frequency regulation resource within the future rolling time window, as well as the unified clearing price corresponding to each time period. The unified clearing price is the Lagrange multiplier corresponding to the power balance constraint at the optimal solution of the optimization problem. Its numerical meaning is the minimum marginal cost that the system needs to pay for each additional unit of power regulation demand under optimal scheduling.

[0085] In this embodiment, it is specifically noted that in step S3, the local control unit configured for each frequency modulation resource makes decisions as an independent intelligent agent. The state observed by each intelligent agent at the decision-making time includes the local state and the neighborhood state.

[0086] The local state includes: the power instruction pre-allocation value for the current time period corresponding to this resource in the power instruction pre-allocation scheme from step S2, and the corresponding unified clearing price for the current time period;

[0087] The real-time status data from step S1 includes, for thermal power units, the first-stage metal temperature of the high-pressure cylinder, and for energy storage devices, the current state of charge; as well as the locally measured or received grid frequency deviation.

[0088] The neighborhood state is a state summary information of adjacent frequency regulation resource agents obtained through communication, including the changing trend of the lifetime loss coefficient of adjacent thermal power units and the current state of charge of adjacent energy storage devices. The adjacent relationship can be predefined according to the electrical connection topology of the power grid or the dispatch management partition. For example, frequency regulation resources that are directly electrically connected or in the same substation are defined as adjacent. The communication cycle of the state summary information is on the order of seconds and is synchronized with the decision time. The local state and the neighborhood state together constitute the complete observation state of the agent.

[0089] Each agent, based on the complete observation state, uses its internal policy network to output a fine-tuning action for the pre-allocated power command value for the current time period. This fine-tuning action is a power offset. The power offset is added to the pre-allocated power command value for the current time period to obtain the final execution command of the frequency modulation resource controlled by the agent. The range of the power offset value is determined by the instantaneous adjustment capability of the frequency modulation resource at the decision time, for example, not exceeding one-tenth of its maximum adjustment rate multiplied by a control cycle.

[0090] Before generating fine-tuning actions, the policy network inside the intelligent body dynamically fuses local and neighborhood information using a graph attention mechanism; the specific process is as follows:

[0091] The agent first calculates the correlation weight between its own state and the state of each neighboring agent, which is called the graph attention weight. This graph attention weight represents the importance of the state information of each neighboring agent to the current decision.

[0092] The calculation process of graph attention weights is as follows: First, a shared linear transformation weight matrix is ​​used to linearly transform the agent's own local state vector and the state vectors of each neighboring agent. The shared linear transformation weight matrix is ​​randomly initialized in the early stage of training and is jointly learned and updated by all agents during training. The agent's own local state vector is composed of its own partial real-time state data from step S1 and the grid frequency deviation obtained by local measurement. Then, a trainable attention vector is concatenated with the vectors formed by the agent's own transformed state and the transformed states of each neighboring agent. A leaky linear rectifier activation function is applied to the inner product result. The negative slope coefficient in the leaky linear rectifier activation function can be set to 0.2. Then, an exponential operation with the natural constant base is performed on the activation result to obtain the unnormalized weights of each neighbor and the agent itself. Finally, all unnormalized weights are summed and each unnormalized weight is divided by the sum to obtain the normalized graph attention weights. The graph attention weights constitute a probability distribution, the sum of all its elements of which is 1.

[0093] Then, the agent performs a linear transformation on the state information of all neighboring agents and itself using a shared linear transformation weight matrix, and then performs a weighted summation according to the corresponding graph attention weights to obtain the aggregated collaborative perception information.

[0094] Ultimately, the policy network, based on the aggregated collaborative perception information and the agent's own local state, jointly decides to generate fine-tuning actions. The policy network is a multilayer perceptron neural network, and its output layer uses the hyperbolic tangent activation function to limit the fine-tuning action value within a preset power offset range. This allows each agent to selectively and adaptively consider the operating state of other frequency modulation resources associated with it when making local decisions, thereby achieving distributed collaboration.

[0095] The reward function used to drive the learning and updating of the agent's policy network consists of four parts:

[0096] The first part is the income incentive item, which is the product of the current clearing price and the value of the fine-tuning action performed by the agent, and is used to incentivize the agent to provide adjustment services.

[0097] The second part is the incremental cost item, which is the difference between the dynamic marginal cost incurred by the agent in executing the final execution instruction and the dynamic marginal cost incurred by only executing the pre-allocated power instruction. The dynamic marginal cost is obtained by calling the dynamic marginal cost calculation model in step S1. This item requires the agent to bear the actual adjustment cost when pursuing revenue. The specific calculation method of the incremental cost item is as follows: take the final execution instruction value as input, call the dynamic marginal cost calculation model to obtain a cost value; then take the pre-allocated power instruction value as input, call the same dynamic marginal cost calculation model to obtain another cost value; calculate the difference between the two cost values.

[0098] The third part is the quality penalty adjustment term, which is the product of a positive weighting coefficient and the square of the local frequency deviation. It is used to bind the agent's reward to the system goal of maintaining frequency stability. The weighting coefficient is used to adjust the proportion of the frequency deviation penalty in the total reward. The typical value range of this weighting coefficient is 0.01 to 0.1, and it needs to be determined through simulation debugging to balance the income incentive and the frequency stability goal.

[0099] The fourth part is the policy regularization term, which is used to constrain the deviation between the agent's current decision policy and a reference policy to ensure the stability of the learning process. The reference policy is a pre-defined benchmark decision policy, such as a policy that outputs zero fine-tuning actions. The policy regularization term quantifies the deviation by calculating the KL divergence between the probability distributions of the current policy and the reference policy and multiplying it by a regularization weight coefficient. The regularization weight coefficient can be set to 0.01 in the early stage of training and gradually decays as training progresses to prevent the policy from becoming rigid too early.

[0100] After each decision and execution, each agent shares its decision experience data to a central experience pool. The decision experience data includes the complete observation state at the moment of decision, the fine-tuning action executed, the immediate reward value calculated according to the reward function, and the complete observation state at the next moment after the action is executed. A central trainer uses a multi-agent reinforcement learning algorithm, which is either a multi-agent proximal policy optimization algorithm or a deep deterministic policy gradient algorithm, to periodically and centrally update the policy network parameters of all agents based on the reward function. The central update period is on the order of minutes, for example, every 5 minutes a mini-batch gradient descent is performed using the data in the experience pool to update all policy network parameters.

[0101] In this embodiment, it is particularly important to explain step S4, where the scheduling center collects the actual running data after the final execution instruction is executed, which specifically includes:

[0102] The actual execution power value corresponding to the final execution command of each frequency regulation resource generated in step S3, the actual operating cost calculated by actual fuel consumption measurement and equipment status monitoring, the actual frequency deviation and actual frequency change rate of the power grid, and data consistent with the real-time status data of each frequency regulation resource collected in step S1.

[0103] The dispatch center aligns all collected actual operation data with a unified time scale at the second level, and then processes it through sliding window filtering and outlier detection to form a time series sample set. The sliding window filtering can use a moving average filter with a length of 5 to 10 sampling points to smooth random noise. Outlier detection can use the three sigma principle based on historical data statistics to identify and remove abnormal data points that deviate significantly from the normal range.

[0104] The dispatch center invokes a newly constructed data value assessment network. This network calculates a scalar value score for each sample in the time series sample set. This value score is the product of three factors: the first factor is a stable value factor, which is calculated by using the natural constant e as the base and dividing the absolute value of the actual frequency deviation of the power grid corresponding to the negative sample by a frequency deviation normalization scaling factor as the exponent. This frequency deviation normalization scaling factor is set according to historical frequency statistics, for example, 0.1 Hz.

[0105] The second factor is the prediction error value factor, which is calculated as follows: add the absolute value of the difference between the actual operating cost of the sample and the predicted cost obtained after inputting the real-time status data of the same sample into the dynamic marginal cost calculation model in step S1, and then divide by the sum of the actual operating cost of the sample and a very small positive number. The very small positive number is used to prevent division by zero; it is usually taken as 10 to the power of negative 6.

[0106] The third factor is the exploratory value factor, which is calculated as follows: the number is added to the information entropy of the output fine-tuning action distribution after inputting the states of the same sample into the multi-agent reinforcement learning decision model in step S3. This information entropy represents the exploratory nature of the policy. The information entropy is obtained by calculating the log probability weighted sum of the action probability distribution of the policy network output layer. The scheduling center selects high-value data samples ranked between 20% and 30% of the value scores to form a high-value sample subset. The ranking is determined according to the order of value scores from high to low.

[0107] The process of online updating and feedback optimization of the dynamic marginal cost calculation model and the multi-agent reinforcement learning decision-making model using actual operational data is as follows:

[0108] First, the dynamic marginal cost calculation model is updated using the elastic weight joint algorithm. The loss function of the elastic weight joint algorithm contains two terms. The first term is the mean squared error term. The calculation process is as follows: In the high-value sample subset, the difference between the actual operating cost of each sample and the predicted cost obtained after inputting the same sample into the dynamic marginal cost calculation model to be updated is calculated. The square of this difference is taken. Then, the squared differences of all samples in the high-value sample subset are added together and divided by the total number of samples in the high-value sample subset to obtain the mean.

[0109] The second term is the elastic constraint term. The calculation process is as follows: For each trainable parameter of the dynamic marginal cost calculation model, calculate the difference between its current update value and its value before the current update (i.e., the old parameter value). Square this difference and multiply it by a diagonal element of the Fisher information matrix corresponding to that parameter. This diagonal element is calculated based on historical data before the current update and is used to quantify the importance of the parameter to the model's original predictive ability. Then, sum the results of all parameters calculated above and multiply by a preset elastic constraint weight coefficient. This weight coefficient is used to balance learning new data and retaining old knowledge. The elastic constraint weight coefficient can range from 0.5 to 2.0 and needs to be determined through cross-validation. The loss function is minimized using the gradient descent method to update the parameters of the dynamic marginal cost calculation model, achieving incremental model learning.

[0110] Secondly, the multi-agent reinforcement learning decision-making model is updated. The scheduling center divides the samples into several batches from easy to difficult based on the value scores or actual frequency deviations of the samples in the high-value sample subset. A course learning strategy is adopted, first using the simple batches of samples with higher value scores or smaller frequency deviations to train the model. After the model performance is stable, the difficult batches of samples with lower value scores or larger frequency deviations are gradually introduced for training, so that the policy network learns complex collaborative behaviors progressively. The batch division can be done using the quantile method. For example, the samples are divided into 4 batches according to the value scores, and the samples with the highest scores in the first batch are the first batch, and so on.

[0111] Meanwhile, based on the recent changes in the mean square value of the power grid frequency deviation relative to the target threshold, the weight coefficient of the adjustment quality penalty term in the reward function in step S3 is dynamically adjusted to achieve the adaptive learning objective.

[0112] The dynamic adjustment method for the weight coefficients of the quality penalty term is as follows: the old weight coefficients of the quality penalty term are added to an adaptive learning rate, and multiplied by the difference between the mean of the squared frequency deviation of the power grid within an evaluation time window and a target threshold for the mean square value of the frequency deviation; the adaptive learning rate is a small positive number, such as 0.001, to prevent excessive fluctuations in the weight coefficients; the evaluation time window length can be several minutes, such as five minutes; the target threshold for the mean square value of the frequency deviation is set according to the frequency control standard; for example, it can be set to 0.01 square Hertz.

[0113] After each model update, the dispatch center will evaluate the performance of the updated dynamic marginal cost calculation model and the multi-agent reinforcement learning decision model on an offline simulation environment or a period of historical operating data. The evaluation indicators include the prediction error reduction rate of the dynamic marginal cost calculation model and the improvement rate of the power grid frequency control performance indicators calculated based on the actual operating data collected in step S4. The performance improvement requirement can be set to a prediction error reduction rate or frequency control performance indicator improvement rate of not less than 2%. If the evaluation result does not meet the preset performance improvement requirement, a safety rollback mechanism will be triggered, the current model update will be abandoned, and the hyperparameters in subsequent model updates will be automatically adjusted. The adjustable hyperparameters include the selection ratio of high-value samples, the elastic constraint weight coefficient, and the learning rate of model training. The adjustment method can be to multiply the corresponding hyperparameter values ​​by a decay factor, such as 0.9, to adopt a more conservative strategy in subsequent updates.

[0114] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0115] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0116] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0117] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0118] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0119] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0120] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for coordinated control of AGC and primary frequency regulation in thermal power units, characterized in that, Specifically, the steps include the following: Step S1: The dispatch center collects real-time status data of various frequency regulation resources in the power grid, including thermal power units, energy storage devices and hydropower units, and calls the dynamic marginal cost calculation model to calculate the dynamic marginal cost curve of various frequency regulation resources based on the real-time status data. The dynamic marginal cost curve is used to characterize the economic cost of providing a unit of regulation power in the future period. Step S2: Based on the dynamic marginal cost curves of various frequency regulation resources obtained in Step S1, the dispatch center constructs a real-time optimization model with the goal of minimizing the total frequency regulation cost of the entire network and solves it to obtain the power command pre-allocation scheme and unified clearing price for each frequency regulation resource within a future rolling time window. Step S3: The local control unit configured for each frequency regulation resource receives the power command pre-allocation scheme and unified clearing price, and based on the pre-built multi-agent reinforcement learning decision model, with the goal of maximizing their respective long-term net benefits, performs distributed collaborative decision-making and dynamic fine-tuning on the received power command to generate the final execution command. Step S4: The scheduling center collects the actual running data after the final execution instruction is executed, and uses the actual running data to update and optimize the dynamic marginal cost calculation model and the multi-agent reinforcement learning decision model online.

2. The method for coordinated control of thermal power unit AGC and primary frequency regulation according to claim 1, characterized in that: In step S1, the process by which the dispatch center collects real-time status data of various frequency regulation resources in the power grid and uses the dynamic marginal cost calculation model to calculate the dynamic marginal cost curves of various frequency regulation resources is as follows: The real-time status data collected includes the real-time output power, regulation rate, main steam pressure, main steam temperature, reheat steam temperature and first-stage metal temperature of the high-pressure cylinder of the thermal power unit; the current state of charge, maximum charge and discharge power and current cycle life count of the energy storage device; and the current output, vibration zone boundary data and available regulating capacity of the reservoir of the hydropower unit. The process of calling the dynamic marginal cost calculation model and calculating the dynamic marginal cost curves of various frequency regulation resources based on real-time status data is as follows: For thermal power units, their dynamic marginal cost consists of fuel cost component and life loss cost component; the fuel cost component is based on the real-time coal consumption characteristic curve of the thermal power unit, and the marginal fuel cost is calculated by differentiating the real-time output power of the current thermal power unit. The calculation of the life loss cost component is achieved through a real-time life loss assessment model. This model takes the operating status sequence data of the thermal power unit within the recent time window as input. The operating status sequence data includes key metal temperature, main steam pressure, main steam temperature, reheat steam temperature, real-time output power, and regulation rate. The real-time life loss assessment model outputs a life loss coefficient. The life loss coefficient is multiplied by the unit capacity investment cost conversion factor to obtain the first product. The first product is then multiplied by the regulation severity additional factor to obtain the life loss cost component. The regulation severity additional factor is the sum of the product of a number 1, the regulation severity weighting coefficient, and the absolute value of the current reference power change rate. The current reference power change rate represents the severity of regulation.

3. The method for coordinated control of thermal power unit AGC and primary frequency regulation according to claim 2, characterized in that: The dynamic marginal cost of the thermal power unit is the sum of the marginal fuel cost, the lifetime loss cost component, and the operation and maintenance marginal cost constant. For energy storage devices, their dynamic marginal cost mainly consists of the cycle life cost component and the state of charge deviation penalty cost component. The cycle life cost component is calculated by dividing the single full cycle cost by the difference between the maximum tolerable cycle life and the current accumulated equivalent cycle number, and then multiplying it by the ratio of the current absolute value of the charge / discharge power to the current charge / discharge efficiency function. The penalty cost component for state of charge deviation is calculated by multiplying the penalty coefficient by the square of the difference between the current state of charge and the preset optimal state of charge. For hydropower units, the calculation of their dynamic marginal cost must avoid the vibration operating zone. When the current output of the hydropower unit is within the range formed by the lower limit and upper limit of the vibration zone power, the dynamic marginal cost is set to a very large positive number. When the current output of the hydropower unit is outside the range formed by the lower limit and upper limit of the vibration zone power, the dynamic marginal cost is obtained by multiplying the water price coefficient, the current net head, and the water consumption rate per unit power.

4. The method for coordinated control of thermal power unit AGC and primary frequency regulation according to claim 3, characterized in that: In step S2, the specific process by which the dispatch center constructs a real-time optimization model with the objective of minimizing the total frequency regulation cost of the entire network is as follows: The scheduling center uses a future rolling time window as the optimization period. The rolling time window is discretized into multiple consecutive optimization periods. An objective function is constructed, which minimizes the sum of the costs of all time periods and all frequency regulation resource calls within the future rolling time window. The cost of calling a single frequency regulation resource in a single time period is obtained by definite integration of the dynamic marginal cost curve of the frequency regulation resource in the time period obtained in step S1 from the power zero point to the planned regulation power value of the resource in the time period. The planned regulation power is the decision variable to be optimized. When calculating the sum of call costs, the scheduling center calculates a regulation contribution quality factor for each frequency regulation resource in each time period, multiplies the call cost by the corresponding regulation contribution quality factor, and then sums them up. The adjustment contribution quality factor is calculated based on the recent historical comprehensive frequency modulation performance index of the frequency modulation resource through an exponential function with the natural constant as the base. The exponent of this exponential function is the product of the negative sensitivity coefficient and the number minus the difference of the historical comprehensive frequency modulation performance index. The higher the historical comprehensive frequency modulation performance index, the less the cost of calling the resource is discounted in the optimization.

5. The method for coordinated control of thermal power unit AGC and primary frequency regulation according to claim 4, characterized in that: The constraints of the real-time optimization model include power balance constraints, operational constraints of each frequency regulation resource, and state coupling constraints of the energy storage device. The power balance constraint requires that the total planned regulation power equals the sum of the total demand under automatic generation control and the predicted primary frequency regulation demand. Operational constraints for each frequency regulation resource include upper and lower power limits based on real-time status data and ramping constraints based on regulation rate; For energy storage devices, additional coupling constraints and state of charge limit constraints are added to characterize the continuous change of the state of charge. The dispatch center solves the real-time optimization model that satisfies all constraints. The optimal solution obtained includes the power command pre-allocation scheme for each time period and each frequency regulation resource within the future rolling time window, as well as the unified clearing price corresponding to each time period. The unified clearing price is the Lagrange multiplier corresponding to the power balance constraint at the optimal solution of the optimization problem. Its numerical meaning is the minimum marginal cost that the system needs to pay for each additional unit of power regulation demand under optimal scheduling.

6. The method for coordinated control of thermal power unit AGC and primary frequency regulation according to claim 5, characterized in that: In step S3, the local control unit configured for each frequency modulation resource acts as an independent intelligent agent to make decisions. The state observed by each intelligent agent at the decision moment includes the local state and the neighborhood state. The local state includes: the power instruction pre-allocation value for the current time period corresponding to this resource in the power instruction pre-allocation scheme from step S2, and the corresponding unified clearing price for the current time period; The real-time status data from step S1 includes, for thermal power units, the first-stage metal temperature of the high-pressure cylinder, and for energy storage devices, the current state of charge; as well as the locally measured or received grid frequency deviation. The neighborhood state is the state summary information of adjacent frequency modulation resource agents obtained through communication, including the changing trend of the lifetime loss coefficient of adjacent thermal power units and the current state of charge of adjacent energy storage devices; the local state and the neighborhood state together constitute the complete observation state of the agent; Each agent, based on the complete observation state, uses its internal policy network to output a fine-tuning action for the pre-allocated power command value for the current time period. This fine-tuning action is a power offset. This power offset is added to the pre-allocated power command value for the current time period to obtain the final execution command for the frequency modulation resource controlled by the agent.

7. The method for coordinated control of thermal power unit AGC and primary frequency regulation according to claim 6, characterized in that: Before generating fine-tuning actions, the policy network within the intelligent body dynamically fuses local and neighboring information using a graph attention mechanism; the specific process is as follows: The agent first calculates the correlation weight between its own state and the state of each neighboring agent, which is called the graph attention weight. This graph attention weight represents the importance of the state information of each neighboring agent to the current decision. The calculation process of graph attention weights is as follows: First, using a shared linear transformation weight matrix, linear transformations are performed on the agent's own local state vector and the state vectors of each neighboring agent. Then, a trainable attention vector is concatenated with the transformed state of the agent and the transformed states of each neighboring agent, and an inner product operation is performed on the inner product result. A leaky linear rectified unit activation function is applied to the activation result, and then an exponential operation with the natural constant base is performed on the activation result to obtain the unnormalized weights of each neighbor and the agent itself. Finally, all unnormalized weights are summed, and each unnormalized weight is divided by the sum to obtain the normalized graph attention weights. Then, the agent performs a linear transformation on the state information of all neighboring agents and itself using a shared linear transformation weight matrix, and then performs a weighted summation according to the corresponding graph attention weights to obtain the aggregated collaborative perception information. Ultimately, the policy network, based on this aggregated collaborative perception information and the agent's own local state, jointly decides to generate fine-tuning actions; enabling each agent to selectively and adaptively consider the operating state of other frequency modulation resources associated with it when making local decisions, thereby achieving distributed collaboration.

8. The method for coordinated control of thermal power unit AGC and primary frequency regulation according to claim 7, characterized in that: The reward function used to drive the learning and updating of the agent's policy network consists of four parts: The first part is the income incentive item, which is the product of the current clearing price and the value of the fine-tuning action performed by the agent; The second part is the incremental cost term, which is the difference between the dynamic marginal cost generated by the agent executing the final execution instruction and the dynamic marginal cost generated by only executing the power instruction pre-allocated value. The specific calculation method of the incremental cost term is as follows: take the final execution instruction value as input, call the dynamic marginal cost calculation model to obtain a cost value; then take the power instruction pre-allocated value as input, call the same dynamic marginal cost calculation model to obtain another cost value; calculate the difference between the two cost values. The third part is the quality adjustment penalty term, which is the product of a positive weighting coefficient and the square of the local frequency deviation; The fourth part is the policy regularization term, which is used to constrain the deviation between the agent's current decision policy and a reference policy; After each decision and execution, each agent shares its decision experience data to a central experience pool, which is then trained by a central trainer using a multi-agent reinforcement learning algorithm. The multi-agent reinforcement learning algorithm is a multi-agent proximal policy optimization algorithm that periodically and centrally updates the policy network parameters of all agents based on the reward function.

9. The method for coordinated control of thermal power unit AGC and primary frequency regulation according to claim 8, characterized in that: In step S4, the scheduling center collects the actual running data after the final execution instruction is executed, specifically including: The dispatch center aligns all collected actual operational data with a unified timescale at the second level, and then processes it through sliding window filtering and outlier detection to form a time series sample set. The dispatch center calls a newly constructed data value assessment network. The data value assessment network calculates a scalar value score for each sample in the time series sample set. The value score is the product of three factors: the first factor is the stable value factor, which is calculated by using the natural constant as the base and the value of the absolute value of the actual frequency deviation of the power grid corresponding to the negative sample divided by a frequency deviation normalization scaling factor as the index. The second factor is the prediction error value factor, which is calculated as follows: add the absolute value of the difference between the actual operating cost of the sample and the predicted cost obtained after inputting the real-time status data of the same sample into the dynamic marginal cost calculation model in step S1, and then divide by the sum of the actual operating cost of the sample and a very small positive number. The third factor is the exploratory value factor, which is calculated by adding the information entropy of the output fine-tuning action distribution after inputting the states of the same sample into the multi-agent reinforcement learning decision model in step S3. This information entropy represents the exploratory nature of the strategy. The scheduling center selects high-value data samples ranked in the top 20% to 30% of the value scores to form a high-value sample subset.

10. The method for coordinated control of thermal power unit AGC and primary frequency regulation according to claim 9, characterized in that: The process of using actual operational data to perform online updates and feedback optimization of the dynamic marginal cost calculation model and the multi-agent reinforcement learning decision-making model is as follows: First, the dynamic marginal cost calculation model is updated by adopting the elastic weight joint algorithm. The loss function of the elastic weight joint algorithm includes two terms: the first term is the mean squared error term, and the second term is the elastic constraint term. Secondly, the multi-agent reinforcement learning decision-making model is updated. The scheduling center divides the samples into several batches from easy to difficult based on the value scores of the samples in the high-value sample subset. The course learning strategy is adopted. The model is first trained using simple batches of samples with higher value scores or smaller frequency deviations. After the model performance is stable, difficult batches of samples with lower value scores or larger frequency deviations are gradually introduced for training, so that the policy network learns complex collaborative behaviors progressively. Meanwhile, based on the recent changes in the mean square value of the power grid frequency deviation relative to the target threshold, the weight coefficient of the adjustment quality penalty term in the reward function is dynamically adjusted to achieve the adaptive learning objective. After each model update, the scheduling center will evaluate the performance of the updated dynamic marginal cost calculation model and the multi-agent reinforcement learning decision model in an offline simulation environment. If the evaluation results do not meet the preset performance improvement requirements, a safety rollback mechanism will be triggered to abandon the current model update and automatically adjust the hyperparameters in subsequent model updates.