Power grid energy scheduling method based on economic-environmental power and combined heat and power generation

Through the Lagrangian method and multi-agent distributed reinforcement learning algorithm, the grid scheduling is optimized, and the problems of high computational complexity and renewable energy volatility are solved, and the stability and efficiency of the power grid system are improved.

CN120471333APending Publication Date: 2025-08-12SHENYANG UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510478483.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing grid scheduling methods are highly complex in processing large-scale power grids, relying on accurate models to make mistakes, and are difficult to co-generation units to co-generation, resulting in system stability and inefficiency.

Method used

The Lagrangian method and multi-agent distributed reinforcement learning algorithm are adopted, combined with the critic neural network and the actor neural network, to optimize the incremental cost of the unit and achieve dynamic scheduling through collaborative learning among agents.

Benefits of technology

It improves the stability, economy and adaptability of the power grid system, can better match the power generation and heating needs of cogeneration units, and significantly improves the system's supply and demand balance capability and scheduling efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471333A_ABST
    Figure CN120471333A_ABST
Patent Text Reader

Abstract

The invention provides a power grid energy scheduling method based on economic-environmental power and combined heat and power generation, and belongs to the technical field of power system optimization scheduling. Quantifying the operation cost of each unit of the power grid system through the mathematical model; determining an objective function of an economic dispatching problem and an objective function of an economic-environment dispatching problem according to the quantized operation cost of each unit; in combination with constraint conditions of power grid operation, representing increment cost of each unit through a Lagrangian method when the power grid operation state is not matched with the power grid demand state; the reviewer neural network outputs a performance index based on a local supply and demand mismatching value of each unit operation state and a power grid demand state and a corresponding increment cost; outputting an operation decision of each unit based on a power grid demand state and a performance index output by a reviewer neural network through an actor neural network; and iterating through a multi-agent distributed reinforcement learning algorithm, and when the incremental cost is converged to be consistent, outputting an optimal control strategy of each unit. And the stability and adaptability of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power system optimization and dispatching, and in particular to a power grid energy dispatching method based on economic-environmental power and cogeneration of heat and power. Background Art

[0002] Economic dispatch (ED) and economic-environmental dispatch (EED) in power systems are core issues in optimizing grid operation. The goal of economic dispatch is to minimize power generation costs by rationally arranging the output of generators while meeting electricity demand. Economic-environmental dispatch, on the other hand, takes environmental factors into account, aiming to reduce pollutant emissions (such as carbon dioxide and sulfur oxides), thereby achieving dual optimization of economy and environmental protection. With the expansion of the power grid and the popularization of renewable energy, scheduling problems have become more complex, involving nonlinear constraints, multi-objective optimization, and uncertain factors (such as the volatility of wind power and photovoltaic power). Traditional scheduling problems are usually modeled as constrained optimization problems, the core of which is to find the optimal solution that meets constraints such as power balance and unit output limits through optimization algorithms.

[0003] Existing scheduling methods fall into two main categories: mathematical methods and meta-heuristic algorithms. Mathematical methods include linear programming, nonlinear programming, and quadratic programming. These methods are based on mathematical models and can quickly find optimal solutions in small-scale power grids. Meanwhile, meta-heuristic algorithms (such as genetic algorithms, particle swarm optimization, and differential evolution) are widely used due to their adaptability to complex problems. These algorithms do not rely on precise mathematical models, can handle nonlinear and non-convex optimization problems, and possess global search capabilities.

[0004] However, existing technologies still have many defects in practical applications. First, when dealing with large-scale power grids, the computational complexity of mathematical methods increases significantly, the algorithm efficiency decreases, and an accurate mathematical model is required. The model error will affect the optimization results. In addition, these methods are mostly based on centralized control and are highly dependent on communication signals. Once the communication link fails, it may cause system instability. Secondly, although the metaheuristic algorithm has strong adaptability, it cannot guarantee the global optimal solution and has poor convergence. In addition, when dealing with cogeneration unit constraints and large-scale power grids, the calculation time is long and the efficiency is low. Finally, the existing methods perform poorly in dealing with the uncertainty of renewable energy (such as the volatility of wind power and photovoltaics) and complex constraints (such as cogeneration units), and it is difficult to meet the real-time scheduling needs of modern power grids. These defects limit the application of traditional methods in practical engineering projects, and a more efficient and flexible optimization method is urgently needed. Summary of the Invention

[0005] In view of this, the present invention provides a grid energy scheduling method based on economic-environmental electricity and cogeneration, which optimizes the incremental cost of the unit by using the Lagrangian method and reinforcement learning algorithm; thereby improving energy utilization efficiency, thereby enhancing the system's adaptability to grid changes, and overcoming the limitations of traditional optimization methods that rely on precise grid models and face complex constraints and uncertainties.

[0006] To this end, the present invention provides the following technical solutions:

[0007] A method for dispatching power grid energy based on economic-environmental electricity and cogeneration of heat and power, comprising:

[0008] Quantify the operating costs of each unit in the power grid system through mathematical models;

[0009] The objective functions of the economic dispatch problem and the economic-environmental dispatch problem are determined based on the operating costs of each unit; and the incremental cost of each unit when the grid operation state does not match the grid demand state is obtained using the Lagrangian method in combination with the constraints of the grid operation.

[0010] The critic neural network outputs the strategic performance index of the actor neural network based on the local supply-demand mismatch value between the operating status of each unit and the grid demand status and the corresponding incremental cost;

[0011] The actor neural network outputs the operation decision of each unit based on the grid demand status and the performance indicators output by the critic neural network;

[0012] Through the iteration of the multi-agent distributed reinforcement learning algorithm, when the incremental cost converges to consistency, the optimal control strategy for each unit is output.

[0013] Furthermore, the objective function of the economic dispatch problem includes: the sum of the heating cost of the heating unit, the cogeneration cost of the cogeneration unit, and the power generation cost of the power generation unit considering only economic dispatch;

[0014] The objective function of the economic-environmental scheduling problem includes: the sum of the power generation cost of the generator set considering only economic scheduling and the scheduling cost of the generator set considering economic-environmental scheduling.

[0015] Furthermore, the constraints on the operation of the power grid include:

[0016] Electrical and thermal power balance constraints:

[0017]

[0018]

[0019] Among them, D P and D Hare power demand and heat demand, respectively; is the total transmission loss; B ij is the loss coefficient between the i-th and j-th units; P i min and P i max are the lower and upper limits of power generation respectively; and are the lower and upper limits of the heat output of the i-th unit respectively; N GH Indicates the number of cogeneration units; N H is the number of heating units; H i is the calorific value of the i-th cogeneration unit; p is the penalty coefficient; N G is the number of generator sets.

[0020] Furthermore, the incremental cost of each unit obtained by the Lagrangian method when the grid operation state does not match the grid demand state includes:

[0021] The incremental cost of the i-th generator set in the economic dispatch problem is:

[0022]

[0023] The incremental cost of the generator set in the i-th economic-environmental dispatch problem is:

[0024]

[0025] The incremental cost of power generation for the i-th cogeneration unit in the economic dispatch problem is:

[0026]

[0027] The incremental heating cost of the i-th cogeneration unit in the economic dispatch problem is:

[0028]

[0029] The incremental cost of heat generation for the heating unit in the i-th economic dispatch problem is:

[0030]

[0031] Among them, a i , b i , g i , h i , α i , β i , γ i , δ i , η i is a positive constant; is the i-th cost function; P i The power generated by the i-th unit; Pi min is the minimum power generation; H i is the heat generated by the i-th unit; f i Indicates P i and λ i The relationship between them.

[0032] Furthermore, the local supply-demand mismatch value includes:

[0033] Initialize unit information and local supply-demand mismatch value;

[0034] The local supply-demand mismatch value for the next step is calculated based on the incremental cost of the previous step, including:

[0035]

[0036] Among them, P i P (k) represents the power generated by the i-th generator set at the k-th step; P i P (k+1) represents the power generated by the i-th generator set at the k+1th step; P i CHP (k) is the power generation capacity of the i-th cogeneration unit at the k-th step; P i CHP (k+1) the power generation of the i-th CHP unit at the k+1th step; is the heating power of the i-th heating unit at the k-th step; is the heating power of the i-th heating unit at the k+1th step; is the heating power of the i-th cogeneration unit at the k-th step; is the heating power of the i-th heating unit at the k+1th step; i PEED (k) is the incremental electricity cost of the i-th generator at step k; is the incremental electricity cost of the i-th CHP unit at the k-th step; is the incremental thermal cost of the i-th CHP unit at step k; is the incremental thermal cost of the i-th heating unit at step k; y i (k) is the local power supply and demand mismatch value of the i-th generator set at the k-th step; y i (k+1) is the local power supply and demand mismatch value of the i-th generator group at the k+1th step; is the local power supply and demand mismatch value of the i-th CHP unit at the k-th step; is the local power supply and demand mismatch value of the i-th CHP unit at the k+1th step; is the local heat supply and demand mismatch value of the i-th CHP unit at the k-th step; is the local heat supply and demand mismatch value of the i-th CHP unit at the k+1th step; The local heat supply and demand mismatch value of the i-th heating unit at the k-th step; The local heat supply and demand mismatch value of the i-th heating unit at the k+1th step; Q(i,j) is the consensus weight matrix element, which represents the connection weight between node i and neighbor j; L P (k) is the total transmission loss at step k.

[0037] Furthermore, the strategic performance indicators of the actor neural network include:

[0038]

[0039] in, and are the weight matrices of the first and second layers respectively; u c (k) is the critic neural network input; ξ i (k) is a random disturbance term.

[0040] Advantages and positive effects of the present invention:

[0041] The method of the present invention combines reinforcement learning, Lagrangian method and distributed control structure. Through collaborative learning and decision-making among intelligent agents, it achieves dynamic optimization of the objective function and precise satisfaction of the constraints, significantly improving the stability, economy and adaptability of the system.

[0042] 1) By clarifying the objective function and constraints, the system provides optimization directions and operational boundaries for grid dispatch. By comprehensively addressing the relationship between the objective function and constraints using the Lagrangian method, local optimality is avoided, and the incremental cost of each unit is accurately calculated, providing a reliable basis for the intelligent agent's decision-making and optimization.

[0043] 2) By calculating the local supply-demand mismatch, we can accurately determine the deviation between the power generation and heating of the CHP units and demand, providing a basis for targeted optimization in subsequent iterations. This allows the CHP units' power generation and heating power to better match demand, significantly improving the system's supply-demand balancing capabilities.

[0044] 3) Utilizing the critic neural network to estimate performance metrics and the actor neural network to calculate group incremental costs, the decision-making process is continuously optimized through iterative training. The introduction of random perturbations increases the model's randomness and exploratory nature, preventing it from falling into local optima and enabling it to explore a more optimal parameter space, thereby maintaining stable system operation.

[0045] 4) A reward-penalty mechanism was introduced to incentivize the unit to quickly learn the optimal decision-making strategy, avoiding lingering in ineffective or harmful decision spaces. This mechanism significantly accelerated the algorithm's convergence process, enabling the system to more quickly reach a stable, optimized operating state and improving overall efficiency.

[0046] 5) Through multiple iterations of a multi-agent distributed reinforcement learning algorithm, the incremental cost of each unit at the optimal power generation and combined heat and power output was ultimately determined, accurately reflecting the cost required to increase unit output. This result provides an important basis for further analysis, such as system cost accounting and resource allocation, and enhances the cost-effectiveness and practicality of the method. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0048] Figure 1 Flowchart of a method for dispatching power grid energy based on economic-environmental electricity and cogeneration of heat and power in an embodiment of the present invention;

[0049] Figure 2 This is a multi-agent signal diagram based on a neural network algorithm in an embodiment of the present invention;

[0050] Figure 3 This is a flowchart of a power dispatch optimization method based on a multi-agent distributed reinforcement learning algorithm in an embodiment of the present invention. DETAILED DESCRIPTION

[0051] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0052] It should be noted that the terms "first," "second," and the like in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or steps is not necessarily limited to those steps or steps explicitly listed, but may include other steps or steps that are not explicitly listed or that are inherent to these processes, methods, products, or apparatuses.

[0053] Traditional methods have many shortcomings in the economic dispatch and economic-environmental dispatch problems of power projects, such as reliance on power grid models and centralized control that is susceptible to communication failures. The present invention provides an energy dispatch method for economic-environmental power and cogeneration based on multi-agent distributed reinforcement learning. It constructs a mathematical model for power grid dispatch and uses the Lagrangian method combined with fuzzy adjustment factors to calculate the incremental cost of the unit, making it dynamically adjustable. Parameters such as unit power are randomly initialized and the local supply and demand mismatch value is calculated. Dynamic risk assessment indicators and risk aversion coefficients are introduced to define unit performance indicators. Performance indicators are estimated through a critic neural network, and incremental costs are calculated through an actor neural network. Both are iteratively optimized and random perturbations are introduced to avoid local optimality. Each unit interacts with the environment, and a reward-penalty strategy is used to increase the algorithm speed. After iteration of the multi-agent distributed reinforcement learning algorithm, the optimal power generation and heating power and incremental cost are output, achieving global optimality, improving system security and adaptability, and being suitable for energy dispatch optimization in various power grid scenarios.

[0054] Combine Figure 1 As shown, the method of the present invention is further described with specific implementation:

[0055] S101. Establish a mathematical model for energy dispatch in the power grid, clarify the objective function and constraints, and use the Lagrangian method to calculate the incremental cost of each unit.

[0056] Quantify the operating costs of each unit in the power grid system through mathematical models;

[0057] For power units, the cost function takes into account not only the power generation cost but also the emission cost; for cogeneration units, the cost function covers the power generation and heating costs as well as the heat-power coupling cost; for heating units, the cost function is the heat generation cost; for economic dispatch problems, the objective function includes the power generation and / or heating costs of all units; for economic-environmental dispatch problems, the objective function takes into account not only the power generation cost of the generator units but also the emission cost;

[0058] 1) In this embodiment, the power generation cost of the i-th generator set in the economic dispatch problem is quantified as:

[0059]

[0060] The emission cost of the i-th generator unit in the economic-environmental scheduling problem is quantified as:

[0061]

[0062] The cogeneration cost of the i-th cogeneration unit is quantified as:

[0063]

[0064] The heating cost of the i-th heating unit is quantified as:

[0065]

[0066] Among them, N GH and N H are the number of combined heat and power generation and heating units respectively; H i is the calorific value of the i-th cogeneration unit. In this embodiment, the objective function of the economic dispatch problem is:

[0067]

[0068] The objective function of the economic-environmental scheduling problem is:

[0069]

[0070] Among them, p is the penalty coefficient; N G is the number of generator sets.

[0071] 2) In this embodiment, the operating constraints of each unit include:

[0072] Constraints on power generation by generator sets and heat generation by heating units include:

[0073] P i min ≤P i ≤P i max ,i=1,2,…,N G

[0074]

[0075] Combined constraints on power generation and heat generation from a cogeneration unit, including:

[0076] P i min (H i )≤P i≤P i max (H i ),i=1,2,…,N GH

[0077]

[0078] Among them, P i min and P i max are the lower and upper limits of power generation respectively; and are the lower limit and upper limit of the heat capacity of the i-th unit respectively.

[0079] In this embodiment, the power and heat balance constraints are:

[0080]

[0081] Among them, D P and D H are power demand and heat demand, respectively; is the total transmission loss; B ij is the loss coefficient between the ith and jth units.

[0082] 3) In this embodiment, the incremental cost of each unit is calculated using the Lagrangian method;

[0083] The incremental cost of the i-th generator set in the economic dispatch problem is:

[0084]

[0085] The incremental cost of power generation for the i-th cogeneration unit in the economic dispatch problem is:

[0086]

[0087] The incremental heating cost of the i-th cogeneration unit in the economic dispatch problem is:

[0088]

[0089] The incremental cost of heat generation for the heating unit in the i-th economic dispatch problem is:

[0090]

[0091] The incremental cost of the generator set in the i-th economic-environmental dispatch problem is:

[0092]

[0093] Among them, a i , b i , gi , h i , α i , β i , γ i , δ i , η i is a positive constant; is the i-th cost function; P i The power generated by the i-th unit; P i min is the minimum power generation; H i is the heat generated by the i-th cogeneration unit; f i Indicates P i and λ i The relationship between them.

[0094] S102. Randomly initialize the generating power, heating power, and incremental cost of all units;

[0095] The initial local supply-demand mismatch value is calculated based on the initial power and demand; then the power generation and heat generation power of the next step and the local supply-demand mismatch value between the demand and the total power generation or heat generation are calculated through the incremental cost of the previous step.

[0096] According to the physical characteristics and operation requirements of the unit, the power generation power of each unit is randomly set to an initial value within its allowed value range to provide the initial state for subsequent iterative calculations;

[0097] Each unit calculates the local supply-demand mismatch value. The degree of difference between the local supply-demand mismatch value and the average demand is an important basis for subsequent optimization and adjustment.

[0098] In this embodiment, the unit information and the local supply-demand mismatch value are initialized:

[0099]

[0100] Where D represents the power demand in the economic-environmental scheduling problem; D P represents the power generation demand in the economic dispatch problem; D H Heating power demand in economic dispatch problem; P i (0) is the initial value of the power generation of the i-th generator set; H i (0) is the initial value of the heating power of the i-th heating unit; λ i (0) is the initial value of the incremental cost of the i-th unit; y i (0) is the mismatch between the generating power or heating power of the i-th unit and the average demand at the initial moment.

[0101] In this embodiment, the power generation and heating power in the next step and the local supply-demand mismatch between the demand and the total power generation or heating power are calculated by the incremental cost of the previous step:

[0102] The mismatch between the generating capacity of the generator set and the local supply and demand:

[0103]

[0104] The mismatch between the power generation capacity of the cogeneration unit and the local supply and demand, as well as the mismatch between the heat generation capacity of the cogeneration unit and the local supply and demand:

[0105]

[0106] The mismatch between the heating power of the heating unit and the local supply and demand:

[0107]

[0108] Among them, P i P (k) represents the power generated by the i-th generator set at the k-th step; P i P (k+1) represents the power generated by the i-th generator set at the k+1th step; P i CHP (k) is the power generation of the i-th cogeneration unit at the k-th step; P i CHP (k+1) the power generation of the i-th CHP unit at the k+1th step; is the heating power of the i-th heating unit at the k-th step; is the heating power of the i-th heating unit at the k+1th step; is the heating power of the i-th cogeneration unit at the k-th step; is the heating power of the i-th heating unit at the k+1th step; i PEED (k) is the incremental electricity cost of the i-th generator at step k; is the incremental electricity cost of the i-th CHP unit at the k-th step; is the incremental thermal cost of the i-th CHP unit at step k; is the incremental thermal cost of the i-th heating unit at step k; y i (k) is the local power supply and demand mismatch value of the i-th generator set at the k-th step; y i (k+1) is the local power supply and demand mismatch value of the i-th generator group at the k+1th step; is the local power supply and demand mismatch value of the i-th CHP unit at the k-th step; is the local power supply and demand mismatch value of the i-th CHP unit at the k+1th step; is the local heat supply and demand mismatch value of the i-th CHP unit at the k-th step; is the local heat supply and demand mismatch value of the i-th CHP unit at the k+1 step; The local heat supply and demand mismatch value of the i-th heating unit at the k-th step; The local heat supply and demand mismatch value of the i-th heating unit at the k+1th step; Q(i,j) is the consensus weight matrix element, which represents the connection weight between node i and neighbor j; L P (k) is the total transmission loss at step k.

[0109] S103, defining the performance index of each unit based on the mismatch value and the incremental cost;

[0110] The critic neural network outputs performance indicators; the actor neural network outputs the decision of each unit and the incremental cost of each unit.

[0111] For each unit in the power grid system, its performance indicators are defined by comprehensively considering the unit's own status and the relationship with adjacent units; and a dynamic risk assessment indicator is introduced. Through this penalty, the performance indicator is lowered to guide adjustment decisions to reduce risks.

[0112] In this embodiment, the performance index of each unit is:

[0113]

[0114] Dynamic risk assessment indicators:

[0115]

[0116] in, e ai (k) represents the actor neural network error; e ai (k) = J i (k-1)-y j (k-1);λ i represents the incremental cost of the i-th unit; y i represents local estimation; k represents the number of algorithm iterations; β is the risk aversion coefficient; R i (t) is the dynamic risk assessment index; F i (t) is the failure rate of unit i at time t; L i (t) is the load fluctuation degree of unit i at time t; M i (t) is the maintenance demand level of unit i at time t; It is a preset weight coefficient used to adjust the impact of each factor on risk assessment.

[0117] The performance indicators of each unit guide the unit to adjust its own behavior in the direction of optimizing the overall performance of the system during the iteration process, prompting the unit to adjust the incremental cost to achieve better coordinated operation.

[0118] Through the calculation of the critic neural network, the output is the estimated value of the performance indicator. Through continuous iterative training, the critic neural network can estimate the performance indicator more accurately and provide an important reference for the team's decision-making; the actor neural network continuously adjusts the weights and calculates the appropriate incremental cost based on the system's state information, enabling the team to make the best decision based on its own situation and the situation of adjacent teams.

[0119] Combine Figure 2 , the multi-agent signal transmission process based on the neural network algorithm in this embodiment is described:

[0120] 1) When the system is initialized, the agent receives the initial incremental cost signal and the initial local mismatch signal.

[0121] 2) The actor neural network generates an initial incremental cost based on the initial incremental cost signal and the initial local mismatch signal and feeds it back into the system.

[0122] 3) The critic neural network evaluates the decision quality of the actor neural network based on the initial incremental cost signal and the initial local mismatch signal.

[0123] 4) Based on the evaluation results of the critic neural network, the actor neural network adjusts its weights to optimize the decision-making in the next round.

[0124] 5) This process is iterated until the incremental cost converges to a consistent solution (global optimum).

[0125] Each agent receives two main input signals:

[0126] ① Incremental cost signal: represented by a solid line, reflecting the changes in costs in the system.

[0127] ② Local mismatch signal: represented by a thick dotted line, which may represent the difference between the local state of the agent and the expected state of the system.

[0128] In this example, the critic neural network outputs an estimate of the performance indicator:

[0129]

[0130] in, and are the weight matrices of the first and second layers respectively; u c (k) is the critic neural network input; ξ i(k) is a random disturbance term.

[0131] The actor neural network calculates the incremental cost for each agent:

[0132]

[0133] in, and are the weight matrices of the first and second layers respectively; u a (k) is the actor neural network input.

[0134] S104. Each intelligent agent interacts with the environment and continuously optimizes its own decision-making (power generation and heat generation), while using a reward and punishment mechanism to increase the speed of the algorithm.

[0135] Each unit in the power grid acts as an intelligent agent, interacting with the environment. During the iteration process, each unit calculates performance indicators. The critic neural network estimates the performance indicators based on the input and calculates the error, and updates the weights through error backpropagation. The actor neural network calculates the incremental cost based on the critic neural network output and performance indicators, and updates the weights based on the error. Each unit continuously adjusts the power generation and heat generation, gradually approaching the minimum value of the system objective function.

[0136] At the same time, in order to improve the convergence speed of the algorithm, a reward-penalty strategy is introduced, which uses the degree of matching between generated power and demand as a reward, and the difference between incremental cost and adjacent intelligent agents as a penalty to improve the learning speed of the algorithm.

[0137] S105. Through the iteration of the multi-agent distributed reinforcement learning algorithm, the incremental cost of each unit converges to a consistent (global optimum) to minimize the system objective function and output the optimal power generation power or heat generation power of each unit.

[0138] Through the continuous iterative optimization of the multi-agent distributed reinforcement learning algorithm steps until the preset incremental cost converges to a consensus (global optimum), the optimal power generation and (or) heating power of each unit is output, realizing the energy optimization scheduling of economic problems and economic-environmental problems in the power grid.

[0139] Combine Figure 3 In this embodiment, the power dispatch optimization process based on the multi-agent distributed reinforcement learning algorithm includes:

[0140] (1) First, perform the initialization operation to set the initial agent parameters and weights.

[0141] (2) Start iteration. In each iteration:

[0142] The grid calculates power and heat generation and their mismatch values;

[0143] The actor neural network updates the weights and calculates the new incremental cost;

[0144] The critic neural network outputs performance metrics and updates weights;

[0145] (3) Iterations continue until all units are calculated and the maximum number of iterations is reached.

[0146] The method of the present invention has significant advantages in technical effects. Through collaborative learning and decision-making among intelligent agents, the system operation data (power generation, incremental cost) is processed with the help of a neural network architecture (critic neural network and actor neural network), thereby realizing optimal control of complex systems. The critic neural network is responsible for evaluating the strategic performance of the intelligent agent, while the actor neural network dynamically adjusts the strategy according to the evaluation results. This decoupling mechanism enables each intelligent agent to learn independently and optimize collaboratively, thereby effectively improving the accuracy and adaptability of the system scheduling strategy without relying on large-scale hardware modifications. In power engineering, the method of the present invention significantly improves the flexibility and economy of the power grid by optimizing the power generation scheduling of the unit, while reducing operating costs. The method of the present invention also has good scalability and versatility, and can be widely used in complex system optimization problems in the fields of electrical and mechanical fields, providing an efficient and reliable technical approach for collaborative decision-making of multi-agent systems.

[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for dispatching power grid energy based on economic-environmental electricity and cogeneration of heat and power, characterized in that: include: Quantify the operating costs of each unit in the power grid system through mathematical models; The objective functions of the economic dispatch problem and the economic-environmental dispatch problem are determined based on the operating costs of each unit; and the incremental cost of each unit when the grid operation state does not match the grid demand state is obtained using the Lagrangian method in combination with the constraints of the grid operation. The critic neural network outputs the strategic performance index of the actor neural network based on the local supply-demand mismatch value between the operating status of each unit and the grid demand status and the corresponding incremental cost; The actor neural network outputs the operation decision of each unit based on the grid demand status and the performance indicators output by the critic neural network; Through the iteration of the multi-agent distributed reinforcement learning algorithm, when the incremental cost converges to consistency, the optimal control strategy for each unit is output.

2. The method for dispatching power grid energy based on economic-environmental electricity and cogeneration according to claim 1, characterized in that: The objective function of the economic dispatch problem includes: the sum of the heating cost of the heating unit, the cogeneration cost of the cogeneration unit, and the power generation cost of the power generation unit considering only economic dispatch; The objective function of the economic-environmental scheduling problem includes: the sum of the power generation cost of the generator set considering only economic scheduling and the scheduling cost of the generator set considering economic-environmental scheduling.

3. The method for dispatching power grid energy based on economic-environmental electricity and cogeneration of heat and power according to claim 1, characterized in that: The constraints on the operation of the power grid include: Electrical and thermal power balance constraints: Among them, D P and D H are power demand and heat demand, respectively; is the total transmission loss; B ij is the loss coefficient between the i-th and j-th units; P i min and P i max are the lower and upper limits of power generation respectively; and are the lower and upper limits of the heat output of the i-th unit respectively; N GH Indicates the number of cogeneration units; N H is the number of heating units; H i is the calorific value of the i-th cogeneration unit; p is the penalty coefficient; N G is the number of generator sets.

4. The method for dispatching power grid energy based on economic-environmental electricity and cogeneration according to claim 1, characterized in that: The incremental cost of each unit obtained by the Lagrangian method when the grid operation state does not match the grid demand state includes: The incremental cost of the i-th generator set in the economic dispatch problem is: The incremental cost of the generator set in the i-th economic-environmental dispatch problem is: +p(2a i P i +b i +n i d i exp(δ i P i )) The incremental cost of power generation for the i-th cogeneration unit in the economic dispatch problem is: The incremental heating cost of the i-th cogeneration unit in the economic dispatch problem is: The incremental cost of heat generation for the heating unit in the i-th economic dispatch problem is: Among them, a i , b i , g i , h i , α i , β i , γ i , δ i , η i is a positive constant; is the i-th cost function; P i The power generated by the i-th unit; P i min is the minimum power generation; H i is the heat generated by the i-th unit; f i Indicates P i and λ i The relationship between them.

5. The method for dispatching power grid energy based on economic-environmental electricity and cogeneration of heat and power according to claim 1, characterized in that: The local supply-demand mismatch value includes: Initialize unit information and local supply-demand mismatch value; The local supply-demand mismatch value for the next step is calculated based on the incremental cost of the previous step, including: Among them, P i P (k) represents the power generated by the i-th generator set at the k-th step; P i P (k+1) represents the power generated by the i-th generator set at the k+1th step; P i CHP (k) is the power generation capacity of the i-th cogeneration unit at the k-th step; P i CHP (k+1) the power generation of the i-th CHP unit at the k+1th step; is the heating power of the i-th heating unit at the k-th step; is the heating power of the i-th heating unit at the k+1th step; is the heating power of the i-th cogeneration unit at the k-th step; is the heating power of the i-th heating unit at the k+1th step; is the incremental electricity cost of the i-th generator at the k-th step; is the incremental electricity cost of the i-th CHP unit at the k-th step; is the incremental thermal cost of the i-th CHP unit at step k; is the incremental thermal cost of the i-th heating unit at step k; y i (k) is the local power supply and demand mismatch value of the i-th generator set at the k-th step; y i (k+1) is the local power supply and demand mismatch value of the i-th generator group at the k+1th step; is the local power supply and demand mismatch value of the i-th CHP unit at the k-th step; is the local power supply and demand mismatch value of the i-th CHP unit at the k+1th step; is the local heat supply and demand mismatch value of the i-th CHP unit at the k-th step; is the local heat supply and demand mismatch value of the i-th CHP unit at the k+1 step; The local heat supply and demand mismatch value of the i-th heating unit at the k-th step; The local heat supply and demand mismatch value of the i-th heating unit at the k+1th step; Q(i,j) is the consensus weight matrix element, which represents the connection weight between node i and neighbor j; L P (k) is the total transmission loss at step k.

6. A method for dispatching power grid energy based on economic-environmental electricity and cogeneration of heat and power according to claim 5, characterized in that: The policy performance indicators of the actor neural network include: in, and are the weight matrices of the first and second layers respectively; u c (k) is the critic neural network input; ξ i (k) is a random disturbance term.