Cogeneration system active environment coordination control method based on distributed reinforcement learning
By using distributed reinforcement learning methods and neural networks to estimate the performance indicators of cogeneration units, adaptive, robust and efficient active power environment coordination control of cogeneration systems is achieved. This solves the problems of high signal path sensitivity and low computational efficiency in traditional control strategies, and ensures the global optimal scheduling of large-scale power grids.
Patent Information
- Application Number
- CN202510929058.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-31
AI Technical Summary
Traditional cogeneration systems rely on centralized control, which makes the control signal path highly sensitive to connectivity issues and prone to interruption due to faults. Furthermore, the computational efficiency decreases as the power grid size increases, making it difficult to achieve globally optimal scheduling in large-scale power grids.
A distributed reinforcement learning approach is adopted, treating the cogeneration units as independent intelligent agents. The performance indicators of each unit are estimated through neural networks. Combined with adaptive control and optimal control strategies, the calculation process is decomposed into multiple modules, and iterative optimization is performed using reward and penalty values.
It improves the robustness and safety of cogeneration systems, enables them to adapt to changes in grid topology, ensures optimal global scheduling, enhances computational efficiency, and solves the interruption problem caused by control signal path failures in traditional methods.
Smart Images

Figure CN120879816A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power system operation and control technology, and specifically relates to a method for coordinated control of active power environment in cogeneration systems based on distributed reinforcement learning. Background Technology
[0002] Under the constraints of dual carbon targets, it is urgent to improve the overall energy efficiency and tap the potential for emission reduction through combined heat and power (CHP) systems. A CHP system includes a fuel supply system, a power generation system, a heat supply system, and a control system. The fuel supply system provides fuel, which is burned to produce high-temperature, high-pressure gas that enters the power generation system and is finally used to generate electricity. During this process, the heat supply system utilizes waste heat for heating, water supply, or other purposes. By using waste heat recovery technology, the efficiency of heat supply is increased, providing not only electricity but also meeting the demand for heating.
[0003] The advantage of combined heat and power (CHP) units is their ability to generate electricity and heat simultaneously with an efficiency of nearly 90%. However, they also face the problem of mutual constraints between power generation and heating capacity, which needs to be considered during production and operation. Active power environment dispatch often employs mathematical optimization methods and metaheuristic algorithms. However, metaheuristic algorithms are not conducive to obtaining globally optimal results, mathematical optimization methods require accurate grid models, and the computational efficiency of these algorithms decreases significantly with the growth of the grid. Furthermore, traditional algorithms all employ centralized control strategies; if one or more control signal paths are disconnected due to a fault, the optimization algorithm will be interrupted.
[0004] Therefore, this invention proposes a coordinated control method for the active power environment of a combined heat and power (CHP) system based on distributed reinforcement learning. The aim is to efficiently solve the active power environment scheduling problem of the CHP system, address the issue of high sensitivity to the connectivity of control signal paths in traditional centralized control strategies, improve computational efficiency, and adapt to changes in grid topology without relying on grid models. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a coordinated control method for the active power environment of a combined heat and power (CHP) system based on distributed reinforcement learning. The method establishes an active power environment scheduling model for the CHP system, determines the objective function for optimizing the system's operating state and minimizing carbon dioxide emissions, and provides constraints based on system active power and heat balance, and unit active power and heat output limits to obtain the marginal function of the CHP unit's active power environment scheduling function. To improve computational efficiency and the robustness and security of active power environment scheduling, a solution strategy based on distributed reinforcement learning is determined. The performance indicators of each unit are identified, and the performance indicators and neural network errors are iteratively updated to optimize the neural network parameters and obtain the optimal command parameters.
[0006] This invention provides a method for coordinated control of the active power environment in a combined heat and power system based on distributed reinforcement learning, comprising the following steps: S1. Establish an active power environment scheduling model for a combined heat and power system; The expression for the active power environmental dispatch model of a combined heat and power system considering environmental effects is as follows: (1); in, For thermal power units Coal consumption, For thermal power units carbon emission coefficient, For combined heat and power units Coal consumption, For pure heating units Coal consumption, For the number of combined heat and power units, This refers to the number of units used solely for heating. This refers to the number of thermal power units. For the unit Those who have made contributions For the unit The heat emitted, As a penalty factor; The constraints are given based on the system power balance and the unit output limits, respectively, to obtain the first... The marginal function of the active power environment dispatch function of a thermal power generating unit, a combined heat and power unit, and a pure heating unit; S2. Determine the solution strategy based on distributed reinforcement learning; S21. Initialize relevant parameters; The generating output and marginal function of the units are initialized based on actual data, and the load is evenly distributed to the generating load adjustment margin of each unit according to the unit capacity. S22. Determine the unit performance indicators; Local performance indicators are used to optimize the power generation output of each unit, and their expressions are as follows: (2); in, For the unit marginal function, For the unit The power generation load regulation margin The time interval for updating the unit's performance indicators. For the number of iterations, For the unit Performance metrics; S3, iteratively update performance metrics and neural network errors; After the training rounds are completed, backpropagation is performed with the goal of maximizing overall performance. This updates the performance metrics and neural network error, optimizes the neural network parameters, and obtains the optimal instruction parameters.
[0007] Preferably, step S1 specifically includes the following steps: S11. Determine the objective function for active power environment scheduling; When considering environmental effects, the objective function expression for active power environmental scheduling of cogeneration units is: (3); in, For coefficients; S12. Determine the constraints of the objective function; S13. Solve for the marginal function of the active power environment scheduling function; No. The marginal function expression of the active power environment dispatch function of the combined heat and power unit is: (4); (5); In a preferred embodiment, step S12 specifically includes the following steps: S121. Regarding active power balance and heat supply balance constraints: (6); (7); in, To meet the system's active power load requirements, To meet the system's heat load requirements, For system transmission losses; S122. Constraints are determined for the limits of active power output and heating output of the unit: (8); (9); in, For the unit The upper limit of contribution. For the unit The lower limit of meritorious service For the unit The lower limit of heating output For the unit The upper limit of contribution.
[0008] In a preferred embodiment, the performance index in expression (2) of step S22 is specifically: (10); in, For the unit The marginal function is the marginal function of the unit. The error of the actor neural network is expressed as: (11).
[0009] In a preferred embodiment, in step S3, which iteratively updates the performance index and neural network error, the output of the commentator neural network is the estimated value of the performance index in expression (2), and the estimation error is: (12); in, For the unit The estimated value of the performance index is expressed as follows: (13); in, Here is the weight matrix of the first layer of the commenter neural network. This is the weight matrix of the second-layer commenter neural network. Input to the commentator's neural network, For the commentator's neural network function.
[0010] Furthermore, the weight matrix of the second-layer neural network is trained using the gradient descent algorithm based on the error value estimated by the performance index, and the expression is: (14); in, This represents the learning rate of the commentator's neural network.
[0011] Furthermore, the actor neural network is trained based on expression (11), and its output expression is: (15); in, As input to the actor neural network, This is the weight matrix of the first layer of the actor neural network, which can be set as a random matrix. This is the weight matrix of the second-layer actor neural network. For the actor neural network function.
[0012] In a preferred embodiment, in step S3, the first In the next iteration, the unit marginal function of the active environment scheduling function The iteration stops when the optimal value is reached.
[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The active power environment coordination control method of the cogeneration system based on distributed reinforcement learning of the present invention adopts a distributed control strategy. By treating the cogeneration units as independent intelligent agents, the performance indicators of each unit are defined and estimated by neural networks. Each unit is relatively independent and has low sensitivity to the connectivity of control signal paths. This solves the problem of optimization algorithm interruption caused by the failure of one or more control signal paths, and improves the robustness and security of the active power environment scheduling method.
[0014] (2) The active power environment coordination control method for cogeneration system based on distributed reinforcement learning of the present invention adopts a combination strategy of adaptive control and optimal control in the reinforcement learning algorithm. It can adapt to the changes in the power grid topology, does not depend on the power grid model, and can ensure that the optimization result is globally optimal. The test results on the actual system prove the effectiveness of the proposed method.
[0015] (3) The active power environment coordination control method of the cogeneration system based on distributed reinforcement learning of the present invention uses both reward value and penalty value in the training process of reinforcement learning algorithm. Moreover, the distributed control strategy decomposes the calculation process into multiple modules, improves the calculation efficiency of the algorithm, and solves the problem that the calculation efficiency of traditional methods decreases significantly with the growth of power grid scale. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the active power environment coordination control method for a combined heat and power system based on distributed reinforcement learning, as described in this invention. Figure 2 The active power environment coordination control method for cogeneration systems based on distributed reinforcement learning, as described in this invention, is the scheduling result of a power grid. Figure 3 This is the convergence curve of the marginal function of the active power environment coordination control method for cogeneration systems based on distributed reinforcement learning, as presented in this invention. Detailed Implementation
[0017] To provide a detailed description of the technical content, objectives, and effects of this invention, the following description will be provided in conjunction with the accompanying drawings.
[0018] This invention provides a method for coordinated control of the active power environment in a combined heat and power (CHP) system based on distributed reinforcement learning, such as... Figure 1 As shown, the coordinated control method includes the following steps: S1. Establish an active power environment scheduling model for a combined heat and power system; Considering the environmental effects of active power environmental dispatch, the optimal system operating state and the minimum carbon dioxide emission level are taken as the overall optimal operating state and serve as the optimization objective function for active power environmental dispatch. Constraints are given for system active power balance and heat supply balance constraints, as well as unit active power output and heat supply output limits.
[0019] S11. Establish the objective function model for active power environment scheduling; When the system includes thermal power units, combined heat and power units, and pure heating units, the active power environment dispatch model can be written in the following form: (1); in, For thermal power units Coal consumption, For thermal power units carbon emission coefficient, For combined heat and power units Coal consumption, For pure heating units Coal consumption, For the number of combined heat and power units, This refers to the number of units used solely for heating. This is a penalty factor.
[0020] S12. Determine the constraints of the objective function; S121. Regarding active power balance and heat supply balance constraints: (6); (7); in, For system transmission losses, This is the loss coefficient.
[0021] S122. Constraints are determined for the limits of active power output and heating output of the unit: (8); (9); For combined heat and power (CHP) units, there is a coupling relationship between their active power output and heating output, expressed as follows: (16); (17); S13. Solve for the marginal function of the active power environment scheduling function; The marginal function of the active power environment dispatch function for the thermal power generating unit is: (18); middle, For coefficients, For the unit The minimum active power output.
[0022] No. The marginal function expression of the active power environment dispatch function of the combined heat and power unit is: (8); (9); No. The marginal function expression of the active power environment dispatch function for a pure heating unit is: (19); Based on the marginal function of the active power environment dispatch function of thermal power generating units, the expression for the active power output regulation capacity of the units during the iterative process can be obtained as follows: (20); in, The expression (18) obtained and The relational function of the theory, For the number of iterations, The estimated power generation load regulation margin for the first generating unit. Let be the adjacency matrix of a directed graph.
[0023] Based on the marginal function of the active power environment dispatch function of thermal power generating units, the expression for the active power output regulation capacity of the units during the iterative process can be obtained as follows: (twenty one); Based on the marginal function of the active power environment dispatch function of thermal power generating units, the expression for the active power output regulation capacity of the units during the iterative process can be obtained as follows: (twenty two) S2. Determine the solution strategy based on distributed reinforcement learning; S21. Initialize relevant parameters; The generating output and marginal function of the generating units are initialized based on the actual data of the previous operating cycle, and the load is evenly distributed to the generating load adjustment margin of each generating unit according to the unit capacity. S22. Determine the unit performance indicators; The training process of the distributed reinforcement learning algorithm is synchronized with the system operation. A unified performance index is defined for each unit so that the proposed algorithm does not depend on the unit parameters. The algorithm includes a commentator neural network and an actor neural network. The commentator neural network is used to estimate local performance indexes to optimize the power generation output of each unit. The actor neural network is used to determine the marginal function of each unit and converge the marginal function to the optimal value, thereby reducing the difference between power generation and load and achieving the goal of optimizing active power environment scheduling.
[0024] Accordingly, the specific expression for the performance metric is: (10); in, The error of the actor neural network is calculated by simultaneously considering both reward and penalty terms, and its expression is: (11); S3, iteratively update performance metrics and neural network errors; The marginal function of each unit is calculated, and after the training rounds are completed, backpropagation is performed with the goal of maximizing overall performance to update performance indicators and neural network errors, optimize neural network parameters, and obtain optimal instruction parameters.
[0025] The output of the commentator neural network is an estimate of the performance index in expression (2), and the estimation error is shown in the following expression: (12); in, The estimated value of the performance index is expressed as follows: (13); in, Here is the weight matrix of the first layer of the commenter neural network. This is the weight matrix of the second-layer commenter neural network. Input to the commentator's neural network.
[0026] The weight matrix of the second-layer neural network is trained using the gradient descent algorithm based on the error value estimated by the performance index. The expression is as follows: (14); in, This represents the learning rate of the commentator's neural network.
[0027] The actor neural network is trained based on expression (11), and its output expression is: (15); in, As input to the actor neural network, This is the weight matrix of the first layer of the actor neural network, which can be set as a random matrix. This is the weight matrix of the second-layer actor neural network.
[0028] The weight matrix of the second-layer neural network is trained using the gradient descent algorithm, and its expression is: (16); in, represents the learning rate of the actor neural network.
[0029] The training rounds for performance metrics and neural network errors are repeated continuously until the cumulative reward value stabilizes, and the unit... marginal function of the active environment scheduling function Training stops when the optimal value is reached. Multiple units can be trained simultaneously to collect diverse data, which is then updated uniformly. Finally, the parameter values of the trained intelligent agent network are saved and can be directly retrieved when needed. For different operating scenarios, inputting real-world scenario data into the intelligent agent generates the optimal control strategy.
[0030] The following section uses a real system with 40 generator sets as an example to illustrate the active power environment coordination control method for cogeneration systems based on distributed reinforcement learning.
[0031] The system contains 40 generator units, with a system load demand of 10,500 MW. (Commentator's Neural Network) and actor neural network The learning rate was determined using a trial-and-error method, with the optimal value being [value to be filled in]. .
[0032] The methods used for active power environment scheduling generally include mathematical optimization methods and metaheuristic algorithms. However, metaheuristic algorithms cannot guarantee a globally optimal result, mathematical optimization methods require an accurate power grid model, and the computational efficiency of the algorithm decreases significantly as the power grid scales up. This invention uses distributed control, which can avoid the problem of optimization algorithm interruption due to the disconnection of one or more control signal paths caused by faults. It also has the advantages of better robustness and fewer inter-regional connections, thus improving computational efficiency.
[0033] The method of this invention is compared with the scheduling results of the harmony search algorithm, Kho algorithm, reinforcement learning algorithm and gravity search algorithm in a power grid as follows: Figure 2 As shown, G represents the active power output of the unit, and H represents the heating output of the unit.
[0034] The results of the method of this invention are compared with other optimization methods such as the flower pollination algorithm, Kho algorithm, squirrel search algorithm and linear programming method, as shown in Table 1.
[0035] Table 1. Comparison of results between the method proposed in this invention and other optimized methods. Flower pollination algorithm Kho algorithm Squirrel Search Algorithm Linear programming method The method proposed in this invention <![CDATA[Active environment scheduling function result (×10 5 )]]> 1.2317 1.2585 1.30 1.3012 1.2415 <![CDATA[Emission function result (×10 5 )]]> 2.0846 2.1084 1.767 1.7812 1.7228 The marginal function convergence characteristic curve of this invention applied to a real power grid system is shown in the figure below. Figure 3 As shown, the curve converges more and more as the number of iterations increases.
[0036] From Table 1, Figure 2 and Figure 3 It can be seen that by applying the solution strategy of the active power environment coordination control method for cogeneration systems based on distributed reinforcement learning of this invention, the active power environment scheduling function is similar to that of other methods, but the emission function is minimized and the system operating state is optimal. It can better guarantee the global optimality of the optimization result than other optimization methods, and improves the robustness and security of the active power environment scheduling method, thus proving the effectiveness of the method of this invention.
[0037] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A method for coordinated control of active power environment in a combined heat and power system based on distributed reinforcement learning, characterized in that, It includes the following steps: S1. Establish an active power environment scheduling model for a combined heat and power system; The expression for the active power environmental dispatch model of a combined heat and power system considering environmental effects is as follows: (1); in, For thermal power units Coal consumption, For thermal power units carbon emission coefficient, For combined heat and power units Coal consumption, For pure heating units Coal consumption, For the number of combined heat and power units, This refers to the number of units used solely for heating. This refers to the number of thermal power units. For the unit Those who have made contributions For the unit The heat emitted, As a penalty factor; The constraints are given based on the system power balance and the unit output limits, respectively, to obtain the first... The marginal function of the active power environment dispatch function of a thermal power generating unit, a combined heat and power unit, and a pure heating unit; S2. Determine the solution strategy based on distributed reinforcement learning; S21. Initialize relevant parameters; The generating output and marginal function of the units are initialized based on actual data, and the load is evenly distributed to the generating load adjustment margin of each unit according to the unit capacity. S22. Determine the unit performance indicators; Local performance indicators are used to optimize the power generation output of each unit, and their expressions are as follows: (2); in, For the unit marginal function, For the unit The power generation load regulation margin The time interval for updating the unit's performance indicators. For the number of iterations, For the unit Performance metrics; S3, iteratively update performance metrics and neural network errors; After the training rounds are completed, backpropagation is performed with the goal of maximizing overall performance. This updates the performance metrics and neural network error, optimizes the neural network parameters, and obtains the optimal instruction parameters.
2. The active power environment coordination control method for a cogeneration system based on distributed reinforcement learning according to claim 1, characterized in that: Step S1 specifically includes the following steps: S11. Determine the objective function for active power environment scheduling; When considering environmental effects, the objective function expression for active power environmental scheduling of cogeneration units is: (3); in, For coefficients; S12. Determine the constraints of the objective function; S13. Solve for the marginal function of the active power environment scheduling function; No. The marginal function expression of the active power environment dispatch function of the combined heat and power unit is: (4); (5)。 3. The active power environment coordination control method for a cogeneration system based on distributed reinforcement learning according to claim 2, characterized in that: Step S12 specifically includes the following steps: S121. Regarding active power balance and heat supply balance constraints: (6); (7); in, To meet the system's active power load requirements, To meet the system's heat load requirements, This refers to system transmission losses; S122. Constraints are determined for the limits of active power output and heating output of the unit: (8); (9); in, For the unit The upper limit of contribution. For the unit The lower limit of meritorious service For the unit The lower limit of heating output For the unit The upper limit of contribution.
4. The active power environment coordination control method for a cogeneration system based on distributed reinforcement learning according to claim 1, characterized in that: The performance metrics in expression (2) of step S22 are as follows: (10); in, For the unit The marginal function is the marginal function of the unit. The error of the actor neural network is expressed as: (11)。 5. The active power environment coordination control method for a cogeneration system based on distributed reinforcement learning according to claim 1, characterized in that: In step S3, during the iterative update of performance metrics and neural network error, the output of the commentator neural network is the estimated value of the performance metrics in expression (2), and the estimation error is: (12); in, For the unit The estimated value of the performance index is expressed as follows: (13); in, Here is the weight matrix of the first layer of the commenter neural network. This is the weight matrix of the second-layer commenter neural network. Input to the commentator's neural network, For the commentator's neural network function.
6. The active power environment coordination control method for a cogeneration system based on distributed reinforcement learning according to claim 5, characterized in that: The weight matrix of the second-layer neural network is trained using the gradient descent algorithm based on the error value estimated by the performance index. The expression is as follows: (14); in, This represents the learning rate of the commentator's neural network.
7. The active power environment coordination control method for a cogeneration system based on distributed reinforcement learning according to claim 5, characterized in that: The actor neural network is trained based on expression (11), and its output expression is: (15); in, As input to the actor neural network, This is the weight matrix of the first layer of the actor neural network, which can be set as a random matrix. This is the weight matrix of the second-layer actor neural network. For the actor neural network function.
8. The active power environment coordination control method for a cogeneration system based on distributed reinforcement learning according to claim 1, characterized in that: In step S3, at the first In the next iteration, the unit marginal function of the active environment scheduling function The iteration stops when the optimal value is reached.