Thermal power plant carbon dioxide emission optimization control system and method based on intelligent feedback

By combining deep learning algorithms and intelligent feedback mechanisms, an intelligent control framework is constructed, which solves the shortcomings of carbon dioxide emission monitoring and control in thermal power plants, realizes accurate monitoring and optimized control of carbon dioxide emissions, improves the system's flexibility and responsiveness, and helps thermal power plants achieve low-carbon transformation.

CN120848415APending Publication Date: 2025-10-28HUANENG POWER INT INC JINGGANGSHAN POWER PLANT +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511003397.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing carbon dioxide emission monitoring and control systems for thermal power plants lack intelligent data analysis capabilities, cannot adapt to changes in power load and environment in real time, and have insufficient feedback mechanisms, resulting in sluggish response and low control efficiency in dynamic environments.

Method used

By employing deep learning algorithms and intelligent feedback mechanisms, an intelligent and dynamically adjustable control framework is constructed. Through multi-dimensional data integration and adaptive feedback mechanisms, real-time monitoring and optimized control of carbon dioxide emissions are achieved, and decision optimization is performed using deep neural networks and policy networks.

Benefits of technology

It improves the accuracy of carbon dioxide emission trend prediction and system flexibility, enhances the ability to respond to emergencies, ensures stable operation and low carbon emissions during power generation, and improves power generation efficiency and overall operational stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120848415A_ABST
    Figure CN120848415A_ABST
Patent Text Reader

Abstract

The invention discloses a thermal power plant carbon dioxide emission optimization control system and method based on intelligent feedback. The system comprises a data acquisition module, a state representation module, a control execution module and an environment feedback module. The data acquisition module acquires multi-dimensional data in real time, the multi-dimensional data comprise fuel components, boiler states, environmental parameters and power generation loads, and outputs standardized multi-dimensional input data to the state representation module; the state representation module receives standardized input data from the data acquisition module and outputs a generated comprehensive state vector for calculation of the decision module; the control execution module receives the state vector from the state representation module and outputs a control command needing to be executed; and the environment feedback module receives the control command generated by the control execution module and the executed system state change, and outputs updated environment feedback data. By improving the resource utilization efficiency and optimizing carbon dioxide emission control, a feasible solution is provided for the thermal power plant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a carbon dioxide emission optimization control system for thermal power plants, specifically to a carbon dioxide emission optimization control system and method for thermal power plants based on intelligent feedback. Background Art

[0002] With increasing global focus on sustainable development and environmental protection, the monitoring and control of carbon dioxide emissions from thermal power plants has become a pressing technological challenge. However, existing monitoring and control systems still face numerous shortcomings in data processing, decision support, and feedback mechanisms. Many current systems rely primarily on traditional physical models and experience-based scheduling methods, resulting in poor adaptability to dynamic environments and difficulty in responding to real-time changes in power load and emission levels. Existing technologies typically employ static monitoring methods, lacking intelligent data analysis and predictive capabilities, and failing to effectively integrate multi-dimensional input data, thus affecting the accurate assessment and control of carbon dioxide emissions. Although some systems attempt to introduce machine learning and optimization algorithms to improve the intelligence of monitoring and control, they often fail to fully consider complex multivariate relationships, leading to a lack of flexibility and real-time performance in the decision-making process. Furthermore, traditional systems also have deficiencies in reward mechanisms and feedback control, lacking adaptive capabilities based on agent-based decision-making. The control logic often cannot dynamically adjust according to real-time data and environmental changes, thus affecting the system's performance under different operating conditions. This presents thermal power plants with significant challenges in reducing carbon dioxide emissions and improving power generation efficiency.

[0003] Currently, thermal power plants face numerous challenges in monitoring and controlling carbon dioxide emissions. First, existing monitoring systems often rely on traditional physical models and empirical rules, lacking intelligent data analysis and predictive capabilities. This results in a lag in the system's response to carbon dioxide emissions in dynamic environments, making it unable to adapt to changes in power load and the environment in real time. Second, existing technologies have limitations in processing multi-dimensional data, failing to effectively integrate various inputs such as fuel composition, boiler status, environmental parameters, and power generation load. This limits the system's ability to accurately assess and control carbon dioxide emissions, making it difficult for thermal power plants to reduce emissions and improve efficiency. Furthermore, the feedback mechanisms and reward calculation logic of traditional systems are insufficient, making it difficult to provide real-time decision support for intelligent agents, resulting in insufficient adaptability in complex environments. Existing technologies typically fail to achieve intelligent utilization of real-time data, affecting the overall efficiency and reliability of the system.

[0004] Therefore, there is an urgent need for a carbon dioxide emission optimization control system for thermal power plants based on an intelligent feedback mechanism, in order to achieve more efficient and flexible energy management and emission control, and promote the transformation of thermal power plants towards low-carbon and intelligent operation. Summary of the Invention

[0005] Addressing the shortcomings of existing technologies, this invention aims to propose a carbon dioxide emission optimization control system and method for thermal power plants based on intelligent feedback. It comprehensively utilizes deep learning algorithms and intelligent feedback mechanisms to construct an intelligent, dynamically adjustable control framework. The system proposed in this invention can effectively process data inputs from multiple dimensions, including fuel composition, boiler status, environmental parameters (such as temperature and humidity), and power generation load. This multi-dimensional data integration enhances the system's real-time monitoring capability for carbon dioxide emissions, making its analysis of emission trends more accurate. By establishing a complex deep learning model, the system can capture nonlinear relationships in the data, thereby achieving high-precision prediction of carbon dioxide emissions. In terms of optimized control, this invention employs an adaptive feedback mechanism, enabling the agent to adjust its decisions in real time based on the current state and environmental feedback. By calculating reward values ​​and updating Q-values ​​in real time, the agent can gradually learn the optimal control strategy to achieve the best power generation efficiency and the lowest carbon dioxide emissions. This implementation of intelligent feedback not only improves the system's flexibility but also enhances its response to emergencies, ensuring that thermal power plants can maintain stable operation in a volatile environment.

[0006] The system proposed in this invention aims to support the goal of low-carbon transformation of thermal power plants. By improving resource utilization efficiency and optimizing carbon dioxide emission control, it provides a practical solution for thermal power plants.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] The intelligent feedback-based carbon dioxide emission optimization control system for thermal power plants includes:

[0009] The data acquisition module is used to collect multi-dimensional data in real time, including fuel composition, boiler status, environmental parameters and power generation load, and outputs standardized multi-dimensional input data to the status representation module.

[0010] The state representation module receives standardized input data from the data acquisition module and outputs a generated comprehensive state vector for calculation by the decision-making module.

[0011] The control execution module is used to receive the state vector from the state representation module and output the control commands to be executed.

[0012] The environment feedback module is used to receive control commands generated by the control execution module and the changes in system state after their execution, and output updated environment feedback data.

[0013] A further improvement of this invention lies in that the data acquisition module outputs standardized multi-dimensional input data and transmits it to the state representation module, wherein each type of data undergoes standardization processing, using the following formula:

[0014]

[0015] Among them, X n This is the raw data, μ x This represents the mean of the corresponding data, σ. x The standard deviation is used to weight the standardized data and generate a comprehensive input vector. The calculation formula is as follows:

[0016] X w =w F *F t +w B *B t +w E *E t +w L *L t (2)

[0017] Among them, w F w B w E w L Weighting coefficients for fuel data, boiler data, environmental data, and load data.

[0018] A further improvement of this invention is that, in the state representation module, the state representation is generated by a deep learning model or feature extraction method to generate a comprehensive state representation, using a deep neural network to generate the state representation:

[0019] S t =DNN(X) w )=f(w*X w +b) (3)

[0020] Where f is the activation function, using the ReLU function to improve the expressive power of nonlinear features; w is the weight term of the DNN unit; and b is the bias term.

[0021] A further improvement of the present invention is that the control execution module receives a state vector S from the state representation module. t The output system needs to execute the control command A. t This includes boiler combustion rate and fuel input. A policy network is used to select the optimal action, as shown in the formula:

[0022] A t =PN(S) t (4)

[0023] Here, PN is the policy network, and the policy network is a deep DQN network, which improves training stability through experience replay and target network.

[0024] A further improvement of this invention is that, in the environmental feedback module, the updated environmental feedback data output includes the carbon dioxide emission level C. new and power generation efficiency E new Environmental feedback is used to calculate the system's reward value, which is measured by reducing carbon dioxide emissions and improving power generation efficiency. The reward formula is as follows:

[0025] R t =α*(E new -E old )-β*(C new -C old (5)

[0026] Where α and β are the weights set to ensure the optimization objective is balanced; the new state S t+1 Updated based on environmental feedback, the formula is:

[0027] S t+1 =f e (S t A t (6)

[0028] Among them, S t A represents the system state at the current moment. t f represents the action taken by the agent at any given time. e This represents the environment state transition function.

[0029] Intelligent feedback-based methods for optimizing carbon dioxide emissions control in thermal power plants include:

[0030] The data acquisition module collects multi-dimensional data in real time, including fuel composition, boiler status, environmental parameters and power generation load, and outputs standardized multi-dimensional input data to the status representation module.

[0031] The state representation module receives standardized input data from the data acquisition module and outputs a generated comprehensive state vector for calculation by the decision module.

[0032] The control execution module receives the state vector from the state representation module and outputs the control commands to be executed.

[0033] The environmental feedback module receives control commands generated by the control execution module and the changes in system state after their execution, and outputs updated environmental feedback data.

[0034] A further improvement of this invention lies in that the data acquisition module outputs standardized multi-dimensional input data and transmits it to the state representation module, wherein each type of data undergoes standardization processing, using the following formula:

[0035]

[0036] Among them, X n This is the raw data, μ x This represents the mean of the corresponding data, σ. x The standard deviation is used to weight the standardized data and generate a comprehensive input vector. The calculation formula is as follows:

[0037] X w =w F *F t +w B *B t +w E *E t +w L *L t (2)

[0038] Among them, w F w B w E w L Weighting coefficients for fuel data, boiler data, environmental data, and load data.

[0039] A further improvement of this invention is that the state representation module generates a comprehensive state representation through a deep learning model or feature extraction method, using a deep neural network to generate the state representation:

[0040] S t =DNN(X) w )=f(w*X w +b) (3)

[0041] Where f is the activation function, using the ReLU function to improve the expressive power of nonlinear features; w is the weight term of the DNN unit; and b is the bias term.

[0042] A further improvement of the present invention is that the control execution module receives a state vector S from the state representation module. t The output system needs to execute the control command A. t This includes boiler combustion rate and fuel input. A policy network is used to select the optimal action, as shown in the formula:

[0043] A t =PN(S) t (4)

[0044] Here, PN is the policy network, and the policy network is a deep DQN network, which improves training stability through experience replay and target network.

[0045] A further improvement of this invention is that the environmental feedback module outputs updated environmental feedback data including carbon dioxide emission level C. new and power generation efficiency E newEnvironmental feedback is used to calculate the system's reward value, which is measured by reducing carbon dioxide emissions and improving power generation efficiency. The reward formula is as follows:

[0046] R t =α*(E new -E old )-β*(C new -C old (5)

[0047] Where α and β are the weights set to ensure the optimization objective is balanced; the new state S t+1 Updated based on environmental feedback, the formula is:

[0048] S t+1 =f e (S t A t (6)

[0049] Among them, S t A represents the system state at the current moment. t f represents the action taken by the agent at any given time. e This represents the environment state transition function.

[0050] Compared with the prior art, the present invention has at least the following beneficial technical effects:

[0051] This invention proposes an intelligent feedback-based optimized control system and method for carbon dioxide emissions in thermal power plants, demonstrating significant technical advantages and application effects. By combining deep learning algorithms and intelligent feedback mechanisms, the system can efficiently process multi-dimensional input data, achieving precise monitoring and optimized control of carbon dioxide emissions. The proposed optimization system significantly improves the ability to predict carbon dioxide emission trends through in-depth analysis of multi-dimensional data such as fuel composition, boiler status, environmental parameters, and power generation load. The introduction of deep learning models effectively captures complex nonlinear relationships in the data, thereby improving prediction accuracy and ensuring more scientific and reasonable emission control during power generation. Through an adaptive feedback mechanism, the agent can receive environmental feedback in real time after executing control commands, quickly updating decision-making strategies, ensuring the system can respond rapidly in dynamic environments, optimizing power generation efficiency and reducing carbon dioxide emissions. For example, when the grid load changes, the agent can adjust the combustion rate and fuel input in a timely manner based on the latest status and reward calculations, minimizing emissions and improving energy efficiency. The optimized control system of this invention has a high level of intelligence, enabling flexible scheduling in complex thermal power plant operating environments, reducing the need for human intervention, and improving the overall operational stability and safety. Through real-time monitoring and adjustments, the system can ensure that carbon dioxide emissions during power generation remain within predetermined environmental standards, thus contributing to the low-carbon transformation of thermal power plants.

[0052] This invention not only effectively solves the shortcomings of existing technologies in carbon dioxide emission monitoring and optimization control, but also significantly improves the environmental protection capabilities and power generation efficiency of thermal power plants through intelligent feedback design. It has good market application prospects and promotion value, and provides strong support for achieving a sustainable energy production model. Attached Figure Description

[0053] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0054] Figure 1 This is a system flowchart for an intelligent feedback-based optimized control system for carbon dioxide emissions in thermal power plants.

[0055] Figure 2 This diagram illustrates the core intelligent agent decision-making mechanism of a thermal power plant carbon dioxide emission optimization control system based on intelligent feedback. Detailed Implementation

[0056] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of the invention. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.

[0057] In the description of this invention,

[0058] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0059] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0060] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0061] The accompanying drawings illustrate various structural schematic diagrams according to embodiments disclosed in this invention. These drawings are not to scale, and some details have been enlarged for clarity, and some details may have been omitted. The shapes of the various regions and layers shown in the drawings, as well as their relative sizes and positional relationships, are merely exemplary and may deviate from reality due to manufacturing tolerances or technical limitations. Furthermore, those skilled in the art can design regions / layers with different shapes, sizes, and relative positions as needed.

[0062] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0063] Example 1

[0064] Please see Figure 1 This invention provides a carbon dioxide emission optimization control system for thermal power plants based on intelligent feedback. It mainly consists of a data acquisition module, a status representation module, a control execution module, and an environmental feedback module. The data acquisition module receives real-time multi-dimensional data, including fuel composition Ft, boiler status Bt, environmental parameters Et (temperature and humidity), and power generation load Lt. It then outputs standardized multi-dimensional input data to the status representation module, which performs standardization processing on each type of data using the following formula:

[0065]

[0066] Where Xn is the original data, μx is the mean of the corresponding data, and σx is the standard deviation; the standardized data are weighted to generate a comprehensive input vector, and the calculation formula is as follows:

[0067] X w =w F *F t +w B *B t +w E *E t +w L *L t (2)

[0068] Among them, w F w B w E w L The weighting coefficients are used for fuel data, boiler data, environmental data, and load data. The weights can be adjusted based on the actual importance of the data.

[0069] The status representation module receives standardized input data X from the data acquisition module. w Output the generated comprehensive state vector S t This is used for computation in the decision-making module. State representations can be generated using deep learning models or feature extraction methods. This system uses deep neural networks (DNNs) to generate state representations.

[0070] S t =DNN(X) w )=f(w*X w +b) (3)

[0071] Where f is the activation function, using the ReLU function (Linear Rectified Unit) to improve the expressive power of nonlinear features; w is the weight term of the DNN unit, and b is the bias term.

[0072] The control execution module receives the state vector S from the state representation module. t The output system needs to execute the control command A. t This includes factors such as boiler combustion rate and fuel input. A strategy network is used to select the optimal action, using the following formula:

[0073] A t =PN(S) t (4)

[0074] Here, PN stands for Policy Network. The policy network in this system is a DQN deep network, which improves training stability through experience replay and target network.

[0075] The input to the environmental feedback module is the control commands generated by the control execution module and the changes in the system state after their execution. t The output is updated environmental feedback data, including carbon dioxide emission level C. new and power generation efficiency E new Environmental feedback is used to calculate the system's reward value, measured by reducing carbon dioxide emissions and improving power generation efficiency. The reward formula is:

[0076] R t =α*(E new -E old )-β*(C new -C old (5)

[0077] Where α and β are the weights set to ensure the optimization objective is balanced; the new state S t+1 Updated based on environmental feedback, the formula is:

[0078] S t+1 =f e (S t At (6)

[0079] Among them, S t A represents the system state at the current moment. t f represents the action taken by the agent at any given time. e The system represents the environmental state transition function. The agent learns the optimal strategy through interaction with the environment to control key parameters such as carbon dioxide emissions. This system achieves closed-loop control through the four modules mentioned above. Combining data acquisition, state representation, control execution, and environmental feedback, it utilizes a deep reinforcement learning policy network for optimization, thereby effectively controlling carbon dioxide emissions from thermal power plants.

[0080] Please see Figure 2 This invention relates to an intelligent feedback-based carbon dioxide emission optimization control system for thermal power plants. Its core intelligent agent decision-making mechanism also includes reward calculation and Q-value updating, belonging to the system environment feedback module and control execution module. Reward calculation, belonging to the control execution module, is performed by the intelligent agent through the execution of a certain action A. t The feedback evaluation of the system performance is used to quantify the effect of the agent's current decisions. This helps the agent understand whether its actions contribute to improving power generation efficiency and reducing carbon dioxide emissions. The reward value is passed back to the agent for subsequent Q-value updates, guiding the agent to choose better actions. Q-value updates are based on the new state and immediate reward, and are performed through a deep DQN (Q-network). The feedback mechanism, by providing new states and rewards, gradually optimizes the Q-value of the control execution module, thereby influencing future action choices. The calculation formula is as follows:

[0081] Q(S t A t ) = r t +γ*max(S t+1 A t+1 (7)

[0082] r t Q(S) represents the immediate reward at the current moment, calculated by the environmental feedback module based on the reduction of carbon dioxide and the improvement of power generation efficiency; γ is a discount factor used to balance the current reward and future rewards; t+1 A t+1 The maximum Q value is the value at the next moment. Based on the calculated Q value, the control execution module selects action A with the maximum Q value. t The policy network determines the optimal action for the current moment, and the Q-value update affects the agent's state S at the next time step. t+1The action selection process is crucial. Reward calculation and Q-value updates together constitute the core decision-making mechanism of the agent. Through close collaboration between the control execution module and the environmental feedback module, the system can optimize its decision-making strategy based on immediate rewards after each action. Reward calculation provides the agent with immediate feedback, helping to evaluate the effectiveness of the current action, while Q-value updates are continuously adjusted through a deep Q-network, ensuring that the agent gradually learns the optimal decision-making method. Ultimately, this closed-loop feedback mechanism enables the system to maximize power generation efficiency and minimize carbon dioxide emissions during continuous optimization, thereby maximizing long-term benefits.

[0083] Example 2

[0084] The present invention provides a method for optimizing and controlling carbon dioxide emissions from thermal power plants based on intelligent feedback, comprising:

[0085] The data acquisition module collects multi-dimensional data in real time, including fuel composition, boiler status, environmental parameters and power generation load, and outputs standardized multi-dimensional input data to the status representation module.

[0086] The state representation module receives standardized input data from the data acquisition module and outputs a generated comprehensive state vector for calculation by the decision module.

[0087] The control execution module receives the state vector from the state representation module and outputs the control commands to be executed.

[0088] The environmental feedback module receives control commands generated by the control execution module and the changes in system state after their execution, and outputs updated environmental feedback data.

[0089] In this embodiment, the data acquisition module outputs standardized multi-dimensional input data to the state representation module. Each type of data undergoes standardization processing, using the following formula:

[0090]

[0091] Among them, X n This is the raw data, μ x This represents the mean of the corresponding data, σ. x The standard deviation is used to weight the standardized data and generate a comprehensive input vector. The calculation formula is as follows:

[0092] X w =w F *F t +w B *B t +w E *E t +w L *L t (2)

[0093] Among them, w F w B w E w L Weighting coefficients for fuel data, boiler data, environmental data, and load data.

[0094] In this embodiment, the state representation module generates a comprehensive state representation through a deep learning model or feature extraction method, using a deep neural network to generate the state representation:

[0095] S t =DNN(X) w )=f(w*X w +b) (3)

[0096] Where f is the activation function, using the ReLU function to improve the expressive power of nonlinear features; w is the weight term of the DNN unit; and b is the bias term.

[0097] In this embodiment, the control execution module receives the state vector S from the state representation module. t The output system needs to execute the control command A. t This includes boiler combustion rate and fuel input. A policy network is used to select the optimal action, as shown in the formula:

[0098] A t =PN(S) t (4)

[0099] Here, PN is the policy network, and the policy network is a deep DQN network, which improves training stability through experience replay and target network.

[0100] In this embodiment, the updated environmental feedback data output by the environmental feedback module includes the carbon dioxide emission level C. new and power generation efficiency E new Environmental feedback is used to calculate the system's reward value, which is measured by reducing carbon dioxide emissions and improving power generation efficiency. The reward formula is as follows:

[0101] R t =α*)E new -E old )-β*(C new -C old (5)

[0102] Where α and β are the weights set to ensure the optimization objective is balanced; the new state S t+1 Updated based on environmental feedback, the formula is:

[0103] S t+1 =f e(S t A t (6)

[0104] Among them, S t A represents the system state at the current moment. t f represents the action taken by the agent at any given time. e This represents the environment state transition function.

[0105] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the scope of the invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

[0106] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can be appropriately combined to form other embodiments that can be understood by those skilled in the art. The above content is only for illustrating the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made based on the technical concept proposed in this invention shall fall within the scope of protection of the claims of this invention.

Claims

1. A carbon dioxide emission optimization control system for thermal power plants based on intelligent feedback, characterized in that, include: The data acquisition module is used to collect multi-dimensional data in real time, including fuel composition, boiler status, environmental parameters and power generation load, and outputs standardized multi-dimensional input data to the status representation module. The state representation module receives standardized input data from the data acquisition module and outputs a generated comprehensive state vector for calculation by the decision-making module. The control execution module is used to receive the state vector from the state representation module and output the control commands to be executed. The environment feedback module is used to receive control commands generated by the control execution module and the changes in system state after their execution, and output updated environment feedback data.

2. The intelligent feedback-based carbon dioxide emission optimization control system for thermal power plants according to claim 1, characterized in that, The data acquisition module outputs standardized, multi-dimensional input data, which is then transmitted to the state representation module. Each data type undergoes standardization processing using the following formula: Among them, X n This is the raw data, μ x This represents the mean of the corresponding data, σ. x The standard deviation is used to weight the standardized data and generate a comprehensive input vector. The calculation formula is as follows: X w =w F *F t +w B *B t +w E *E t +w L *L t (2) Among them, w F w B w E w L Weighting coefficients for fuel data, boiler data, environmental data, and load data.

3. The intelligent feedback-based carbon dioxide emission optimization control system for thermal power plants according to claim 1, characterized in that, In the state representation module, a comprehensive state representation is generated through a deep learning model or feature extraction method. Specifically, a deep neural network is used to generate the state representation. S t =DNN(X w )=f(w*X w +b) (3) Where f is the activation function, using the ReLU function to improve the expressive power of nonlinear features; w is the weight term of the DNN unit; and b is the bias term.

4. The intelligent feedback-based carbon dioxide emission optimization control system for thermal power plants according to claim 3, characterized in that, In the control execution module, the state vector S is received from the state representation module. t The output system needs to execute the control command A. t This includes boiler combustion rate and fuel input. A policy network is used to select the optimal action, as shown in the formula: A t =PN(S t )(4) Here, PN is the policy network, and the policy network is a deep DQN network, which improves training stability through experience replay and target network.

5. The intelligent feedback-based carbon dioxide emission optimization control system for thermal power plants according to claim 4, characterized in that, The environmental feedback module outputs updated environmental feedback data, including carbon dioxide emission levels C. new and power generation efficiency E new Environmental feedback is used to calculate the system's reward value, which is measured by reducing carbon dioxide emissions and improving power generation efficiency. The reward formula is as follows: R t =α*(E new -E old )-β*(C new -C old ) (5) Where α and β are the weights set to ensure the optimization objective is balanced; the new state S t+1 Updated based on environmental feedback, the formula is: S t+1 =f e (S t ,A t )(6) Among them, S t A represents the system state at the current moment. t f represents the action taken by the agent at any given time. e This represents the environment state transition function.

6. A method for optimizing carbon dioxide emission control in thermal power plants based on intelligent feedback, characterized in that, include: The data acquisition module collects multi-dimensional data in real time, including fuel composition, boiler status, environmental parameters and power generation load, and outputs standardized multi-dimensional input data to the status representation module. The state representation module receives standardized input data from the data acquisition module and outputs a generated comprehensive state vector for calculation by the decision module. The control execution module receives the state vector from the state representation module and outputs the control commands to be executed. The environmental feedback module receives control commands generated by the control execution module and the changes in system state after their execution, and outputs updated environmental feedback data.

7. The method for optimizing carbon dioxide emission control in thermal power plants based on intelligent feedback according to claim 6, characterized in that, The data acquisition module outputs standardized, multi-dimensional input data, which is then transmitted to the state representation module. Each data type undergoes standardization processing using the following formula: Among them, X n This is the raw data, μ x This represents the mean of the corresponding data, σ. x The standard deviation is used to weight the standardized data and generate a comprehensive input vector. The calculation formula is as follows: X w =w F *F t +w B *B t +w E *E t +w L *L t (2) Among them, w F w B w E w L Weighting coefficients for fuel data, boiler data, environmental data, and load data.

8. The method for optimizing carbon dioxide emission control in thermal power plants based on intelligent feedback according to claim 6, characterized in that, The state representation module generates a comprehensive state representation through deep learning models or feature extraction methods, using deep neural networks to generate the state representation: S t =DNN(X w )=f(w*X w +b) (3) Where f is the activation function, using the ReLU function to improve the expressive power of nonlinear features; w is the weight term of the DNN unit; and b is the bias term.

9. The method for optimizing carbon dioxide emissions from thermal power plants based on intelligent feedback according to claim 8, characterized in that, The control execution module receives the state vector S from the state representation module. t The output system needs to execute the control command A. t This includes boiler combustion rate and fuel input. A policy network is used to select the optimal action, as shown in the formula: A t =PN(S t ) (4) Here, PN is the policy network, and the policy network is a deep DQN network, which improves training stability through experience replay and target network.

10. The method for optimizing carbon dioxide emission control in thermal power plants based on intelligent feedback according to claim 9, characterized in that, The environmental feedback module outputs updated environmental feedback data, including carbon dioxide emission levels (C). new and power generation efficiency E new Environmental feedback is used to calculate the system's reward value, which is measured by reducing carbon dioxide emissions and improving power generation efficiency. The reward formula is as follows: R t =α*(E new -E old )-β*(C new -C old ) (5) Where α and β are the weights set to ensure the optimization objective is balanced; the new state S t+1 Updated based on environmental feedback, the formula is: S t+1 =f e (S t ,A t ) (6) Among them, S t A represents the system state at the current moment. t f represents the action taken by the agent at any given time. e This represents the environment state transition function.

Citation Information

Patent Citations

  • Optimization control method and control system for flue gas recirculation and boiler coupling system

    CN118818959A

  • Virtual power plant peak regulation optimization scheduling method and system, electronic equipment and medium

    CN119783997A

  • Intelligent scheduling and optimizing method for thermal power plant

    CN119886608A

  • Energy-saving control system and method for boiler in thermal power plant

    US12353177B1