A method for optimizing the combustion efficiency of a heating stove
By introducing the Q-Learning algorithm with dynamic exploration tendency coefficient into the heating boiler, the problems of slow learning and control disturbance in traditional algorithms under load changes are solved, achieving efficient and stable combustion optimization control and improving the system's energy efficiency and safety.
Patent Information
- Application Number
- CN202511624217.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-11-07
AI Technical Summary
Traditional Q-Learning algorithms in heating boilers cannot quickly adapt to load changes due to a fixed exploration rate, resulting in slow learning or unnecessary disturbances in control parameters, which affects the stability and efficiency of the combustion process.
A dynamic exploration tendency coefficient is introduced, and the exploration behavior of the Q-Learning algorithm is dynamically adjusted through adaptive learning urgency, performance deviation and safety constraint factors. Combined with real-time load changes and operating status, fuel supply and combustion air volume are optimized.
It achieves high combustion efficiency quickly and stably under various operating conditions, improving system energy efficiency, adaptability and safety, and avoiding unnecessary control disturbances.
Smart Images

Figure CN121067349B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of control optimization, and specifically to a method for optimizing the combustion efficiency control of a heating boiler. Background Technology
[0002] As an important heating device, the combustion efficiency of a boiler directly affects energy consumption and pollutant emissions. By monitoring the combustion status in real time, such as flue gas temperature and oxygen content, and dynamically adjusting the fuel supply and combustion air volume, the boiler can always operate at a highly efficient and clean operating point, resulting in significant economic and environmental benefits.
[0003] To achieve the above goals, reinforcement learning algorithms, especially Q-Learning algorithms, have been attempted to be applied to combustion optimization control. This algorithm, through continuous trial and error and learning in the state space, can autonomously find the mapping strategy from the current combustion state to the optimal control action, such as adjusting the damper and fuel valve. However, traditional Q-Learning algorithms have significant drawbacks. The core issue is that traditional Q-Learning algorithms typically use a fixed exploration rate or a simple preset decay strategy to balance the relationship between executing the currently known optimal action to obtain rewards. This static exploration mechanism is difficult to adapt to the dynamic and variable operating environment of the heating boiler. On the one hand, when the user's heat load fluctuates drastically, a fixed low exploration rate leads to slow algorithm learning, making it unable to quickly adapt to new operating conditions and resulting in low optimization efficiency. On the other hand, when the load is stable and the system is already in a relatively optimal state, a fixed high exploration rate will cause unnecessary disturbances to control parameters, affecting the stability of the combustion process and actual operating efficiency. Therefore, how to dynamically adjust the algorithm's exploration behavior according to the real-time operating status of the heating boiler and changes in external load is a technical problem that urgently needs to be solved by existing technologies. Summary of the Invention
[0004] To address the problem that a fixed exploration rate cannot quickly adapt to new operating conditions or cause unnecessary disturbances to control parameters, this invention proposes a method for optimizing the combustion efficiency control of a heating boiler. The method includes: acquiring the current operating parameters of the heating boiler, including flue gas oxygen content, exhaust gas temperature, and current heating load; applying a dynamic exploration tendency coefficient as the exploration rate to a Q-Learning control algorithm to adjust fuel supply and combustion air; in each control cycle, selecting to execute an exploration action or a utilization action based on the exploration rate and a generated random number, and updating the Q value; the dynamic exploration tendency coefficient is based on the base exploration rate and incorporates adaptive learning urgency, performance deviation, and safety constraints. The value of the sub-calculated value is positively correlated with the difference between the preset maximum exploration rate and the basic exploration rate, as well as the product of the latter three. The adaptive learning urgency is used to characterize the degree of change of the current heating load compared with the recent average heating load. The performance deviation is used to characterize the deviation of the current flue gas oxygen content and exhaust temperature from the optimal flue gas oxygen content and optimal exhaust temperature corresponding to the current heating load, respectively. The value of the safety constraint factor is close to 1 under the safe operating condition that the current flue gas oxygen content is higher than the set minimum safety threshold and the exhaust temperature is lower than the set maximum safety threshold, and rapidly approaches 0 when the current operating parameters approach any set safety threshold to suppress exploration behavior.
[0005] Compared to traditional Q-Learning algorithms that employ fixed control parameters or fixed exploration rates, this invention introduces a dynamic exploration tendency coefficient. This coefficient intelligently adjusts the balance between exploration and utilization in the control algorithm based on real-time load changes, operational performance deviations, and safety boundaries of the heating boiler. When system operating conditions change significantly or efficiency is low, it proactively enhances exploration to quickly find a better control strategy; when the system operates smoothly and efficiently, it reduces unnecessary exploration to maintain stability; and when approaching safety limits, it suppresses exploration behavior to ensure safety. This adaptive control method enables the heating boiler to achieve and maintain high combustion efficiency more quickly and stably under various operating conditions, thereby improving the overall system energy efficiency, adaptability, and safety.
[0006] Furthermore, the specific method for calculating the dynamic exploration tendency coefficient is as follows:
[0007] ;
[0008] in This represents the dynamic exploration tendency coefficient at the current moment; The base exploration rate is a small constant. This indicates the maximum preset exploration rate, which defines the upper limit of the exploration intensity; This represents the security constraint factor; This indicates the urgency of the adaptive learning; This indicates the performance deviation; This indicates the set exploration gain coefficient; It is a natural exponential function.
[0009] Furthermore, the safety constraint factor is the product of two functions: the value of the first function decreases from 1 to 0 as the difference between the current oxygen content of the flue gas and the minimum safety threshold decreases; the value of the second function decreases from 1 to 0 as the difference between the maximum safety threshold and the current exhaust temperature decreases.
[0010] By designing the safety constraint factor as a function related to the safety thresholds of flue gas oxygen content and exhaust temperature, the system's operational safety is significantly improved compared to control algorithms without built-in safety feedback mechanisms. When the control system is exploring and optimizing, it can proactively avoid control actions that might lead to dangerous conditions such as incomplete combustion or overheating. This is equivalent to adding a safety mechanism to the intelligent algorithm, ensuring that the optimization process of combustion efficiency always takes place within safe and reliable boundaries.
[0011] Further, the performance deviation is the L2 norm of the performance deviation vector; the performance deviation vector includes the following components: the first component is the deviation between the current flue gas oxygen content and the optimal flue gas oxygen content multiplied by a first weight; the second component is the deviation between the current exhaust temperature and the optimal exhaust temperature multiplied by a second weight; the sum of the first weight and the second weight is 1.
[0012] Compared to control methods that only consider a single indicator, this approach can more comprehensively and accurately assess the gap between the current combustion state and the optimal state. This integrated evaluation method makes the control system more sensitive and complete in its perception of performance degradation, thereby enabling it to more precisely drive the Q-Learning algorithm to adjust towards the direction of optimal overall performance, avoiding the problems that may arise from single-objective optimization.
[0013] Furthermore, the method for calculating the adaptive learning urgency is as follows:
[0014] ;
[0015] in This indicates the urgency of adaptive learning at the current moment; This indicates the set gain coefficient; Indicates the current heating load; This represents the average heating load in the recent period; It is a very small positive number, used to prevent the recent average load from being used as a reference. A division by zero error occurs when the result is zero.
[0016] By introducing adaptive learning urgency to assess the degree of change in heating load, the control system gains the ability to predict and quickly respond to changes in external demand. When user demand changes significantly, the system identifies it as an urgent situation and intensifies its exploration efforts, thereby quickly adapting to the new load point and finding the most efficient combustion mode, significantly improving the system's response speed and operating efficiency under dynamic conditions.
[0017] Furthermore, the method for obtaining the current heating load specifically involves: measuring the mass flow rate of circulating water using a flow meter; measuring the supply water temperature and return water temperature using a temperature sensor and calculating the supply and return water temperature difference; and using the product of the circulating water mass flow rate, the specific heat capacity of water, and the supply and return water temperature difference as the current heating load.
[0018] The current heating load is obtained by measuring the circulating water flow rate and the temperature difference between the supply and return water. Compared with the method of estimating or indirectly inferring the load, this direct measurement method provides real-time and accurate input data for the upper-level control algorithm, ensuring the reliability of the adaptive learning urgency calculation, which is the foundation for the effective implementation of the entire adaptive control strategy.
[0019] Furthermore, the recent average heating load It is calculated using the exponential moving average method.
[0020] Furthermore, the optimal flue gas oxygen content and optimal flue gas temperature are obtained by looking up a pre-built lookup table; the lookup table stores the mapping relationship between multiple discrete heating load points and the corresponding optimal flue gas oxygen content and optimal flue gas temperature.
[0021] Furthermore, the method for constructing the lookup table is as follows: under multiple different heating load conditions, the fuel supply and combustion air volume of the heating boiler are manually adjusted to determine the flue gas oxygen content and exhaust temperature at which the highest combustion efficiency is achieved under the current load, and these are recorded as the preset optimal values for the corresponding load points.
[0022] Furthermore, the reward function used by the Q-Learning control algorithm is set based on at least one of the real-time combustion efficiency, energy consumption, and pollutant emissions of the heating boiler.
[0023] The technical effects of this invention are as follows:
[0024] This invention proposes a boiler combustion efficiency optimization control method based on the Q-Learning algorithm and incorporating a dynamic exploration tendency coefficient. Unlike traditional reinforcement learning which uses a fixed exploration rate, this method dynamically adjusts the algorithm's exploration intensity by comprehensively considering load changes, operational state deviations, and safety constraints. This enables the control system to proactively explore better strategies when operating conditions change, maintain high efficiency and stability during smooth operation, and suppress risky behaviors when approaching dangerous boundaries. Thus, while ensuring safety, it achieves adaptive and refined control of boiler combustion efficiency, significantly improving system energy efficiency and environmental adaptability. Attached Figure Description
[0025] Figure 1 This is a schematic flowchart illustrating a method for optimizing and controlling the combustion efficiency of a heating furnace according to an embodiment of the present invention.
[0026] Figure 2 This is a schematic diagram illustrating the combustion efficiency variation curve and the scatter plot showing the relationship between efficiency and flue gas oxygen content in an embodiment of the present invention.
[0027] Figure 3 This is a schematic diagram illustrating the combined curve of load comparison and adaptive learning urgency change in an embodiment of the present invention;
[0028] Figure 4 This is a schematic diagram illustrating the comprehensive curve of the dynamic exploration tendency coefficient generation process in an embodiment of the present invention. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0031] An example of a method for optimizing and controlling the combustion efficiency of a heating boiler:
[0032] like Figure 1 As shown, the combustion efficiency optimization and control method for a heating boiler according to the present invention includes:
[0033] S1. The operating status of the heating boiler is sensed in real time through multiple sensors, and the current heating load and real-time combustion efficiency are accurately calculated using flue gas analysis and thermal formulas.
[0034] In this embodiment, to achieve precise monitoring and subsequent optimized control of the heating boiler's operating status, it is first necessary to collect several key operational data in real time using multiple sensors deployed on the heating boiler system. In one embodiment, these data may include: exhaust gas temperature obtained through a flue gas analyzer. and oxygen content in furnace flue gas The instantaneous flow rate of fuel supply obtained through the fuel flow meter The opening degree of the combustion fan damper is obtained through the damper opening sensor. and the water supply temperature obtained through temperature sensors. and return water temperature .
[0035] Based on the collected data, the current heating load and real-time combustion efficiency are calculated to provide a basis for subsequent dynamic strategy adjustments. The heating load reflects the current user's immediate heat demand, and the calculation method is as follows:
[0036] ;
[0037] in This indicates the current heating load, expressed in kilowatts (kW). ); The specific heat capacity of water is a physical constant, which can be taken as in this embodiment. ; This indicates the mass flow rate of circulating water measured by the flow meter, in kilograms per second (kg / s). ); This indicates the current water supply temperature measured by the temperature sensor. This represents the current return water temperature measured by the temperature sensor. From this formula, the circulating water mass flow rate can be determined. Or supply and return water temperature difference The increase will directly lead to a decrease in heating load. The increase directly reflects the increase in users' calorie demand.
[0038] In this embodiment, to evaluate the quality of the combustion process, the real-time combustion efficiency is calculated using flue gas analysis. The specific calculation method is as follows:
[0039] ;
[0040] in This represents the real-time combustion efficiency at the current moment, expressed as a dimensionless percentage. This represents the heat loss from flue gas, which is the flue gas temperature. and oxygen content in flue gas The function can be derived from the standard thermal calculation formula; This indicates that the heat input to the fuel is equal to the fuel flow rate. The product of the lower heating value and the lower heating value of the fuel, which represents the heat released when water vapor in the flue gas remains in a gaseous state during the complete combustion of a unit mass or unit volume of fuel, is expressed in kilojoules per kilogram or kilojoules per cubic meter.
[0041] The above-mentioned flue gas heat loss The calculation method is as follows:
[0042] First, obtain the excess air coefficient, which represents the ratio of the actual amount of air supplied during combustion to the theoretically required amount of air for complete combustion. It directly reflects the air-fuel ratio during combustion and can be obtained through the oxygen content of the flue gas. To calculate, we have: ;in The excess air coefficient at the current moment is dimensionless. 21 represents the oxygen content of the flue gas obtained by the flue gas analyzer; 21 represents the approximate volume percentage of oxygen in the air. This represents the parameter tuning factor, used to avoid cases where the denominator is 0.
[0043] Then calculate the flue gas heat loss. ,have:
[0044] ;
[0045] in This indicates the heat loss from flue gas exhaust, which is the heat lost per unit time as the flue gas is discharged. The unit is kilowatt. Indicates the instantaneous flow rate of fuel supply; The theoretical flue gas volume refers to the actual volume of smoke generated when a unit of fuel is completely burned, and is a constant related to the fuel composition. This represents the average isobaric specific heat capacity of the flue gas, which is a constant related to the fuel and temperature. This indicates the excess air coefficient mentioned above; Theoretical air demand refers to the volume of air theoretically required for the complete combustion of a unit of fuel, and is a constant related to the fuel composition. This represents the average isobaric specific heat capacity of air; This indicates the exhaust gas temperature obtained through the flue gas analyzer; This represents the ambient temperature, typically the temperature at which fuel or combustion air enters the heating furnace. The acquisition and calculation of these parameters are mature technologies in this field and will not be elaborated upon here.
[0046] like Figure 2 As shown, it illustrates the key performance indicators calculated above, and the top subplot displays the real-time combustion efficiency calculated using flue gas analysis throughout the entire monitoring period. The fluctuations in the curve reflect the varying quality of the combustion process under different operating conditions. The bottom subplot shows the oxygen content in the flue gas, a key input parameter. With real-time combustion efficiency The relationship between the oxygen content and efficiency levels can be seen from the graph. Different oxygen contents correspond to different efficiency levels.
[0047] S2. By calculating the degree of deviation of the current load from the recent stable level, the urgency of adaptive learning is determined to assess the dynamic adjustment requirements of the algorithm strategy.
[0048] This step calculates the adaptive learning urgency to reflect the degree of deviation of the current heating load from the recent stable level. The specific calculation method is as follows:
[0049] ;
[0050] in This indicates the urgency of adaptive learning at the current moment, representing the degree to which the algorithm needs to adjust its strategy. The gain coefficient is an adjustable parameter. In this embodiment, the gain coefficient... Can be set as experience value . The larger the value, the more sensitive the system is to load fluctuations; that is, even small changes will produce a large response. Values encourage algorithms to explore more actively; The heating load at the current moment obtained in step S1; For past time windows The time-weighted average of heating loads within the specified time frame is used to characterize the recent stable load level. As a preferred approach, an exponential moving average can be used for calculation. ,in This is a smoothing coefficient. In this embodiment, the smoothing coefficient... Preferred Hyperbolic tangent function Make It grows slowly when load fluctuations are small, but rapidly saturates and approaches zero when fluctuations are severe. This forms a bounded learning requirement signal that is sensitive to mutation signals; For example, a very small positive number Used to prevent when the recent average load A division by zero error occurs when the result is zero.
[0051] like Figure 3 As shown, the top subplot compares the heating load at the current moment. (Blue line) and the time-weighted average of heating loads used to characterize recent stable load levels (Red line) The difference between the two curves visually reflects the real-time fluctuation of the load. The bottom subplot shows the final calculated adaptive learning urgency. (Green line). By comparing the two charts, it can be seen that when the gap between the current load (blue line) and the weighted average load (red line) in the top chart increases, that is, when the load fluctuates drastically, the learning urgency (green line) in the bottom chart also increases.
[0052] S3. A multi-dimensional risk assessment is performed by integrating the performance deviation scalar and the safety constraint function to obtain the intelligent dynamic exploration tendency coefficient used to guide the Q-Learning algorithm.
[0053] This step is used to evaluate the system’s multi-dimensional performance deviations and, in conjunction with safety constraints, ultimately obtain an intelligent dynamic exploration tendency coefficient.
[0054] S3.1. A performance deviation vector is constructed by comparing real-time flue gas data with the optimal value under load, and its L2 norm is calculated to calculate the comprehensive performance deviation scalar of the system deviation.
[0055] First, we define a two-dimensional performance deviation vector, whose components represent the degree to which the current combustion state deviates from the optimal operating point in key dimensions.
[0056] ;
[0057] in and For real-time measurement of flue gas oxygen content and exhaust temperature; and At the current heating load Below, the heating load is represented by a function derived from offline calibration or theoretically optimized values provided by engineers. and ; A mapping function can be represented by: This function will determine the deviation value. Mapped to interval, It's about adjusting parameters to control the function's sensitivity to deviations. The smaller the value, the more sensitive the function is to small deviations; These are the weighting coefficients, and This reflects the importance of oxygen content deviation and flue gas temperature deviation on overall performance, and can be set based on experience, for example... .
[0058] In a preferred embodiment, the above and This can be obtained through a lookup table. First, engineers conduct calibration experiments, manually adjusting the damper and fuel supply at different load points (e.g., 20%, 40%, 60%, 80%, 100% load) to find the flue gas oxygen content and exhaust temperature at the point where combustion efficiency is highest and operation is most stable under the current load. Then, these "load → optimal value" correspondences are recorded in a table, as shown in the following example:
[0059] |Load rate Optimal flue gas oxygen content |Optimal smoke exhaust temperature |;
[0060] | 20% | 4.5% | 135°C |;
[0061] | 40% | 4.0% | 145°C |;
[0062] | 60% | 3.5% | 155°C |;
[0063] | 80% | 3.2% | 165°C |;
[0064] | 100% | 3.0% | 180°C |;
[0065] In actual operation, the control system will adjust according to the current load. The current time can be determined by looking up a table or by linear interpolation. and .
[0066] Then, the L2 norm of the above performance deviation vector is calculated to obtain a comprehensive performance deviation scalar, specifically:
[0067] ;
[0068] The value in Within the interval, this value is used to assess the overall distance of the current state from the optimal point. A larger value indicates a more severe deviation and worse system performance.
[0069] S3.2. Obtain the safety constraint manifold function based on the safety threshold of flue gas oxygen content and exhaust temperature. When the system approaches the danger boundary, its value approaches zero to suppress exploration.
[0070] To ensure the safety of the exploration process, a safety constraint manifold function, also denoted as a safety constraint factor, is needed as a constraint on the exploration behavior, specifically:
[0071] ;
[0072] in This represents a safety-constrained manifold function with a range of . When the system is far from the danger boundary, its value is close to This allows for ample exploration; its value will rapidly approach a critical boundary when approaching any dangerous boundary. This inhibits exploratory behavior; Represents the Sigmoid function; This indicates the minimum safe threshold for flue gas oxygen content to prevent flameout, for example... ; This indicates the maximum safe threshold for exhaust gas temperature to prevent equipment from overheating or excessive heat loss, for example... . The slope factor, representing a positive value, determines how steeply the function decreases as it approaches the threshold. Smaller values... A value that makes the function curve extremely steep near the threshold creates a hard safety margin, while a larger value... The value provides a transition zone.
[0073] S3.3. By combining the learning urgency, performance deviation and safety constraint function, a dynamic exploration tendency coefficient that can respond to changes in working conditions and is constrained by safety boundaries is obtained, and this coefficient is used as the actual exploration rate of the Q-Learning algorithm.
[0074] Finally, obtain the dynamic exploration tendency coefficient. Specifically:
[0075] ;
[0076] in The dynamic exploration tendency coefficient at the current moment will be used as the actual exploration rate of the Q-Learning algorithm; Represents the base exploration rate, a small constant, for example This is used to ensure that the system still has a slight continuous learning ability in a stable state; This represents the maximum exploration rate and defines the upper limit of exploration intensity, for example... ; Represents a safety-constrained manifold function; Indicates the urgency of adaptive learning; Indicates a scalar value representing performance deviation; This represents the exploration gain coefficient, a positive adjustable parameter used to adjust the exploration intensity in response to exploration demands. The response sensitivity is preferably set in the range of 2-5, and in this embodiment it can be set to an empirical value of 3; It is a natural exponential function.
[0077] As can be seen from the above formula, when the system is in a steady state, and All are relatively small, and the values inside the square brackets are close to , making Falling back to base exploration rate When the load fluctuates drastically (i.e. (increase) and combustion performance deviates from optimal (i.e.) When the product term increases, Significantly increased, leading to The term approaches At this time, the tendency to explore Towards the maximum exploration rate Improvement. However, if the system state approaches the safety boundary, safety constraints... It will approach Forcibly suppressing the tendency to explore to The nearest location is prioritized to ensure system security.
[0078] like Figure 4 As shown, the top subplot displays the overall performance deviation scalar. (Red line) represents the total distance of the current combustion state from the theoretical optimal operating point. The higher the curve value, the more serious the system performance deviation and the more urgent the need for optimization.
[0079] The middle subgraph shows the safety-constrained manifold function. (Green line) When its value is close to 1, it means that the system is operating within the safe range; when its value drops significantly, as shown in the time point 400-1000 in the figure, it means that the system state is approaching the safety boundary, such as the minimum flue gas oxygen content or the maximum exhaust temperature.
[0080] The bottom subplot shows the final dynamic exploration tendency coefficient. (Blue line) This curve embodies the core logic of this invention:
[0081] When the system performance deviation is large (high red line) and it is in the safe zone (high green line), the exploration tendency coefficient (blue line) will increase accordingly, prompting the algorithm to actively explore.
[0082] When the system approaches the danger boundary (the green line drops significantly), even if performance deviations still exist, safety constraints... It will also forcibly reduce the exploration tendency to near the basic level, thereby prioritizing system security.
[0083] S4. The dynamic exploration tendency coefficient is embedded into the Q-Learning algorithm to continuously optimize combustion efficiency within the safe boundary by adjusting the exploration and utilization strategy in real time.
[0084] This step embeds the dynamic exploration tendency coefficient obtained in the previous steps into the action selection strategy of the Q-Learning algorithm in real time.
[0085] In one embodiment, the controller's state space It consists of discrete furnace temperature, flue gas oxygen content, current load, etc., denoted as Action space Discrete commands, including those for fine-tuning the fuel supply rate and damper opening, are denoted as... At each control decision moment The controller performs the following operations:
[0086] First, obtain the current state. Then, based on steps S1 to S3 above, calculate the dynamic exploration tendency coefficient at the current moment. Then generate one in Uniformly distributed random numbers within the interval Continue with action selection: If Then, exploration is performed in the action space. Randomly select an action .like Then, the following will be executed: select the option that allows the current state to be optimized. The action with the largest Q value, i.e. Then perform the selected action on the heating furnace. After waiting for one control cycle, observe the new state the system transitions to. And calculate the reward based on the preset reward function. The reward function can be designed as an increasing function of the real-time combustion efficiency calculated in step S1, and a decreasing function of energy consumption and pollutant emissions; finally, according to the classic Q-Learning update rule, the Q values of the corresponding state-action pairs in the Q-value table are updated, resulting in:
[0087] ;
[0088] in In the state Next action Q value; The learning rate; To perform the action The immediate reward obtained afterward; Discount factor; In the new state The maximum possible future Q value. The specific implementation of the above Q-Learning is a well-known technology and will not be described in detail here.
[0089] By repeatedly executing the above steps, the control system can intelligently adjust its exploration behavior according to changes in external load and its own operating status, thereby quickly adapting to new operating conditions and continuously optimizing combustion efficiency while ensuring safety.
Claims
1. A method for optimizing and controlling the combustion efficiency of a heating boiler, characterized in that, The method includes: Obtain the current operating parameters of the heating boiler, including flue gas oxygen content, flue gas temperature, and current heating load; apply the dynamic exploration tendency coefficient as the exploration rate to the Q-Learning control algorithm to adjust fuel supply and combustion air; in each control cycle, select to execute an exploration action or a utilization action based on the exploration rate and a generated random number, and update the Q value; The dynamic exploration tendency coefficient is calculated based on the baseline exploration rate, combined with adaptive learning urgency, performance deviation, and safety constraint factors. Specifically: ; in This represents the dynamic exploration tendency coefficient at the current moment; The base exploration rate is a small constant. This indicates the maximum preset exploration rate, defining the upper limit of exploration intensity; This represents the security constraint factor; This indicates the urgency of the adaptive learning; This indicates the performance deviation; This indicates the set exploration gain coefficient; It is a natural exponential function; The adaptive learning urgency is used to characterize the degree of change of the current heating load compared to the recent average heating load; the performance deviation is used to characterize the degree of deviation of the current flue gas oxygen content and exhaust gas temperature from the optimal flue gas oxygen content and optimal exhaust gas temperature corresponding to the current heating load, respectively. The value of the safety constraint factor approaches 1 under safe operating conditions where the current flue gas oxygen content is higher than the set minimum safety threshold and the exhaust gas temperature is lower than the set maximum safety threshold, and rapidly approaches 0 when the current operating parameters approach any set safety threshold to suppress exploration behavior.
2. The method for optimizing and controlling the combustion efficiency of a heating boiler according to claim 1, characterized in that, The security constraint factor is the product of two functions: The value of the first function decreases from 1 to 0 as the difference between the current oxygen content in the flue gas and the minimum safety threshold decreases; The second function value decreases from 1 to 0 as the difference between the highest safety threshold and the current exhaust temperature decreases.
3. The method for optimizing and controlling the combustion efficiency of a heating boiler according to claim 1, characterized in that, The performance deviation is the L2 norm of the performance deviation vector; The performance deviation vector includes the following components: The first component is the deviation between the current flue gas oxygen content and the optimal flue gas oxygen content multiplied by the first weight. The second component is the deviation between the current exhaust temperature and the optimal exhaust temperature multiplied by the second weight; The sum of the first weight and the second weight is 1.
4. The method for optimizing and controlling the combustion efficiency of a heating boiler according to claim 1, characterized in that, The method for calculating the adaptive learning urgency is as follows: ; in This indicates the urgency of adaptive learning at the current moment; This indicates the set gain coefficient; Indicates the current heating load; This represents the average heating load in the recent period; It is a very small positive number, used to prevent the recent average load from being used as a reference. A division by zero error occurs when the result is zero.
5. The method for optimizing and controlling the combustion efficiency of a heating boiler according to claim 4, characterized in that, The method for obtaining the current heating load is as follows: The mass flow rate of circulating water is measured using a flow meter. The supply and return water temperatures are measured using temperature sensors, and the supply and return water temperature difference is calculated. The product of the circulating water mass flow rate, the specific heat capacity of the water, and the temperature difference between the supply and return water is taken as the current heating load.
6. The method for optimizing and controlling the combustion efficiency of a heating boiler according to claim 4, characterized in that, The recent average heating load It is calculated using the exponential moving average method.
7. The method for optimizing and controlling the combustion efficiency of a heating boiler according to claim 1, characterized in that, The optimal flue gas oxygen content and optimal exhaust temperature are obtained by looking up a pre-built lookup table; The lookup table stores the mapping relationship between multiple discrete heating load points and the corresponding optimal flue gas oxygen content and optimal flue gas temperature.
8. The method for optimizing and controlling the combustion efficiency of a heating boiler according to claim 7, characterized in that, The method for constructing the lookup table is as follows: Under multiple different heating load conditions, the fuel supply and combustion air volume of the heating boiler are manually adjusted to determine the flue gas oxygen content and exhaust temperature at which the highest combustion efficiency is achieved under the current load, and these values are recorded as the preset optimal values for the corresponding load points.
9. The method for optimizing and controlling the combustion efficiency of a heating boiler according to claim 1, characterized in that, The reward function used in the Q-Learning control algorithm is set based on at least one of the boiler's real-time combustion efficiency, energy consumption, and pollutant emissions.
Citation Information
Patent Citations
Wall-hanging stove constant-temperature energy-saving control system based on enhanced autonomous learning and SAC algorithm
CN119983568A
Water electrolysis efficiency dynamic scheduling method and system based on reinforcement learning
CN120485871A