Waste incineration field task optimization method and system applying reinforcement learning
By optimizing the waste incineration process through reinforcement learning strategy networks, the problem of combustion state imbalance in traditional methods is solved, achieving a balance between energy efficiency improvement, pollutant emission reduction and equipment maintenance, and dynamically adjusting the control strategy to adapt to process changes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-03-17
AI Technical Summary
Existing technologies struggle to achieve a dynamic balance between improving energy efficiency, reducing pollutant emissions, and maintaining equipment during waste incineration, and traditional control methods are prone to causing imbalances in the combustion state.
A reinforcement learning strategy network is used to encode the real-time operation data of the waste incineration plant to generate a state representation vector. The incineration temperature and oxygen supply are optimized through a dual-channel output structure. The multi-objective reward value is calculated by combining the reward function to generate optimized control commands.
It achieves improved energy utilization efficiency, reduced pollutant emissions, and extended equipment lifespan during waste incineration, and dynamically adjusts control strategies to adapt to process changes.
Smart Images

Figure CN121676971A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing, and more specifically, to a method and system for optimizing waste incineration plant tasks using reinforcement learning. Background Technology
[0002] In the field of waste incineration, optimizing the incineration process is crucial for improving energy efficiency, controlling pollutant emissions, and ensuring stable equipment operation. Currently, the common approach is to collect waste composition data and incinerator operating parameters using sensors, and then generate control commands such as temperature and oxygen supply based on preset control logic or simplified mathematical models. However, the time-varying nature of waste composition and the complexity of physicochemical reactions within the incinerator make it difficult for preset models to accurately capture the dynamic changes in the combustion state. This leads to deviations between control commands and actual combustion requirements. Furthermore, when addressing multi-objective optimization problems such as improving thermal efficiency, reducing pollutant emissions, and maintaining equipment, traditional methods struggle to achieve a dynamic balance between these objectives, often resulting in compromises and impacting the overall optimization effect of the incineration process. Summary of the Invention
[0003] This invention provides a method and system for optimizing waste incineration plant tasks using reinforcement learning.
[0004] In a first aspect, embodiments of the present invention provide a method for optimizing a waste incineration plant task using reinforcement learning. The method includes: acquiring a real-time operational data set of the waste incineration plant, the real-time operational data set including waste composition data and incinerator operating parameter data recorded by the incinerator control system; the waste composition data including measured values of waste moisture content and waste combustible material content; and the incinerator operating parameter data including measured values of furnace temperature and flue gas oxygen content; based on the real-time operational data set, performing state encoding processing to extract waste calorific value distribution features and incinerator operating state features, and performing vectorized fusion processing on the waste calorific value distribution features and incinerator operating state features to generate a state representation vector; and invoking a pre-trained reinforcement learning policy. The reinforcement learning policy network performs policy reasoning on the state representation vector. The network employs a dual-channel output structure, where the main channel outputs the incineration temperature setpoint and the secondary channel outputs the oxygen supply adjustment value. The outputs of the main and secondary channels are combined to generate a control action sequence. Based on this sequence, the performance of the incineration process is evaluated using a reward function. This reward function calculates a weighted sum of thermal efficiency optimization indicators, pollutant emission control indicators, and equipment lifespan maintenance indicators to generate a multi-objective reward value. Based on this multi-objective reward value, the policy network parameter update process is performed, adjusting the weight parameters of the reinforcement learning policy network and generating optimized control commands. These commands are then sent to the waste incinerator control system to adjust the incinerator's combustion conditions.
[0005] Secondly, embodiments of the present invention provide a computer system, comprising: a memory storing a computer program; and a processor for loading the computer program to implement the above-described method for optimizing waste incineration plant tasks using reinforcement learning.
[0006] This invention provides a task optimization method for waste incineration plants using reinforcement learning. By receiving a real-time operational data set composed of waste composition analysis data and incinerator operating parameter data, it accurately focuses on the key impacts of waste characteristics and equipment operating status on the incineration process, improving the matching degree between data and core related factors of the incineration process. Based on this real-time operational data set, it performs state encoding processing, extracting waste calorific value distribution features and incinerator operating status features, and then vectorizing and fusing them to generate state representation vectors. This achieves a unified representation of waste characteristics and equipment operating status across physical objects, avoiding the limitations of single-dimensional features in describing the overall state of the incineration process. A pre-trained reinforcement learning policy network is applied to perform policy inference on the state representation vectors, and a dual-channel output structure is used to output the respective... The system generates a control action sequence by adjusting the incineration temperature setpoint and oxygen supply, ensuring coordinated adaptation of core control parameters and avoiding combustion imbalances caused by traditional single-parameter adjustments. Based on the control action sequence, a multi-objective reward value is generated by calculating the weighted sum of thermal efficiency optimization indicators, pollutant emission control indicators, and equipment life maintenance indicators through a reward function. This balances the economic, environmental, and safety requirements of the incineration process, avoiding the problem of neglecting one aspect for another caused by single-objective optimization. Based on the multi-objective reward value, the system updates the strategy network parameters and generates optimized control commands, which are then sent to the incinerator control system. This allows the control strategy to dynamically adjust with changes in the incineration process, improving its adaptability to fluctuations in waste composition and changes in equipment operating status, thereby effectively optimizing the combustion conditions of the incinerator. Attached Figure Description
[0007] Figure 1 This is a flowchart of a waste incineration plant task optimization method using reinforcement learning, provided by an embodiment of the present invention.
[0008] Figure 2 This is a schematic diagram of the composition of a computer system provided in an embodiment of the present invention. Detailed Implementation
[0009] Please see Figure 1 , Figure 1 A flowchart illustrating a task optimization method for waste incineration plants using reinforcement learning, provided as an embodiment of the present invention, is included. This method can be executed by a computer system, such as an industrial computer system controlling the incineration process in a waste incineration plant. The method may include the following steps: Step S100: Obtain the real-time operation data set of the waste incineration plant. The real-time operation data set includes waste composition data and incinerator operation parameter data recorded by the incinerator control system. The waste composition data includes the measured values of waste moisture content and waste combustible content. The incinerator operation parameter data includes the measured values of furnace temperature and flue gas oxygen content.
[0010] Waste composition data reflects the physical and chemical properties of the waste itself. Among these, the moisture content measurement indicates the amount of water contained in the waste. Water evaporation requires heat absorption, so excessively high moisture content increases energy consumption during incineration and reduces incineration efficiency. Moisture detection equipment can be used to obtain waste moisture content measurements. For example, a moisture sensor based on the principle of capacitance can be used. When a waste sample comes into contact with the sensor, the capacitance value changes due to variations in moisture content. By measuring the capacitance value and applying a calibration algorithm, the moisture content of the waste can be determined.
[0011] The combustible material content of waste indicates the proportion of substances in the waste that can participate in the combustion reaction. These combustible materials are the main source of heat generated during waste incineration. Thermogravimetric analysis (TGA) can be used to determine the combustible material content of waste. The waste sample is placed in a thermogravimetric analyzer; during the heating process, the combustible materials in the waste gradually burn and decompose. By measuring the change in sample mass, the combustible material content can be calculated.
[0012] The measured temperature inside the furnace directly reflects the thermal environment inside the incinerator. To measure the furnace temperature, high-temperature resistant temperature sensors, such as thermocouple temperature sensors, can be used. Thermocouples are based on the thermoelectric effect; when the furnace temperature changes, the thermocouple generates a corresponding electromotive force. By measuring and converting the electromotive force, an accurate furnace temperature value can be obtained.
[0013] The measured oxygen content in flue gas reflects the oxygen supply and the degree of combustion during incineration. Sufficient oxygen is essential for complete combustion of waste. If the oxygen content is too low, incomplete combustion may occur, producing pollutants such as carbon monoxide; if the oxygen content is too high, energy will be wasted. A zirconia oxygen sensor can be used to measure the oxygen content in flue gas. Zirconia has the property of conducting oxygen ions at high temperatures. When the oxygen concentrations on both sides of the sensor are different, an oxygen concentration gradient electromotive force is generated. By measuring this electromotive force, the oxygen content in the flue gas can be calculated.
[0014] Step S200: Based on the real-time operation data set, perform state coding processing to extract the waste calorific value distribution characteristics and incinerator operation status characteristics, and perform vectorized fusion processing on the waste calorific value distribution characteristics and incinerator operation status characteristics to generate a state representation vector.
[0015] Status coding processing is a process of in-depth mining and transformation of real-time operational data sets. Its purpose is to extract features from raw, complex data that reflect key information about the waste incineration process. Waste calorific value distribution characteristics describe the distribution of calorific values in different regions or time periods. It is closely related to waste composition data because different waste components have different calorific values. By analyzing and processing waste composition data, the distribution pattern of waste calorific values can be inferred. Incinerator operating status characteristics reflect the incinerator's operational status at the current moment, determined by incinerator operating parameter data. For example, the stability of the furnace temperature and fluctuations in flue gas oxygen content can reflect whether the incinerator is operating well.
[0016] In one implementation, step S200 may include the following steps S210-S260: Step S210: Divide the real-time operation data set into continuous sampling window units. Each sampling window unit contains waste composition data and incinerator operation parameter data within a preset time period, and establish a data time-series correlation index.
[0017] A sampling window unit is a data unit obtained by dividing continuous real-time operational data according to a preset duration. The preset duration needs to be determined based on the actual operation of the waste incineration plant and the analysis requirements. For example, if a more detailed analysis of the dynamic changes in the waste incineration process is required, the preset duration can be set to be shorter; if only the overall trend over a longer period is considered, the preset duration can be set to be longer.
[0018] Each sampling window unit contains waste composition data and incinerator operating parameter data within a preset time period. This division allows for the discretization of continuous data, facilitating subsequent analysis and processing. The data time-series correlation index is created to establish the temporal order relationship between different sampling window units and within each sampling window unit. It can use timestamps to mark the start and end times of each sampling window unit, and simultaneously mark the corresponding time for each data point. Through the data time-series correlation index, data from different time points can be easily found and processed, ensuring the correct temporal order of the data, thus providing accurate time references for subsequent analysis.
[0019] Step S220: Perform regional correlation analysis on the measured values of waste moisture content and waste combustible content in each sampling window unit, calculate the calorific value contribution weight of different stacking areas, and generate regional calorific value change curves.
[0020] Regional correlation analysis delves into the intrinsic relationship between the moisture content and combustible material content of waste in different landfill areas, and the impact of these factors on the calorific value of waste. Due to differences in factors such as source, landfill time, and environment, the moisture content and combustible material content of waste in different landfill areas can vary significantly. These differences directly result in different calorific value contributions from waste in different areas.
[0021] Calorific value contribution weight is an indicator that measures the degree to which waste from each storage area contributes to the total calorific value of the entire waste incineration plant. By calculating the calorific value contribution weight of different storage areas, the importance of waste from each area in the incineration process can be clarified, providing a basis for optimizing incineration strategies.
[0022] In one implementation, step S220 may include the following steps S221-S226: Step S221: Divide the waste feeding area of the waste incineration plant into multi-row, multi-column sub-areas according to the grid division rules, and collect the measured values of waste moisture content and waste combustible content in the sub-areas.
[0023] Grid partitioning is a method of uniformly dividing the waste feeding area. By dividing it into multiple sub-regions, the characteristics of waste in different areas can be analyzed in more detail. For example, the waste feeding area can be divided into regular grids based on its actual size and shape, with each grid being a sub-region.
[0024] Multiple sampling points can be set up within each sub-region. For collecting measurements of waste moisture content, moisture detection equipment can be used to collect a certain amount of waste sample at each sampling point for testing. For collecting measurements of waste combustible material content, waste samples can be collected at each sampling point, and methods such as thermogravimetric analysis can be used to determine the combustible material content.
[0025] To ensure the accuracy and representativeness of the data, sampling points can be set at different locations in the sub-region (such as the four corners and the center), and the collected data can be averaged to obtain the measured values of the moisture content and combustible content of the waste in that sub-region.
[0026] Step S222: For each sub-region, based on the preset calorific value estimation model, the measured values of the waste moisture content and the waste combustible content of the sub-region are converted into calorific value per unit mass. The calorific value estimation model is obtained by fitting historical experimental data. The input is moisture content and combustible content, and the output is calorific value per unit mass.
[0027] The pre-defined calorific value estimation model is a mathematical model built upon a large amount of historical experimental data, describing the quantitative relationship between the moisture content, combustible material content, and calorific value per unit mass of waste. Calorific value per unit mass refers to the heat released when a unit mass of waste is completely burned; it is an indicator of waste combustion performance. The model was established by collecting a large number of waste samples with different moisture contents and combustible material contents, conducting combustion experiments, and measuring the calorific value per unit mass for each sample. Then, statistical analysis methods were used to fit these experimental data to obtain a mathematical model that accurately describes the relationship between the three factors. For example, multiple linear regression can be used to fit the linear relationship between moisture content, combustible material content, and calorific value per unit mass using the least squares method. The measured moisture content and combustible material content of waste in each sub-region are input into the pre-defined calorific value estimation model, and the model calculates the calorific value per unit mass for that sub-region based on its internal algorithm and parameters.
[0028] Step S223: Calculate the product of the unit mass calorific value of each sub-region and the amount of waste fed into that sub-region to obtain the calorific value contribution of that sub-region. Calculate the sum of the calorific value contributions of all sub-regions as the total calorific value of the sampling window unit.
[0029] The calorific value contribution of a sub-region represents the actual contribution of the waste in that sub-region to the total calorific value during incineration. The calorific value per unit mass reflects the combustion capacity of the waste in that sub-region, while the waste feed rate indicates the quantity of waste participating in incineration in that sub-region. Multiplying the calorific value per unit mass by the waste feed rate yields the calorific value contribution of that sub-region. To accurately obtain the waste feed rate of each sub-region, a high-precision weighing device can be installed at the waste inlet to record the waste feed weight of each sub-region in real time. Then, multiplying the calorific value per unit mass of each sub-region by the corresponding waste feed rate yields the calorific value contribution of each sub-region. The sum of the calorific value contributions of all sub-regions gives the total calorific value of the sampling window unit. The total calorific value reflects all the heat released by waste incineration within that sampling window unit. By calculating the total calorific value, the energy output of the waste incineration plant during that time period can be understood.
[0030] Step S224: Divide the calorific value contribution of each sub-region by the total calorific value to obtain the calorific value contribution weight of that sub-region, and construct a regional calorific value contribution weight matrix, where the matrix elements are the calorific value contribution weights of the corresponding sub-regions.
[0031] The regional calorific value contribution weight matrix is a matrix used to describe the calorific value contribution weights of all sub-regions. The rows and columns of the matrix correspond to different sub-regions, and the matrix elements are the calorific value contribution weights of the corresponding sub-regions. In actual matrix construction, a two-dimensional array can be used to store these weight values. The calorific value contribution of each sub-region is divided by the total calorific value of the sampling window cells to obtain the calorific value contribution weight of that sub-region, and then this weight value is stored in the corresponding position in the matrix.
[0032] Step S225: Extract the regional heat value contribution weight matrix of each window in sequence according to the timestamp order of the sampling window units, calculate the weight change rate of the same sub-region in the continuous window, and generate a set of regional heat value change curves containing multiple rows and columns of curves.
[0033] By sequentially extracting the regional calorific value contribution weight matrix for each window according to timestamps, the changes in the calorific value contribution weight of each sub-region over different time periods can be clearly observed. The weight change rate refers to the proportion of change in the calorific value contribution weight of the same sub-region in two consecutive sampling window units, reflecting the dynamic changes in the waste calorific value contribution of that sub-region over different time periods. Calculating the weight change rate can help operators promptly detect abnormal fluctuations in the calorific value contribution of sub-regions, so that appropriate measures can be taken for adjustment.
[0034] The purpose of generating a set of regional calorific value change curves containing multiple rows and columns is to visually demonstrate how the calorific value contribution of waste in each sub-region changes over time. Each row of curves corresponds to a sub-region, and each column of curves corresponds to a time dimension. By observing the trends and fluctuations of these curves, one can quickly understand the changing trends of waste calorific value contribution in different sub-regions over different time periods.
[0035] Step S226: The set of regional calorific value change curves is denoised using a curve smoothing algorithm to eliminate high-frequency noise interference, retain the main trend of regional calorific value change, and generate the final regional calorific value change curve.
[0036] In a set of regional calorific value variation curves, high-frequency noise interference may appear due to factors such as measurement errors and random fluctuations, affecting the observation and analysis of the main trends in regional calorific value changes. Curve smoothing algorithms (such as moving average algorithms and Savitzky-Golay filtering algorithms) can remove noise interference from the curves, making them smoother.
[0037] Step S230: Based on the spatiotemporal distribution patterns of regional calorific value change curves and furnace temperature measurements, construct a physical field coupling matrix. The row dimension of the physical field coupling matrix corresponds to the waste dumping area number, the column dimension corresponds to the incinerator operating parameter type, and the matrix elements represent the correlation strength between regional calorific value and operating parameters.
[0038] The regional calorific value variation curve reflects the change of calorific value over time in different waste dumping areas, while the furnace temperature measurement values reflect the distribution of the incinerator's internal temperature at different locations and times. The spatiotemporal distribution pattern refers to the characteristics and interrelationships of these two factors in time and space.
[0039] The physics coupling matrix is used to describe the relationship between regional calorific value and incinerator operating parameters. The row dimension corresponds to the waste dumping zone number, meaning each row represents a waste dumping zone. The column dimension corresponds to the type of incinerator operating parameter, such as furnace temperature and flue gas oxygen content; each column represents a single operating parameter. Matrix elements represent the correlation strength between regional calorific value and operating parameters; a stronger correlation indicates a greater impact of the waste calorific value on the corresponding operating parameter.
[0040] Step S240: Use the physical field coupling matrix to perform feature filtering on waste composition data and incinerator operation parameter data, retain feature dimensions with correlation strength exceeding a preset threshold, and generate dimensionality-reduced waste calorific value distribution features and incinerator operation status features.
[0041] The physical field coupling matrix reflects the correlation strength between regional calorific value and incinerator operating parameters. Feature filtering of waste composition data and incinerator operating parameter data using the physical field coupling matrix aims to select feature dimensions closely correlated with regional calorific value from the original, high-dimensional data, while removing those with weaker correlations, thus achieving data dimensionality reduction. A preset threshold is a pre-defined critical value for correlation strength. Only when the correlation strength between a feature dimension and regional calorific value exceeds this threshold will it be retained. This aims to reduce data redundancy and improve data quality and analysis efficiency. During feature filtering, each element of the physical field coupling matrix is traversed. For each element, if its value exceeds the preset threshold, the corresponding feature dimension of the waste composition data or incinerator operating parameter data will be retained; if its value is below the preset threshold, the feature dimension will be deleted. After feature filtering, the generated dimensionality-reduced waste calorific value distribution features and incinerator operating status features only contain key information strongly correlated with regional calorific value.
[0042] Step S250: Calculate the fusion coefficient of the waste calorific value distribution characteristics and the incinerator operating status characteristics through a dynamic weight allocation algorithm. The fusion coefficient is dynamically adjusted according to the timestamp of the sampling window unit to reflect the influence weight of the two types of characteristics on the incineration process in different time periods.
[0043] The dynamic weighting algorithm is an algorithm that dynamically adjusts the fusion coefficient based on the influence of the waste calorific value distribution characteristics and the incinerator operating status characteristics on the incineration process over different time periods. The fusion coefficient is used to measure the relative importance of the waste calorific value distribution characteristics and the incinerator operating status characteristics in the fusion process.
[0044] The composition of waste and the operating status of the incinerator may change over different time periods, thus affecting the weight of these two characteristics on the incineration process. For example, the calorific value distribution of waste may have a greater impact on the incineration process during certain time periods, while the operating status of the incinerator may be more critical during other time periods.
[0045] The dynamic weight allocation algorithm dynamically adjusts the fusion coefficient based on the timestamp of the sampling window unit, combining historical and real-time data. In practical applications, machine learning methods, such as constructing a neural network model, can be used. The neural network is trained using historical data on waste calorific value distribution characteristics, incinerator operating status characteristics, and corresponding incineration effects from different periods. After training, the model learns the mapping relationship between different features and incineration effects. When real-time data for the current period is input, the neural network outputs the corresponding fusion coefficient. If the current waste composition fluctuates significantly, the influence of calorific value distribution characteristics is enhanced, and the algorithm increases their corresponding weights; if the incinerator is unstable, the weight of operating status characteristics will be increased, thus achieving dynamic adjustment of the fusion coefficient.
[0046] Step S260: Based on the fusion coefficient, the dimensionality-reduced waste calorific value distribution characteristics and incinerator operating status characteristics are weighted and vectorized to generate a state representation vector containing spatiotemporal correlation. Each dimension of the state representation vector corresponds to the weighted fused feature value.
[0047] Weighted vector fusion is the process of integrating the dimensionality-reduced waste calorific value distribution characteristics and incinerator operating status characteristics according to a fusion coefficient. The fusion coefficient reflects the relative importance of these two characteristics to the incineration process within the current time period. The dimensionality-reduced waste calorific value distribution characteristics and incinerator operating status characteristics are represented as vectors. Then, these two vectors are weighted according to the fusion coefficient. For the waste calorific value distribution characteristic vector, the value of each dimension is multiplied by the weight corresponding to that characteristic in the fusion coefficient; for the incinerator operating status characteristic vector, the value of each dimension is multiplied by the weight corresponding to that characteristic in the fusion coefficient. Finally, the values of the corresponding dimensions of the two weighted vectors are summed to obtain a new vector, which is the state representation vector containing spatiotemporal correlations. Each dimension of the state representation vector corresponds to a weighted fused feature value, integrating information from both waste calorific value distribution characteristics and incinerator operating status characteristics, and considering the influence weights of the two types of features within different time periods.
[0048] Step S300: Call the pre-trained reinforcement learning policy network to perform policy reasoning processing on the state representation vector. The reinforcement learning policy network adopts a dual-channel output structure, in which the main channel outputs the incineration temperature setpoint and the secondary channel outputs the oxygen supply adjustment value. The outputs of the main channel and the secondary channel are combined to generate a control action sequence.
[0049] In one implementation, step S300 may include the following steps S310-S360: Step S310: Input the state representation vector into the feature divide-and-conquer layer of the reinforcement learning policy network, and divide the state representation vector into calorific value feature sub-vector and operation feature sub-vector according to the physical meaning of the feature dimensions. The calorific value feature sub-vector contains dimensions related to the calorific value distribution of waste, and the operation feature sub-vector contains dimensions related to the operating status of the incinerator.
[0050] The feature divide-and-conquer layer is used to decompose and classify the input state representation vector. The state representation vector contains various information such as the calorific value distribution characteristics of waste and the operating status characteristics of the incinerator. In order to facilitate subsequent processing and analysis, it needs to be divided into different sub-vectors.
[0051] Based on the physical meaning of the feature dimensions, the state representation vector is divided into a calorific value feature sub-vector and an operational feature sub-vector. The calorific value feature sub-vector contains dimensional information related to the calorific value distribution of waste, reflecting the combustion performance and energy distribution of the waste. The operational feature sub-vector contains dimensional information related to the operating status of the incinerator, such as furnace temperature and flue gas oxygen content, reflecting the actual operating condition of the incinerator.
[0052] In the feature divide-and-conquer layer, the state representation vector can be divided programmatically. For example, based on the index and corresponding physical meaning of each dimension in the state representation vector, elements belonging to the dimensions related to the calorific value distribution of waste can be extracted to form a calorific value feature sub-vector; elements belonging to the dimensions related to the incinerator's operating state can be extracted to form an operating feature sub-vector.
[0053] Step S320: Calculate the mutual information entropy between the heat value feature vector and the running feature vector through the cross-attention module of the reinforcement learning policy network. Generate a feature association weight matrix based on the mutual information entropy. The row dimension of the feature association weight matrix corresponds to the dimension of the heat value feature vector, the column dimension corresponds to the dimension of the running feature vector, and the matrix elements represent the degree of association between the two types of feature dimensions.
[0054] In one implementation, step S320 may specifically include the following steps S321-S325: Step S321: Use the heat value feature vector as the query feature vector and the running feature vector as the key feature vector. Perform dimension mapping on the query feature vector and the key feature vector respectively through a linear transformation matrix so that the mapped query feature vector and the key feature vector have the same dimension.
[0055] In the cross-attention mechanism, the heat value feature vector is used as the query feature vector, and the running feature vector is used as the key feature vector. This is because the query feature vector is used to find related information, while the key feature vector is used to provide information available for querying.
[0056] To enable efficient comparison and correlation calculations between query feature vectors and key feature vectors, a linear transformation matrix is needed to map their dimensions. This linear transformation matrix is a predefined matrix that adjusts the dimensions of the query and key feature vectors to ensure they have the same dimensions.
[0057] Specifically, the query feature vector is multiplied by a linear transformation matrix to obtain the mapped query feature vector; the key feature vector is multiplied by another linear transformation matrix to obtain the mapped key feature vector. The design of these two linear transformation matrices needs to be adjusted based on the original and target dimensions of the query and key feature vectors to ensure that the mapped vectors have the same dimensions, thus facilitating subsequent calculations.
[0058] Step S322: Calculate the dot product of the mapped query feature vector and the key feature vector to obtain the initial attention score matrix. The elements of the initial attention score matrix represent the original association scores of the i-th dimension of the query feature vector and the j-th dimension of the key feature vector.
[0059] The initial attention score matrix is obtained by calculating the dot product of the mapped query feature vector and the key feature vector. This initial attention score matrix is a two-dimensional matrix, where the row dimensions correspond to the dimensions of the query feature vectors and the column dimensions correspond to the dimensions of the key feature vectors. Each element in the matrix represents the original association score between the i-th dimension of the query feature vector and the j-th dimension of the key feature vector. This score reflects the similarity or correlation between the two dimensions; a higher score indicates a stronger association between the two dimensions.
[0060] Step S323: Perform row normalization on the initial attention score matrix to obtain the normalized attention score matrix. Calculate the mutual information entropy of the query feature vector and the key feature vector based on the normalized attention score matrix. The mutual information entropy is calculated by multiplying each element of the normalized attention score matrix by its logarithm and summing the results, taking the negative value as the mutual information entropy value.
[0061] During row normalization, for each row of the initial attention score matrix, each element of that row is divided by the sum of all elements in that row to obtain the normalized element value. After row normalization, the normalized attention score matrix is obtained.
[0062] The mutual information entropy of the query feature vector and the key feature vector is calculated based on the normalized attention score matrix. The mutual information entropy is calculated by multiplying each element of the normalized attention score matrix by its logarithm, summing the results, and then taking the negative value as the mutual information entropy value. Mutual information entropy reflects the uncertainty and correlation between the query feature vector and the key feature vector. A larger mutual information entropy value indicates a stronger correlation between the two vectors; a smaller mutual information entropy value indicates a weaker correlation between the two vectors.
[0063] Step S324: Compare the mutual information entropy value with the preset benchmark mutual information entropy, calculate the mutual information entropy deviation rate, and adjust the scaling factor of the normalized attention score matrix based on the mutual information entropy deviation rate so that the adjusted matrix element values reflect the actual correlation strength.
[0064] The preset baseline mutual information entropy is a pre-defined reference value used to measure the degree of normal correlation between query feature vectors and key feature vectors. The calculated mutual information entropy value is compared with the preset baseline mutual information entropy to calculate the mutual information entropy deviation rate. The mutual information entropy deviation rate reflects the degree of difference between the actual mutual information entropy value and the baseline mutual information entropy value.
[0065] The scaling factor of the normalized attention score matrix is adjusted based on the mutual information entropy deviation rate. The scaling factor adjusts the element values of the normalized attention score matrix to more accurately reflect the actual correlation strength between the query feature vector and the key feature vector. If the mutual information entropy deviation rate is large, it indicates a significant difference between the actual correlation and the baseline correlation; in this case, the scaling factor needs to be adjusted to increase or decrease the matrix element values accordingly. If the mutual information entropy deviation rate is small, it indicates that the actual correlation is close to the baseline correlation; in this case, the scaling factor can be kept unchanged or adjusted slightly.
[0066] Step S325: Use the adjusted normalized attention score matrix as the feature association weight matrix, where the matrix elements represent the association weights between the heat value feature sub-vector dimension and the running feature sub-vector dimension.
[0067] After adjusting the scaling factor, the normalized attention score matrix is used as the feature association weight matrix. The elements of the feature association weight matrix represent the association weights between the calorific value feature sub-vector dimension and the operational feature sub-vector dimension. These weight values reflect the actual degree of correlation between the two types of feature dimensions. By combining the feature association weight matrix with the calorific value feature sub-vector and the operational feature sub-vector, the two types of feature information can be integrated more rationally, enabling the fused features to more accurately reflect the actual situation of the waste incineration process.
[0068] Step S330: Input the heat value feature sub-vector and the running feature sub-vector into the feature fusion layer, and fuse them with the feature association weight matrix to generate a fused feature vector. The dimension of the fused feature vector is the same as the dimension of the state representation vector.
[0069] In the feature fusion layer, the calorific value feature vector and the operational feature vector are fused using a feature association weight matrix. Specifically, matrix operations are performed between the feature association weight matrix and the calorific value feature vector and the operational feature vector. Each element in the feature association weight matrix represents the association weight between the dimensions of the calorific value feature vector and the operational feature vector. Based on these weight values, the elements of the calorific value feature vector and the operational feature vector are summed using a weighted summation method. This weighted summation integrates the information from the calorific value feature vector and the operational feature vector to generate a fused feature vector. The dimension of the fused feature vector is the same as the dimension of the state representation vector. This ensures data dimensionality consistency in subsequent processing, facilitating unified analysis and processing.
[0070] Step S340: Input the fused feature vector into the temperature decision network of the main channel, perform feature mapping through multi-layer fully connected layers and residual connection structure, and after each fully connected layer, perform batch normalization processing and activation function to output the incineration temperature setpoint. The range of the incineration temperature setpoint is limited to the safe operating range by a preset temperature constraint function.
[0071] In one implementation, step S340 may include the following steps S341-S346: Step S341: Input the fused feature vector into the first fully connected layer of the main channel temperature decision network. The input dimension is the dimension of the fused feature vector, and the output dimension is the preset dimension of the first hidden layer. The first hidden layer features are generated by linearly transforming the weight matrix and the fused feature vector.
[0072] The first fully connected layer of the main channel temperature decision network is the entry layer of the entire network. Its main function is to perform preliminary feature transformation on the input fused feature vector. The input dimension is the dimension of the fused feature vector, which is to ensure the dimensionality consistency of the input data. The output dimension is the preset dimension of the first hidden layer, which is pre-set according to the network design and actual needs.
[0073] In the first fully connected layer, a linear transformation is performed between the weight matrix and the fused feature vector. The weight matrix is a pre-trained matrix containing the network's parameter information. Multiplying the weight matrix by the fused feature vector yields a new vector, which is the first hidden layer feature. The first hidden layer feature is the result of preliminary processing of the fused feature vector and contains higher-level feature information.
[0074] Step S342: Perform batch normalization on the first hidden layer features, calculate the mean and variance of the first hidden layer features for all samples in the batch, standardize the first hidden layer features to feature values with a mean of 0 and a variance of 1, and generate normalized first hidden layer features.
[0075] Batch normalization is used to accelerate the training process of neural networks. When performing batch normalization on the features of the first hidden layer, the mean and variance of the features of all samples in the batch are first calculated. The mean is the average of the features of all samples in the batch, and the variance measures how much these feature values deviate from the mean. Then, the features of the first hidden layer are standardized to feature values with a mean of 0 and a variance of 1. Specifically, for each element in the features of the first hidden layer, the mean of the batch is subtracted, and then divided by the standard deviation of the batch to obtain the standardized feature value. After batch normalization, the normalized features of the first hidden layer are generated.
[0076] Step S343: Input the normalized first hidden layer features into the activation function, use the LeakyReLU function for nonlinear mapping, retain the non-zero gradient of the negative half axis, and generate the activated first hidden layer features.
[0077] Activation functions are used to introduce non-linearity, enabling neural networks to learn more complex feature patterns. The LeakyReLU function is used to perform a non-linear mapping on the normalized first hidden layer features. The LeakyReLU function is an improved ReLU function that preserves non-zero gradients on the negative half-axis. For each element in the normalized first hidden layer features, if the element is greater than 0, the output of the LeakyReLU function is equal to that element; if the element is less than 0, the output of the LeakyReLU function is equal to that element multiplied by a small positive number (e.g., 0.01). This avoids the gradient vanishing problem on the negative half-axis, allowing the network to better learn feature information on the negative half-axis. After the non-linear mapping using the LeakyReLU function, the activated first hidden layer features are generated.
[0078] Step S344: Perform residual connection between the activated first hidden layer features and the fused feature vector, and directly superimpose the fused feature vector onto the activated first hidden layer features through skip connections to generate residual-enhanced first hidden layer features.
[0079] In this embodiment of the invention, the activated first hidden layer features are residually connected to the fused feature vector. Through skip connections, the fused feature vector is directly superimposed onto the activated first hidden layer features. Based on this, the network can more easily learn the differences between features. The fused feature vector contains the original feature information, while the activated first hidden layer features are feature information obtained after a series of processing steps. By adding the two through residual connections, residual-enhanced first hidden layer features are generated. These residual-enhanced first hidden layer features contain both the information of the original features and the processed feature information, enabling the network to better capture feature changes and differences.
[0080] Step S345: Repeat the above fully connected, batch normalization, activation and residual connection process, process the residual enhancement of the first hidden layer features in sequence, generate the second hidden layer features and the third hidden layer features, and reduce the output dimension of each layer according to a preset ratio.
[0081] To further extract and transform feature information, the fully connected, batch normalization, activation, and residual connection processes need to be repeated. Starting with residual enhancement of the first hidden layer features, the processes are performed sequentially to generate the second and third hidden layer features.
[0082] In each layer's processing, fully connected layers perform linear transformations and feature extraction on the input features; batch normalization normalizes the features, making the data distribution more stable; activation functions introduce non-linear factors, enabling the network to learn more complex feature patterns; and residual connections solve the vanishing gradient problem, improving the network's training efficiency and performance. The output dimension of each layer decreases by a preset ratio. This is to gradually reduce the dimensionality of the features, allowing the network to focus on more important feature information.
[0083] Step S346: Input the third hidden layer features into the fully connected layer of the output layer. The input dimension is the dimension of the third hidden layer features, and the output dimension is 1. Generate the original temperature setpoint through linear transformation. Substitute the original temperature setpoint into the temperature constraint function and output the incineration temperature setpoint that meets the safe operating range.
[0084] The fully connected output layer is the last layer of the temperature decision network. Its main task is to output the final incineration temperature setpoint based on the input features from the third hidden layer. The input dimension is the same as the dimension of the third hidden layer features, and the output dimension is 1 because a single incineration temperature setpoint is required as the final output. The original temperature setpoint is generated by multiplying the third hidden layer features by the weight matrix of the fully connected output layer through a linear transformation, plus a bias term. This original temperature setpoint is an unconstrained temperature value and may exceed the safe operating range.
[0085] To ensure that the output incineration temperature setpoint is within the safe operating range, the original temperature setpoint is substituted into a preset temperature constraint function. The temperature constraint function adjusts the original temperature setpoint according to the incinerator's safe operating requirements. If the original temperature setpoint is higher than the upper limit of the safe operating range, it is adjusted to the upper limit value; if the original temperature setpoint is lower than the lower limit of the safe operating range, it is adjusted to the lower limit value. After processing by the temperature constraint function, the output incineration temperature setpoint conforms to the safe operating range.
[0086] Step S350: Input the fused feature vector into the oxygen supply decision network of the sub-channel, and extract temporal features through a convolution-recurrent hybrid structure. First, extract local features through a one-dimensional convolutional processing layer, and then model the temporal dependency through a long short-term memory network layer to output the oxygen supply adjustment value. The change range of the oxygen supply adjustment value is limited to a preset single adjustment threshold through a gradient clipping mechanism.
[0087] The oxygen supply decision network is a core module in the subchannel of the reinforcement learning policy network. Its main task is to output an appropriate oxygen supply adjustment value based on the input fused feature vector. The convolutional-recurrent hybrid structure is the main architecture of the oxygen supply decision network, combining the advantages of convolutional neural networks and recurrent neural networks, and can effectively extract temporal features.
[0088] The one-dimensional convolutional processing layer, the first part of the convolution-recurrent hybrid structure, is used to extract local features from the fused feature vector. The one-dimensional convolutional processing layer slides a convolution kernel across the fused feature vector, performing convolution operations on features in local regions to extract local feature information. These local features reflect local patterns and feature variations in the fused feature vector at different locations. The Long Short-Term Memory (LSTM) network layer, the second part of the convolution-recurrent hybrid structure, is used to model temporal dependencies. LSTM is a special type of recurrent neural network capable of effectively handling long-term dependencies in sequential data. In the LSTM layer, through a gating mechanism, the network can selectively remember and forget past information, thereby better capturing the temporal information in the fused feature vector.
[0089] The variation in the output oxygen supply adjustment value is limited to a preset single-adjustment threshold using a gradient clipping mechanism. Gradient clipping is a technique used to prevent gradient explosion. During training, if the gradient variation exceeds the preset single-adjustment threshold, the gradient is clipped to ensure its variation does not exceed that threshold. This ensures that the oxygen supply adjustment value does not change too drastically, avoiding instability to the incinerator's operation.
[0090] In one implementation, step S350 may specifically include the following steps S351-S356: Step S351: Reconstruct the dimensions of the fused feature vector and convert it into a one-dimensional feature sequence. The sequence length is the dimension of the fused feature vector, and each time step corresponds to a dimension value of the fused feature vector.
[0091] To facilitate subsequent processing by the one-dimensional convolutional layers, the fused feature vector needs to be reconstructed in terms of dimensions, transforming it into a one-dimensional feature sequence. The fused feature vector is a multi-dimensional vector containing feature information across multiple dimensions. Dimensional reconstruction unfolds this dimensional information into a one-dimensional sequence. The sequence length is equal to the dimension of the fused feature vector; that is, each element in the one-dimensional feature sequence corresponds to a dimensional value of the fused feature vector. Each time step corresponds to a dimensional value of the fused feature vector. Based on this, the fused feature vector is converted into time-series data, facilitating subsequent convolutional and temporal modeling processing.
[0092] Step S352: Input the one-dimensional feature sequence into the one-dimensional convolutional processing layer of the secondary channel oxygen supply decision network. The one-dimensional convolutional processing layer uses a preset number of convolutional kernels. The size of each convolutional kernel is a preset window size. Local feature extraction is performed on the one-dimensional feature sequence through a sliding window to generate a multi-channel convolutional feature map. The number of channels of the multi-channel convolutional feature map is equal to the number of convolutional kernels, and the length is the length of the one-dimensional feature sequence minus the window size plus 1.
[0093] One-dimensional convolutional processing layers are an important component of the oxygen supply decision network, used to extract local features from a one-dimensional feature sequence. These layers employ a predetermined number of convolutional kernels, each with a predefined window size. The kernels are small weight matrices that slide across the one-dimensional feature sequence, performing convolution operations on features in local regions.
[0094] By using a sliding window approach, the convolutional kernel moves segment by segment along the one-dimensional feature sequence, taking a step at a time. At each window position, the kernel performs a convolution operation with the features within that window, yielding a convolution result. Combining the convolution results from all window positions generates a multi-channel convolutional feature map. The number of channels in the multi-channel convolutional feature map equals the number of convolutional kernels, with each channel corresponding to the output of one kernel. The length of the multi-channel convolutional feature map is the length of the one-dimensional feature sequence minus the window size plus one, because the sliding of the convolutional kernel during the convolution operation reduces the length of the output.
[0095] Step S353: Perform time-dimension max pooling on the multi-channel convolutional feature map, with the pooling window size set to a preset value and the stride equal to the pooling window size. Take the maximum value of the local features of each channel to generate the dimensionality-reduced convolutional feature sequence.
[0096] Max pooling is a technique used to reduce data dimensionality and extract important features. In this embodiment of the invention, max pooling is performed on the multi-channel convolutional feature map along the temporal dimension. The pooling window size is a preset value, and the stride is equal to the pooling window size.
[0097] In the max pooling process, the pooling window slides along the time dimension of the multi-channel convolutional feature map. For each local feature within the pooling window, the maximum value is taken as the output of that pooling window. In this way, the local features of each channel are dimensionality-reduced, generating a dimensionality-reduced convolutional feature sequence. This dimensionality-reduced convolutional feature sequence retains the feature information from the multi-channel convolutional feature map while reducing the data dimensionality. This reduces the computational complexity of subsequent processing and improves the training efficiency and performance of the model.
[0098] Step S354: Input the dimensionality-reduced convolutional feature sequence into the Long Short-Term Memory (LSTM) network layer. The LSM network layer contains a preset number of hidden units. It learns the long-term dependencies of the feature sequence through a gating mechanism, outputs the hidden state vector at each time step, and uses the hidden state vector at the last time step as the output of the LSM network layer.
[0099] Long Short-Term Memory (LSTM) network layers are used to model the temporal dependencies in the dimensionality-reduced convolutional feature sequences. LSM network layers contain a preset number of hidden units, which control the flow of information through a gating mechanism.
[0100] Gating mechanisms are a core feature of Long Short-Term Memory (LSTM) networks, including input gates, forget gates, and output gates. Input gates control the input of new information; forget gates control the forgetting of past information; and output gates control the output of the current hidden state. Through these gating mechanisms, LTM networks can selectively remember and forget past information, thereby better capturing long-term dependencies in the dimensionality-reduced convolutional feature sequences.
[0101] In the Long Short-Term Memory (LSTM) network layer, a hidden state vector is output at each time step. This hidden state vector contains information from the current time step and relevant information from past time steps. The hidden state vector from the last time step is used as the output of the LSM layer. This output vector contains temporal information and long-term dependencies from the dimensionality-reduced convolutional feature sequence, providing crucial feature representations for subsequent processing.
[0102] Step S355: Input the hidden state vector output by the Long Short-Term Memory network layer into the fully connected layer, and map the hidden state vector into a single value through linear transformation, which serves as the original oxygen supply adjustment value.
[0103] The fully connected layer is the final part of the oxygen supply decision network, used to map the hidden state vector output by the Long Short-Term Memory (LSTM) network layer to a single numerical value. The fully connected layer performs a linear transformation, multiplying the hidden state vector by the weight matrix of the fully connected layer and adding a bias term to obtain a single numerical value.
[0104] This single value is the original oxygen supply adjustment value. The original oxygen supply adjustment value is the result obtained after a series of processing steps, such as one-dimensional convolutional processing layers and long short-term memory network layers, based on the information in the fused feature vector. It reflects the magnitude of the oxygen supply adjustment that needs to be made at the current moment.
[0105] Step S356: Substitute the original oxygen supply adjustment value into the adjustment threshold constraint function. If the absolute value of the adjustment value exceeds the preset single adjustment threshold, then truncate it to the threshold boundary value; otherwise, keep the original value and generate an oxygen supply adjustment value that meets the adjustment range limit.
[0106] To ensure that the variation in the oxygen supply adjustment value remains within a safe range, the original oxygen supply adjustment value is substituted into a preset adjustment threshold constraint function. This adjustment threshold constraint function is a predefined function that limits the variation in the oxygen supply adjustment value based on the incinerator's operational requirements.
[0107] If the absolute value of the original oxygen supply adjustment exceeds the preset single adjustment threshold, it is truncated to the threshold boundary value. That is, if the original oxygen supply adjustment is greater than the upper threshold, it is adjusted to the upper threshold value; if the original oxygen supply adjustment is less than the lower threshold, it is adjusted to the lower threshold value. If the absolute value of the original oxygen supply adjustment does not exceed the preset single adjustment threshold, the original value remains unchanged.
[0108] After adjusting the threshold constraint function, an oxygen supply adjustment value that meets the adjustment range limit is generated. This ensures that the oxygen supply adjustment is not too drastic, avoiding instability to the incinerator operation.
[0109] Step S360: Combine the combustion temperature setpoint output by the main channel and the oxygen supply adjustment value output by the secondary channel in a one-to-one correspondence according to the timestamp order to generate a control action sequence containing timestamps. Each element of the control action sequence contains the combustion temperature setpoint, oxygen supply adjustment value and timestamp information at the corresponding time.
[0110] The combustion temperature setpoint output from the main channel and the oxygen supply adjustment value output from the secondary channel are matched one-to-one according to the timestamp sequence in order to ensure that the combustion temperature setpoint and the oxygen supply adjustment value at each time point can be accurately matched.
[0111] Each element of the control action sequence includes the incineration temperature setpoint, oxygen supply adjustment value, and timestamp information for the corresponding moment. This structure clearly records the control action information at each point in time, facilitating real-time adjustments and control by the waste incinerator's control system. In practical applications, the control action sequence is sent to the waste incinerator's control system, which automatically adjusts the incinerator's combustion parameters based on the information in the sequence to achieve optimized incineration results.
[0112] Step S400: Based on the control action sequence, the performance of the incineration process is evaluated by calculating the reward function. The reward function calculates the weighted sum of the thermal efficiency optimization index, the pollutant emission control index, and the equipment life maintenance index to generate a multi-objective reward value.
[0113] Reward functions are tools used to evaluate the performance of incineration processes. By comprehensively considering multiple indicators, they quantify the effectiveness of control action sequences. Thermal efficiency optimization indicators reflect the energy utilization efficiency during incineration, measuring the ratio of energy generated by waste incineration to input energy. Improving thermal efficiency can reduce energy waste and lower operating costs.
[0114] Pollutant emission control indicators reflect the environmental impact of the incineration process, focusing on the emission of various pollutants generated during incineration. Strict control of pollutant emissions can reduce environmental pollution and meet environmental protection requirements.
[0115] Equipment life maintenance indicators take into account the long-term operation of the incinerator equipment and reflect the impact of control action sequences on equipment life. Reasonable control actions can reduce equipment wear and damage, and extend the equipment's service life.
[0116] The reward function generates a multi-objective reward value by calculating a weighted sum of thermal efficiency optimization indicators, pollutant emission control indicators, and equipment life maintenance indicators. The weighting method can be adjusted based on the relative importance of each indicator. For example, if environmental requirements are prioritized, the weight of pollutant emission control indicators can be appropriately increased; if the long-term operating costs of the equipment are of greater concern, the weight of equipment life maintenance indicators can be appropriately increased.
[0117] Multi-objective reward values are comprehensive evaluation metrics that reflect the combined effect of a control action sequence across multiple aspects. By calculating the reward function, different control action sequences can be compared and evaluated, providing a basis for subsequent policy optimization.
[0118] In one implementation, step S400 may specifically include the following steps S410-S460: Step S410: According to the timestamp sequence of the control action sequence, each action element is sent to the incinerator simulation model in sequence. The incinerator simulation model is constructed through physical field modeling. The input is the incineration temperature setpoint and the oxygen supply adjustment value. The output is the thermal efficiency related data, pollutant emission data and equipment operation status data within the preset time.
[0119] An incinerator simulation model is a model used to simulate the actual operation of an incinerator. It is constructed through physical field modeling. Physical field modeling is based on physical principles and mathematical models to describe and simulate physical processes such as heat transfer, combustion reaction, and fluid flow within the incinerator.
[0120] Following the timestamp sequence of the control actions, each action element is sent sequentially to the incinerator simulation model. Each action element includes the corresponding incineration temperature setpoint and oxygen supply adjustment value. These setpoints and adjustments are used as inputs to the incinerator simulation model, which, based on the principles of physics modeling, simulates the incinerator's operation under these input conditions. The outputs of the incinerator simulation model are thermal efficiency-related data, pollutant emission data, and equipment operating status data for a preset time period. Thermal efficiency-related data includes information such as the incinerator's input and output heat, used to calculate thermal efficiency optimization indicators. Pollutant emission data includes information such as the concentration and emission rate of various pollutants generated during incineration, used to calculate pollutant emission control indicators. Equipment operating status data reflects the state of the incinerator equipment during operation, such as the temperature and vibration of key components, used to calculate equipment lifespan maintenance indicators.
[0121] Step S420: Calculate the thermal efficiency optimization index based on thermal efficiency related data. The thermal efficiency related data includes the input heat measurement value and output heat measurement value of the incinerator, which are calculated by comparing the input-output heat ratio and the historical best thermal efficiency.
[0122] Thermal efficiency optimization is an indicator that measures the energy utilization efficiency of the incineration process. Relevant thermal efficiency data includes the input and output heat measurements of the incinerator. Input heat measurement refers to the total energy input into the incinerator during the incineration process, primarily from waste combustion and other auxiliary energy sources. Output heat measurement refers to the actual useful energy output by the incinerator during operation, such as the heat generated from the steam.
[0123] In one implementation, step S420 may specifically include the following steps S421-S425: Step S421: Extract the input heat measurement value and the output heat measurement value from the thermal efficiency related data. The input heat measurement value is the total heat released by the waste combustion, which is calculated by multiplying the total mass of waste by the calorific value per unit mass. The output heat measurement value is the heat of steam generated by the incinerator, which is calculated by the difference between the steam flow rate, the steam enthalpy, and the inlet water enthalpy.
[0124] In thermal efficiency data, the input heat measurement is a key metric for measuring the energy provided by waste combustion. The total heat released by waste combustion depends on the total mass of the waste and its calorific value per unit mass. The calorific value per unit mass refers to the heat released when a unit mass of waste is completely burned, and it can be calculated using the previously described preset calorific value estimation model based on measurements of the waste's moisture content and combustible material content. Multiplying the total mass of waste by the calorific value per unit mass yields the input heat measurement.
[0125] The output heat measurement mainly reflects the heat of steam generated by the incinerator. The calculation of steam heat requires consideration of steam flow rate, steam enthalpy, and inlet water enthalpy. Steam flow rate can be measured in real time using a flow sensor installed on the steam pipeline, reflecting the amount of steam passing through the pipeline per unit time. Steam enthalpy is a thermodynamic state parameter of steam, representing the energy possessed by a unit mass of steam, and can be determined using parameters such as steam temperature and pressure, employing a steam enthalpy table or relevant calculation software. Inlet water enthalpy is the enthalpy of the water entering the incinerator, and can also be calculated based on parameters such as inlet water temperature and pressure. The output heat measurement value can be obtained by multiplying the steam flow rate by the difference between the steam enthalpy and the inlet water enthalpy. For example, at a certain moment, if the measured steam flow rate is a fixed value, and the enthalpy of the steam and the inlet water are determined at that moment, multiplying the steam flow rate by the difference between the two enthalpy values yields the steam heat generated by the incinerator at that moment, i.e., the output heat measurement value.
[0126] Step S422: Calculate the ratio of the output heat measurement value to the input heat measurement value to obtain the actual thermal efficiency value. The actual thermal efficiency value reflects the energy conversion efficiency of the incinerator under the current control action.
[0127] The actual thermal efficiency value is obtained by dividing the output heat measurement value calculated in step S421 by the input heat measurement value. This ratio directly reflects the efficiency with which the incinerator converts input energy into useful output energy under the current control action. The higher the actual thermal efficiency value, the more effectively the incinerator can convert the energy released from waste combustion into useful energy such as steam under the current control action, and the better the energy conversion efficiency. Conversely, the lower the actual thermal efficiency value, the more energy loss occurs during the energy conversion process, and the incinerator's operating efficiency needs to be improved.
[0128] Step S423: Retrieve the highest thermal efficiency record value under the same operating condition from the historical database as the thermal efficiency benchmark value. The same operating condition is determined by matching the waste composition parameters, incineration load parameters and ambient temperature parameters.
[0129] The historical database stores operating data of the incinerator under different operating conditions, including thermal efficiency values. To accurately evaluate the thermal efficiency optimization effect of current control actions, it is necessary to retrieve the highest thermal efficiency record value under the same operating condition as the thermal efficiency benchmark. Determining the same operating condition requires comprehensive consideration of waste composition parameters, incineration load parameters, and ambient temperature parameters. Waste composition parameters, such as the moisture content and combustible material content mentioned earlier, result in different combustion characteristics and heat release, thus affecting thermal efficiency. The incineration load parameter reflects the amount of waste processed by the incinerator per unit time; excessively high or low loads can both affect thermal efficiency. Ambient temperature parameters also have a certain impact on the incineration process; for example, at lower ambient temperatures, the incinerator needs to consume more energy to preheat the waste and air.
[0130] By accurately measuring and recording the waste composition parameters, incineration load parameters, and ambient temperature parameters under the current operating conditions, and then matching and searching in the historical database, the highest thermal efficiency record value under the same operating conditions is found and used as the thermal efficiency benchmark value. For example, if the current incinerator is processing waste with a moisture content of X, a combustible content of Y, an incineration load of Z, and an ambient temperature of T, the historical database is searched for operating condition records with the same or similar parameter values, and the record value with the highest thermal efficiency is selected as the thermal efficiency benchmark value.
[0131] Step S424: Calculate the ratio of the actual thermal efficiency value to the thermal efficiency benchmark value to obtain the relative thermal efficiency value. When the relative thermal efficiency value is greater than 1, it means that the actual thermal efficiency exceeds the historical best level, and when it is less than 1, it means that the historical best level has not been reached.
[0132] Dividing the actual thermal efficiency value calculated in step S422 by the thermal efficiency benchmark value determined in step S423 yields the relative thermal efficiency value. The relative thermal efficiency value is an important reference indicator, clearly reflecting the relationship between the current thermal efficiency and the historical best thermal efficiency. When the relative thermal efficiency value is greater than 1, it indicates that the actual thermal efficiency under the current control action has exceeded the highest thermal efficiency level under the same historical operating conditions. This suggests that the current control strategy has achieved good results in thermal efficiency optimization, possibly due to reasonable control of the incineration temperature setpoint and oxygen supply adjustment value, resulting in more complete waste combustion and more efficient energy conversion. When the relative thermal efficiency value is less than 1, it means that the current actual thermal efficiency has not reached the historical best level, requiring further analysis of the reasons and adjustment of the control actions. This may be due to factors such as changes in waste composition, aging of the incinerator equipment, or unreasonable control parameter settings leading to a decrease in thermal efficiency.
[0133] Step S425: Multiply the relative value of thermal efficiency by a preset thermal efficiency benchmark coefficient. The thermal efficiency benchmark coefficient is set according to the design thermal efficiency of the incinerator. The higher the design thermal efficiency, the larger the benchmark coefficient. Use the product result as the thermal efficiency optimization index. The larger the thermal efficiency optimization index value, the better the optimization effect of the current control action on thermal efficiency.
[0134] The preset thermal efficiency benchmark coefficient is set based on the incinerator's design thermal efficiency. Design thermal efficiency reflects the incinerator's energy conversion capability under ideal conditions. Different incinerators have different design thermal efficiencies due to differences in design structure, technology, and other factors. An incinerator with a higher design thermal efficiency should be able to achieve a higher thermal efficiency level during normal operation, and therefore, its corresponding thermal efficiency benchmark coefficient will be larger.
[0135] The relative thermal efficiency value calculated in step S424 is multiplied by a preset thermal efficiency benchmark coefficient, and the resulting product is the thermal efficiency optimization index. The thermal efficiency optimization index comprehensively considers the comparison between the current thermal efficiency and the historical best thermal efficiency, as well as the design thermal efficiency of the incinerator. A larger thermal efficiency optimization index value indicates a better effect of the current control actions in improving thermal efficiency, and a more complete utilization of the incinerator's energy conversion potential.
[0136] Step S430: Calculate pollutant emission control indicators based on pollutant emission data. The pollutant emission data includes the concentration measurements of various pollutants. The emission compliance rate is calculated by comparing the concentrations of each pollutant with the emission standards, and then obtained by weighted averaging.
[0137] Pollutant emission control targets are indicators used to measure the degree of environmental pollution caused by the incineration process. Pollutant emission data includes measured concentrations of various pollutants generated during incineration, such as sulfur dioxide, nitrogen oxides, and particulate matter. The emission of these pollutants harms the atmospheric environment and human health, therefore, their emissions must be strictly controlled.
[0138] The emission compliance rate is calculated by comparing the measured concentrations of each pollutant with the emission standard concentrations specified in environmental protection standards. For each pollutant, if its measured concentration is less than or equal to the emission standard concentration, the emission compliance rate is 100%; if the measured concentration exceeds the emission standard concentration, the emission compliance rate is calculated based on the degree of exceedance. For example, if the emission standard concentration of a pollutant is A and the actual measured concentration is B, if B ≤ A, the emission compliance rate is 1; if B > A, the emission compliance rate = A / B. After obtaining the emission compliance rate for each pollutant, a weighted average is required. Different pollutants pose different degrees of harm to the environment and human health, therefore different weights are assigned to each pollutant. The emission control weights for each pollutant are retrieved from a dynamic weight library; pollutants with higher levels of harm have larger weights. The emission compliance rate for each pollutant is multiplied by its corresponding weight, and then all results are summed to obtain the pollutant emission control index. The larger the pollutant emission control index value, the better the current control action is on pollutant emissions, and the lower the degree of environmental pollution caused by the incineration process.
[0139] In one implementation, step S430 may specifically include the following steps S431-S436: Step S431: Extract concentration measurements of multiple pollutants from pollutant emission data. The concentration of each pollutant includes continuous measurements within a preset time period. Calculate the average concentration of each pollutant as the representative concentration of that pollutant.
[0140] Pollutant emission data is collected in real time by pollutant monitoring devices installed on the flue gas emission ducts of the incinerator. These devices can continuously measure the concentration of multiple pollutants and record continuous measurements over a preset time period. For example, for pollutants such as sulfur dioxide and nitrogen oxides, their concentrations are measured at regular intervals (e.g., every minute) within an hour, resulting in a series of measurements. To more accurately assess the emissions of each pollutant, the average concentration of each pollutant needs to be calculated. The average concentration of each pollutant is obtained by summing all continuous measurements for each pollutant over the preset time period and then dividing by the number of measurements. This average concentration is used as the representative concentration of that pollutant.
[0141] Step S432: Obtain the emission standard concentrations of each pollutant specified in the environmental protection standards as the benchmark limits for the corresponding pollutants.
[0142] Environmental standards are mandatory standards established to protect the environment and human health. They clearly define the maximum permissible concentrations of various pollutants in flue gas emitted from incinerators. These standards are formulated based on scientific research and environmental quality requirements, and are authoritative and instructive. The emission standard concentrations of each pollutant are obtained from environmental standard documents or relevant databases and used as the baseline limits for the corresponding pollutants. These baseline limits are the basis for determining whether pollutant emissions meet the standards.
[0143] Step S433: For each pollutant, calculate the ratio of its representative concentration to the benchmark limit to obtain the pollution exceedance multiple. If the representative concentration is less than or equal to the benchmark limit, the pollution exceedance multiple is 0; otherwise, it is the ratio of the difference between the representative concentration and the benchmark limit to the benchmark limit.
[0144] For each pollutant, the representative concentration calculated in step S431 is compared and calculated with the baseline limit determined in step S432. If the representative concentration is less than or equal to the baseline limit, the emission of the pollutant complies with the standard, and the pollution exceedance multiple is 0. For example, if the representative concentration of a pollutant is D1 and the baseline limit is C1, and D1 ≤ C1, then the pollution exceedance multiple = 0. If the representative concentration is greater than the baseline limit, the pollution exceedance multiple needs to be calculated. The pollution exceedance multiple reflects the degree to which the pollutant emission exceeds the standard, and its calculation formula is: Pollution exceedance multiple = (Representative concentration - Baseline limit) / Baseline limit. For example, if the representative concentration of a pollutant is D2 and the baseline limit is C2, and D2 > C2, then the pollution exceedance multiple = (D2 - C2) / C2.
[0145] Step S434: Calculate the emission compliance rate based on the pollution exceedance multiple. The formula for calculating the emission compliance rate is 1 divided by 1 plus the sum of the pollution exceedance multiples. Ensure that the compliance rate is 1 when the pollution exceedance multiple is 0, and the larger the exceedance multiple, the smaller the compliance rate.
[0146] Calculate the emission compliance rate based on the pollution exceedance multiple calculated in step S433. The formula for calculating the emission compliance rate is, for example: Emission compliance rate = 1 / (1 + pollution exceedance multiple). When the pollution exceedance multiple is 0, it means the concentration is less than or equal to the benchmark limit, and the emission compliance rate = 1 / (1+0) = 1, which is 100%, indicating that the emission of the pollutant fully complies with the standard.
[0147] As the exceedance factor of pollution increases, the compliance rate gradually decreases. For example, if the exceedance factor is 1, the compliance rate is 1 / (1+1) = 0.5, or 50%, meaning that only half of the pollutant emissions meet the standards. This calculation method accurately quantifies the compliance status of each pollutant's emissions, intuitively reflecting the degree of environmental impact.
[0148] Step S435: Retrieve the emission control weights of each pollutant from the dynamic weight library. The weights of different pollutants are set according to their degree of environmental hazard. The higher the degree of hazard, the greater the weight.
[0149] The dynamic weight database is a database that stores the emission control weights for various pollutants. These weights are set based on the degree of harm that different pollutants pose to the environment and human health. Different pollutants have different chemical properties and toxicity, and their impacts on the atmospheric environment, soil, water bodies, and human health vary. For example, sulfur dioxide contributes to acid rain, acidifies soil and water bodies, and affects the ecological balance; nitrogen oxides participate in the formation of photochemical smog, damaging the human respiratory system. The higher the degree of harm of a pollutant, the greater its emission control weight. When calculating pollutant emission control targets, the emission control weights for each pollutant need to be retrieved from the dynamic weight database. The dynamic weight database is updated and adjusted in real time based on current environmental conditions, policy requirements, and other factors to ensure the rationality and accuracy of the weights.
[0150] Step S436: Multiply the emission compliance rate of each pollutant by its corresponding weight, and sum them to obtain the pollutant emission control index. The larger the emission control index value, the better the control effect of the current control action on pollutant emissions.
[0151] Multiply the emission compliance rate of each pollutant calculated in step S434 by the corresponding emission control weight retrieved in step S435, and then add up the product results of all pollutants. The sum obtained is the pollutant emission control index.
[0152] Pollutant emission control targets comprehensively consider the compliance status of various pollutant emissions with standards and their degree of environmental harm. A higher emission control target value indicates that, under the current control measures, the overall emissions of various pollutants during incineration are more in line with standards, resulting in lower environmental pollution; in other words, the current control measures are more effective in controlling pollutant emissions. Pollutant emission control targets can be used to evaluate the effectiveness of different control measures in reducing pollutant emissions, providing a reference for optimizing control strategies.
[0153] Step S440: Calculate equipment life maintenance indicators based on equipment operating status data. The equipment operating status data includes temperature measurement values, vibration amplitude measurement values, and cumulative operating time values of key components. These are converted into life loss rates through the equipment aging model and then normalized.
[0154] Equipment life maintenance indicators are metrics that measure the impact of current control actions on the lifespan of the incinerator equipment. Equipment operating status data includes measurements of temperature, vibration amplitude, and cumulative operating time for key incinerator components. These data reflect the equipment's condition and performance during operation.
[0155] Temperature measurements reflect the thermal state of critical components during operation. Excessively high temperatures accelerate component aging and damage, reducing equipment lifespan. For example, prolonged operation of incinerator components such as the grate and combustion chamber in high-temperature environments may lead to material degradation, deformation, and cracking. Vibration amplitude measurements reflect the equipment's stability during operation. Excessive vibration can cause fatigue damage to the equipment's structure, affecting its normal operation and lifespan. Cumulative operating time records the total operating time from commissioning to the present moment; as operating time increases, the degree of wear and aging of the equipment gradually intensifies.
[0156] Equipment aging models convert equipment operating status data into lifespan attrition rates. These models are multi-input, single-output models trained on a large amount of historical equipment failure data. The model comprehensively considers the impact of factors such as temperature, vibration, and operating time on equipment lifespan. Inputs include statistical characteristics of temperature, vibration, and operating time; the output is the overall equipment lifespan attrition rate. After obtaining the lifespan attrition rate, normalization processing is required to eliminate differences in data dimensions and ranges, ensuring comparability of lifespan attrition rates for different equipment or different time periods.
[0157] In one implementation, step S440 may specifically include the following steps S441-S445: Step S441: Extract the temperature measurement value, vibration amplitude measurement value and cumulative running time value of key components from the equipment operation status data. Each parameter includes the maximum value, minimum value and average value within a preset time period, and construct a set of parameter statistical features.
[0158] The system extracts temperature measurements, vibration amplitude measurements, and cumulative operating time values of key components from equipment operating status data. For each parameter, not only is its current value recorded, but the maximum, minimum, and average values within a preset time period are also calculated. For example, for temperature measurements, measurements are taken at regular intervals within an hour to obtain a series of temperature data; the maximum and minimum values are then identified, and the average value is calculated. Similarly, vibration amplitude measurements and cumulative operating time values undergo similar statistical processing.
[0159] The maximum, minimum, and average values are combined to construct a set of parameter statistical features. This set comprehensively reflects the changes in the operating status of key components over a preset time period. For example, the maximum temperature value reflects the thermal load of the equipment under extreme conditions, the minimum value reflects the performance of the equipment in low-temperature environments, and the average value reflects the overall thermal state of the equipment. By constructing this set of parameter statistical features, richer and more accurate input data can be provided for subsequent equipment aging models.
[0160] Step S442: Input the set of parameter statistical features into the equipment aging model. The equipment aging model is a multi-input single-output model, which is trained by historical equipment failure data. The inputs are temperature statistical features, vibration statistical features, and operating time statistical features, and the output is the comprehensive life loss rate of the equipment.
[0161] The calculation process of the equipment aging model may specifically include: comparing the maximum value of the temperature measurement with the safe temperature threshold to calculate the proportion of time the temperature exceeds the limit; comparing the average value of the vibration amplitude measurement with the fatigue limit to calculate stress cycle damage; and comparing the cumulative value of the running time with the rated life time to calculate the life consumption ratio.
[0162] The equipment aging model is a multi-input single-output model trained based on historical equipment failure data. Taking the set of parameter statistical features constructed in step S441 as input, the model calculates the overall equipment lifespan loss rate according to its internal algorithm and parameters.
[0163] In the calculation of the equipment aging model, the maximum value of the measured temperature is first compared with the safe temperature threshold. The safe temperature threshold is the highest temperature that the equipment can withstand during normal operation; exceeding this threshold will damage the equipment. The proportion of time the maximum measured temperature exceeds the safe temperature threshold is calculated out of a preset time, i.e., the temperature over-limit time proportion. The larger the temperature over-limit time proportion, the longer the equipment has been operating in a high-temperature environment, and the greater the impact on the equipment's aging.
[0164] Next, the average value of the vibration amplitude measurements is compared with the fatigue limit. The fatigue limit is the maximum vibration amplitude that the equipment can withstand under long-term vibration; exceeding this limit will lead to fatigue cracks and damage to the equipment components. Stress cyclic damage is calculated based on the difference between the average vibration amplitude measurements and the fatigue limit. Stress cyclic damage reflects the cumulative damage caused by vibration to the equipment structure. Finally, the cumulative operating time is compared with the rated lifespan. The rated lifespan is the normal operating time specified in the equipment's design. Dividing the cumulative operating time by the rated lifespan yields the lifespan consumption ratio.
[0165] The overall equipment lifespan loss rate is calculated by comprehensively considering the proportion of time spent exceeding temperature limits, stress cycle damage, and lifespan loss according to certain weights. The overall equipment lifespan loss rate ranges from 0 to 1, where 0 indicates that the equipment is undamaged and in brand-new condition, and 1 indicates that the equipment is completely damaged and cannot operate normally.
[0166] Step S443: Multiply the proportion of temperature over-limit duration, stress cycle damage, and life consumption by the preset damage weights respectively, and sum them to obtain the overall life loss rate of the equipment. The life loss rate ranges from 0 to 1, where 0 represents no loss and 1 represents complete loss.
[0167] In step S442, the proportion of time exceeding temperature limits, stress cycle damage, and lifespan consumption were obtained. To comprehensively consider the impact of these three factors on equipment lifespan, preset damage weights need to be assigned to each of them. These damage weights are determined based on the degree of influence of temperature, vibration, and operating time on equipment aging. Different equipment has different damage weights due to differences in structure and working principle.
[0168] Multiply the proportion of time the temperature exceeds the limit by the corresponding damage weight, multiply the stress cycle damage by its corresponding damage weight, and multiply the life consumption proportion by its corresponding damage weight. Then add these three products together to get the total sum, which is the overall life loss rate of the equipment.
[0169] Step S444: Smooth the overall life loss rate of the equipment. Use a moving average algorithm to eliminate instantaneous fluctuations and generate a smoothed life loss rate. The sliding window size is a preset time length.
[0170] The overall equipment lifespan depreciation rate may fluctuate during calculation due to instantaneous factors. To more accurately reflect the aging trend of equipment, the overall equipment lifespan depreciation rate needs to be smoothed. By continuously moving the sliding window, the entire series of overall equipment lifespan depreciation rates is processed to eliminate the impact of instantaneous fluctuations, generating a smoothed lifespan depreciation rate. The smoothed lifespan depreciation rate can more clearly show the long-term trend of equipment aging, providing a more reliable basis for equipment maintenance and management.
[0171] Step S445: Subtract the smoothed life loss rate from 1 as the equipment life maintenance index. The larger the equipment life maintenance index value, the better the maintenance effect of the current control action on the equipment life.
[0172] Subtracting the smoothed life loss rate obtained in step S444 from 1, the difference is the equipment life maintenance index. The equipment life maintenance index reflects the effectiveness of current control actions in maintaining the equipment's lifespan. The overall equipment life loss rate represents the proportion of the equipment's lifespan already lost; subtracting this proportion from 1 yields the equipment life maintenance index, which represents the proportion of the equipment's remaining maintainable lifespan. A higher equipment life maintenance index value indicates that current control actions are better able to reduce equipment aging and wear, extending the equipment's service life.
[0173] Step S450: Obtain the weights of thermal efficiency optimization index, pollutant emission control index, and equipment life maintenance index through the dynamic weight library of the reinforcement learning environment. The dynamic weight library automatically adjusts the weight values according to the current operating conditions of the incineration process, and increases the weight of emission control index when the operating conditions are abnormal.
[0174] The dynamic weight library of the reinforcement learning environment is a database used to store and manage the weights of indicators such as thermal efficiency optimization, pollutant emission control, and equipment life maintenance. These weights determine the relative importance of each indicator in the overall evaluation when calculating the multi-objective reward value.
[0175] The dynamic weight library automatically adjusts weight values based on the current operating conditions of the incineration process. These current conditions include factors such as waste composition, incineration load, and ambient temperature. Different conditions have varying impacts on the thermal efficiency, pollutant emissions, and equipment lifespan of the incineration process. For example, when the waste has a high moisture content, it may be necessary to focus more on improving thermal efficiency, thus appropriately increasing the weight of the thermal efficiency optimization index. Conversely, when the ambient temperature is low, it may be necessary to strengthen equipment protection, thus increasing the weight of the equipment lifespan maintenance index.
[0176] In abnormal operating conditions, such as a sudden increase in pollutant emissions, the dynamic weight library will increase the weight of emission control indicators. This is to emphasize the control of pollutant emissions, ensure that the incineration process meets environmental protection requirements, and reduce environmental pollution. By dynamically adjusting the weight values, the multi-objective reward values can more accurately reflect the actual situation of the current incineration process, providing more reasonable feedback for the training of the reinforcement learning strategy network.
[0177] Step S460: Multiply the thermal efficiency optimization index by its weight, the pollutant emission control index by its weight, and the equipment life maintenance index by its weight, and add the three together to obtain the multi-objective reward value. The range of the multi-objective reward value is limited to a preset range through standardization.
[0178] Step S500: Based on the multi-objective reward value, perform policy network parameter update processing, adjust the weight parameters of the reinforcement learning policy network, generate optimization control instructions, and send the optimization control instructions to the waste incinerator control system to adjust the combustion conditions of the incinerator.
[0179] The multi-objective reward value is a quantitative evaluation of the overall effect of the current control action. Based on this reward value, the parameters of the reinforcement learning policy network need to be updated. The purpose of updating the policy network parameters is to enable the reinforcement learning policy network to learn better control strategies and improve its decision-making ability under different operating conditions.
[0180] Through a series of algorithms and optimization methods, the weight parameters of the reinforcement learning policy network are adjusted based on the multi-objective reward value. Adjusting the weight parameters affects the network's output, enabling it to produce control actions that better meet actual needs in subsequent decisions. For example, if the multi-objective reward value is high, it indicates that the current control strategy is effective, and the network will appropriately increase the weight of relevant parameters; if the multi-objective reward value is low, the network will adjust the parameters to try and find a better control strategy.
[0181] After updating the strategy network parameters, optimized control commands are generated. These commands include adjusted incineration temperature setpoints and oxygen supply adjustments, derived from the updated strategy network outputs, enabling more effective optimization of the incineration process. The optimized control commands are then sent to the waste incinerator control system, which automatically adjusts the incinerator's combustion conditions based on these commands, such as adjusting burner power and fan airflow, thereby optimizing the waste incineration process, improving thermal efficiency, controlling pollutant emissions, and extending equipment lifespan.
[0182] In one implementation, step S500 may include the following steps S510-S560: Step S510: Store the current state representation vector, control action sequence, and multi-objective reward value into the experience replay pool. The experience replay pool uses a first-in-first-out mechanism to manage the data, and the storage capacity is the preset maximum number of samples to ensure the diversity of data distribution.
[0183] The experience replay pool is a buffer used to store historical data, providing diverse data samples for training reinforcement learning policy networks. The current state representation vector, control action sequence, and multi-objective reward value are stored as a single sample in the experience replay pool. The state representation vector contains information about the current state of the waste incineration process, the control action sequence represents the decisions made based on that state, and the multi-objective reward value is an evaluation of the effectiveness of those decisions.
[0184] The experience replay pool uses a first-in, first-out (FIFO) mechanism to manage data. When the number of data samples in the pool reaches the preset maximum number, the earliest sample entering the pool is removed to make room for new samples. This mechanism ensures the timeliness of the data in the experience replay pool while also ensuring the diversity of data distribution. Diverse data samples allow the reinforcement learning policy network to be exposed to state, action, and reward information under different conditions during training, improving the network's generalization ability and learning performance.
[0185] Step S520: Randomly sample a preset number of samples from the experience replay pool. Each sample contains a state representation vector, a control action sequence, a multi-objective reward value, and a next state representation vector to construct a training sample batch.
[0186] To train the reinforcement learning policy network, training data needs to be obtained from the experience replay pool. A predetermined number of samples are randomly sampled, and each sample contains a state representation vector, a control action sequence, a multi-objective reward value, and a next state representation vector. The state representation vector describes the waste incineration state at a certain moment, the control action sequence is the decision made for that state, the multi-objective reward value is the evaluation of that decision, and the next state representation vector indicates the next state the waste incineration process enters after taking the control action.
[0187] A predetermined number of randomly sampled samples are combined to construct a training sample batch. This training sample batch provides a representative set of data for training the reinforcement learning policy network. By learning from and analyzing these samples, the network can continuously adjust its parameters and improve its decision-making ability. Random sampling avoids correlation between samples, enabling the network to learn a wider range of state-action-reward relationships and enhancing its robustness.
[0188] Step S530: Based on the training sample batch, the target Q value of the policy network is calculated using the temporal difference algorithm. The target Q value is calculated by adding the multi-objective reward value to the product of the time discount factor and the Q value of the next state. The time discount factor is set according to the long-term reward importance of the task.
[0189] Temporal difference algorithms are an efficient method for calculating Q-values in reinforcement learning. Based on the training sample batches constructed in step S520, the temporal difference algorithm is used to calculate the target Q-value of the policy network. The target Q-value is an estimate of the future reward based on the current state-action relationship.
[0190] The target Q-value is calculated as the product of the multi-objective reward value, the time discount factor, and the next-state Q-value. The multi-objective reward value is an immediate assessment of the current control action, reflecting its effect at the current moment. The next-state Q-value is an estimate of the expected reward in the next state after taking the current control action. The time discount factor is a value between 0 and 1, set according to the importance of the task's long-term reward. A larger time discount factor indicates a greater focus on long-term rewards; a smaller time discount factor indicates a greater focus on short-term rewards.
[0191] Step S540: Input the state representation vector into the current policy network to obtain the current Q value, and calculate the mean square error between the current Q value and the target Q value as the loss function value. The loss function value reflects the prediction bias of the current policy network.
[0192] The state representation vectors from the training sample batch are input into the current reinforcement learning policy network. The network calculates the current Q-value based on its internal parameters and algorithm. The current Q-value is the policy network's prediction of the future reward based on the current state-action pair.
[0193] Calculate the mean squared error between the current Q-value and the target Q-value calculated in step S530, and use this mean squared error as the loss function value. The loss function value reflects the prediction bias of the current policy network, that is, the degree of difference between the network's prediction of future returns and the actual target value. The larger the loss function value, the more inaccurate the current policy network's prediction is, and the more necessary it is to adjust the network's parameters; the smaller the loss function value, the closer the network's prediction is to the actual target value, and the better the network's performance. By continuously adjusting the network parameters, the loss function value is gradually reduced, thereby improving the policy network's predictive ability and decision accuracy.
[0194] Step S550: The adaptive learning rate optimization algorithm is used to update the weight parameters of the policy network. The adaptive learning rate is dynamically adjusted according to the rate of change of the loss function value. The learning rate is increased when the loss decreases quickly and decreased when the loss decreases slowly.
[0195] In one implementation, step S550 may specifically include the following steps S551-S556: Step S551: Initialize the learning rate to a preset initial learning rate, set the learning rate adjustment period to a preset number of iterations, record the loss function value for each iteration, and construct a loss change sequence.
[0196] When updating the weight parameters of the policy network, the learning rate is first initialized to a preset initial learning rate. The initial learning rate is a suitable starting value determined based on experience and experiments, and it affects the training speed and convergence performance of the network.
[0197] In each training iteration, the loss function value is recorded. These loss function values are then arranged sequentially according to the iteration order to construct a loss change sequence. The loss change sequence can visually demonstrate the trend of the loss function value during training, providing a basis for subsequent learning rate adjustments.
[0198] Step S552: Calculate the average rate of change of the loss function value within the current iteration period. The average rate of change is the ratio of the difference between the average loss of the current period and the average loss of the previous period to the average loss of the previous period, reflecting the rate of loss reduction.
[0199] After each learning rate adjustment period, the average rate of change of the loss function value within the current iteration period is calculated. First, the average of all loss function values within the current period is calculated to obtain the average loss for the current period; similarly, the average of all loss function values within the previous period is calculated to obtain the average loss for the previous period.
[0200] The average rate of change reflects how quickly the loss function value decreases within the current iteration period. A positive and large average rate of change indicates that the loss function value is decreasing rapidly; a positive and small average rate of change indicates that the loss function value is decreasing slowly; and a negative average rate of change indicates that the loss function value may have increased, requiring further examination of the network's training status.
[0201] Step S553: Compare the average rate of change with the preset first threshold and second threshold. If the first threshold is greater than the second threshold, the average rate of change is less than the second threshold, indicating that the loss is decreasing too slowly. If the average rate of change is greater than the first threshold, the loss is decreasing too quickly.
[0202] Step S554: When the average rate of change is greater than the first threshold, the current learning rate is multiplied by a preset increase coefficient, which is greater than 1, to speed up the parameter update speed; when the average rate of change is less than the second threshold, the current learning rate is multiplied by a preset decrease coefficient, which is less than 1, to slow down the parameter update speed.
[0203] Step S555: Apply boundary constraints to the adjusted learning rate to ensure that the learning rate does not exceed the preset maximum learning rate and is not lower than the preset minimum learning rate, so as to avoid oscillation caused by an excessively large learning rate or convergence stagnation caused by an excessively small learning rate.
[0204] If the adjusted learning rate exceeds the preset maximum learning rate, it is truncated to the maximum learning rate. The maximum learning rate is to prevent the learning rate from being too large, which could cause the network to oscillate wildly during training and fail to converge to the optimal solution. If the adjusted learning rate is lower than the preset minimum learning rate, it is truncated to the minimum learning rate. The minimum learning rate is to avoid the learning rate being too small, which could lead to excessively small parameter update steps, resulting in extremely slow network convergence or even stagnation. By limiting the learning rate, we can ensure its rationality and stability, thereby improving the training effect of reinforcement learning policy networks.
[0205] Step S556: Based on the adjusted learning rate and the gradient of the loss function, update the weight parameters of the policy network using gradient descent. The update amount of the weight parameters is the learning rate multiplied by the negative of the gradient, completing one parameter update iteration.
[0206] Step S560: After the update is completed, the output of the current strategy network is used as an optimization control command, which includes the incineration temperature setpoint and oxygen supply adjustment value. This command is sent to the waste incinerator control system via the industrial bus protocol. The control system adjusts the operating parameters of the burner and blower according to the command.
[0207] After updating the weight parameters of the strategy network, the output of the current strategy network becomes the optimized control command. The optimized control command includes the adjusted incineration temperature setpoint and oxygen supply adjustment value. These values are obtained after the strategy network has learned and optimized them, which can more effectively optimize the waste incineration process.
[0208] Optimized control commands are sent to the waste incinerator control system via an industrial bus protocol. The industrial bus protocol is a standard protocol for communication between industrial devices, ensuring accurate transmission and reliable reception of commands. Upon receiving the optimized control commands, the waste incinerator control system automatically adjusts the operating parameters of the burner and blower based on the incineration temperature setpoint and oxygen supply adjustment value specified in the command. For example, if the incineration temperature setpoint in the optimized control command increases, the control system increases the fuel supply to the burner, improving combustion intensity and bringing the furnace temperature to the setpoint; if the oxygen supply adjustment value increases, the control system adjusts the blower's airflow to increase oxygen supply, ensuring complete combustion of waste. By continuously updating the strategy network parameters and sending optimized control commands, real-time optimization and control of the waste incineration process are achieved, improving thermal efficiency, reducing pollutant emissions, and extending equipment lifespan.
[0209] It is understood that the various algorithms involved in the above descriptions of the embodiments of the present invention can all be obtained from relevant content in the prior art. To save space, they will not be elaborated on in the embodiments of the present invention. In addition, those skilled in the art can supplement the details based on common knowledge in the art when implementing the solutions of the present invention. For example, they can use normalization to eliminate dimensional conflicts before feature fusion, use interpolation to eliminate dimensional differences, reasonably set thresholds based on historical data, experience or business scenario requirements, train the model based on a general model training method, set the number of layers in the model structure based on actual needs, select activation functions, etc. The present invention will not provide redundant descriptions of overly detailed implementation processes here.
[0210] Please see Figure 2 , Figure 2This is a schematic diagram of a computer system provided in an embodiment of the present invention. The computer system includes at least a processor 101, a communication interface 102, and a memory 103. The processor 101, communication interface 102, and memory 103 can be connected via a bus or other means. The processor 101 (or Central Processing Unit, CPU) is the computing and control core of the computer system, capable of parsing various instructions and processing various data within the computer system. The communication interface 102 may optionally include a standard wired interface or a wireless interface (such as Wi-Fi, mobile communication interface, etc.), and can be used to send and receive data under the control of the processor 101; the communication interface 102 can also be used for data transmission and interaction within the computer system. The memory 103 is a storage device in the computer system used to store programs and data. It is understood that the memory 103 here can include the computer system's built-in memory, or it can include extended memory supported by the computer system. The memory 103 provides storage space, which stores the computer system's operating system; this invention does not limit this storage space.
[0211] In one embodiment, the processor 101 executes the waste incineration plant task optimization method using reinforcement learning provided above in the embodiments of the present invention by running a computer program in the memory 103.
Claims
1. A method for optimizing tasks in a waste incineration plant using reinforcement learning, characterized in that, The method comprises: acquiring a real-time operation data set of a waste incineration site, the real-time operation data set comprising waste composition data and incinerator operation parameter data recorded by an incinerator control system, the waste composition data comprising waste moisture content measurement values and waste combustible content measurement values, and the incinerator operation parameter data comprising in-furnace temperature measurement values and flue gas oxygen content measurement values; based on the real-time operation data set, performing state encoding processing, extracting waste calorific value distribution features and incinerator operation state features, and performing vectorization fusion processing on the waste calorific value distribution features and the incinerator operation state features to generate a state representation vector; calling a pre-trained reinforcement learning policy network to perform policy inference processing on the state representation vector, the reinforcement learning policy network adopting a double-channel output structure, wherein a main channel outputs an incineration temperature set value, and a secondary channel outputs an oxygen supply amount adjustment value, and a control action sequence is generated in combination with the outputs of the main channel and the secondary channel; according to the control action sequence, evaluating the incineration process performance through reward function calculation processing, wherein the reward function calculates a weighted sum of a thermal efficiency optimization indicator, a pollutant emission control indicator, and an equipment life maintenance indicator to generate a multi-objective reward value; based on the multi-objective reward value, performing policy network parameter update processing, adjusting the weight parameters of the reinforcement learning policy network, and generating an optimized control instruction, which is sent to the waste incinerator control system to adjust the combustion conditions of the incinerator.
2. The method of claim 1, wherein, The method comprises: dividing the real-time operation data set into continuous sampling window units, each sampling window unit containing waste composition data and incinerator operation parameter data within a preset time period, and establishing a data time sequence association index; performing regional association analysis on the waste moisture content measurement values and the waste combustible content measurement values in each sampling window unit, calculating the heat value contribution weights of different stacking regions, and generating regional heat value change curves; based on the regional heat value change curves and the spatiotemporal distribution law of the in-furnace temperature measurement values, constructing a physical field coupling matrix, the row dimension of the physical field coupling matrix corresponding to the waste stacking region number, the column dimension corresponding to the incinerator operation parameter type, and the matrix elements representing the association strength of the regional heat value and the operation parameters; performing feature selection on the waste composition data and the incinerator operation parameter data through the physical field coupling matrix, retaining the feature dimensions with an association strength exceeding a preset threshold, and generating reduced waste calorific value distribution features and incinerator operation state features; calculating the fusion coefficients of the waste calorific value distribution features and the incinerator operation state features through a dynamic weight distribution algorithm, the fusion coefficients being dynamically adjusted with the timestamp of the sampling window unit, and reflecting the influence weights of the two types of features on the incineration process in different time periods. The reduced dimensionality waste heat value distribution features and the incinerator operation state features are weighted and vectorized based on the fusion coefficient to generate a state representation vector containing spatiotemporal correlation, each dimension of the state representation vector corresponding to a weighted and fused feature value.
3. The method of claim 2, wherein, The waste moisture content measurement value and the waste combustible content measurement value in each sampling window unit are subjected to regional correlation analysis to calculate the heat value contribution weight of different stacking regions and generate a regional heat value change curve, including: The waste feeding area of the waste incineration site is divided into multiple rows and multiple columns of sub-regions according to grid division rules, and the waste moisture content measurement value and the waste combustible content measurement value of the sub-regions are collected; For each sub-region, the waste moisture content measurement value and the waste combustible content measurement value of the sub-region are converted into unit mass heat value based on a preset heat value estimation model; The product of the unit mass heat value of each sub-region and the waste feeding amount of the sub-region is calculated to obtain the heat value contribution of the sub-region, and the sum of the heat value contributions of all sub-regions is calculated as the total heat value of the sampling window unit; The heat value contribution of each sub-region is divided by the total heat value to obtain the heat value contribution weight of the sub-region, and a regional heat value contribution weight matrix is constructed, with the matrix elements being the heat value contribution weights of the corresponding sub-regions; According to the timestamp order of the sampling window unit, the regional heat value contribution weight matrix of each window is extracted in sequence, the weight change rate of the same sub-region in consecutive windows is calculated, and a set of regional heat value change curves containing multiple rows and multiple columns of curves is generated; The set of regional heat value change curves is subjected to noise reduction processing to generate a final regional heat value change curve.
4. The method of claim 1, wherein, The state representation vector is subjected to policy inference processing by calling a pre-trained reinforcement learning policy network, and the reinforcement learning policy network adopts a double-channel output structure, wherein the main channel outputs an incineration temperature set value, and the auxiliary channel outputs an oxygen supply amount adjustment value, and the control action sequence is generated by combining the outputs of the main channel and the auxiliary channel, including: The state representation vector is input into the feature division and resolution layer of the reinforcement learning policy network, and the state representation vector is divided into a heat value feature sub-vector and an operation feature sub-vector according to the physical meaning of the feature dimension; The mutual information entropy of the heat value feature sub-vector and the operation feature sub-vector is calculated through the cross-attention module of the reinforcement learning policy network, and a feature correlation weight matrix is generated based on the mutual information entropy, wherein the row dimension of the feature correlation weight matrix corresponds to the dimension of the heat value feature sub-vector, the column dimension corresponds to the dimension of the operation feature sub-vector, and the matrix elements represent the correlation closeness between the two types of feature dimensions; The heat value feature sub-vector and the operation feature sub-vector are input into the feature fusion layer, and are fused by combining the feature correlation weight matrix to generate a fused feature vector, wherein the dimension of the fused feature vector is the same as that of the state representation vector; The fused feature vector is input into the temperature decision network of the main channel, and feature mapping is performed through multiple fully connected layers and residual connection structures, and batch normalization processing and an activation function are connected after each fully connected layer to output an incineration temperature set value; The fusion feature vector is input into an oxygen supply decision network of a secondary channel, time sequence feature extraction is performed through a convolution-circulation hybrid structure, local features are extracted through a one-dimensional convolution processing layer, a long short-term memory network layer is used to model time sequence dependency, and an oxygen supply adjustment value is output; The incineration temperature set value output by the primary channel and the oxygen supply adjustment value output by the secondary channel are correspondingly combined in time stamp order to generate a control action sequence containing time labels, each element of the control action sequence containing an incineration temperature set value, an oxygen supply adjustment value and time stamp information at a corresponding time.
5. The method of claim 4, wherein, The cross-attention module through the network calculates the mutual information entropy of the calorific value feature sub-vector and the operation feature sub-vector, and generates a feature correlation weight matrix based on the mutual information entropy, including: The calorific value feature sub-vector is taken as a query feature sub-vector, and the operation feature sub-vector is taken as a key feature sub-vector, and linear transformation matrices are used to map the dimensions of the query feature sub-vector and the key feature sub-vector respectively, so that the mapped query feature sub-vector and the key feature sub-vector have the same dimension; The dot product of the mapped query feature sub-vector and the key feature sub-vector is calculated to obtain an initial attention score matrix; The initial attention score matrix is subjected to row normalization processing to obtain a normalized attention score matrix, and the mutual information entropy of the query feature sub-vector and the key feature sub-vector is calculated based on the normalized attention score matrix; The mutual information entropy value is compared with a preset reference mutual information entropy, a mutual information entropy deviation rate is calculated, and the scaling coefficient of the normalized attention score matrix is adjusted based on the mutual information entropy deviation rate, so that the element value of the adjusted matrix reflects the actual correlation strength; The adjusted normalized attention score matrix is taken as the feature correlation weight matrix, and the matrix element represents the correlation weight of the calorific value feature sub-vector dimension and the operation feature sub-vector dimension.
6. The method of claim 4, wherein, The fusion feature vector is input into the temperature decision network of the primary channel, feature mapping is performed through multiple full connection layers and residual connection structures, a batch normalization processing and an activation function are connected after each full connection layer, and an incineration temperature set value is output, including: The fusion feature vector is input into the first full connection layer of the temperature decision network of the primary channel, the input dimension is the dimension of the fusion feature vector, the output dimension is a preset first hidden layer dimension, a first hidden layer feature is generated through linear transformation of the fusion feature vector by a weight matrix; The first hidden layer feature is subjected to batch normalization processing, the mean and variance of the first hidden layer feature of all samples in the batch are calculated, the first hidden layer feature is standardized to a feature value with a mean of 0 and a variance of 1, and a normalized first hidden layer feature is generated; The normalized first hidden layer feature is input into an activation function for non-linear mapping, the non-zero gradient of the negative half axis is retained, and an activated first hidden layer feature is generated; The activated first hidden layer feature is connected in residual with the fusion feature vector, the fusion feature vector is directly superimposed on the activated first hidden layer feature through a jump connection to generate a residual enhanced first hidden layer feature; The full connection, batch normalization, activation and residual connection process is repeatedly performed to sequentially process the residual enhanced first hidden layer features to generate second hidden layer features and third hidden layer features, and the output dimension of each layer is decreased by a preset ratio; The third hidden layer features are input into the output layer full connection layer, the input dimension is the dimension of the third hidden layer features, the output dimension is 1, and the original temperature setting value is generated by linear transformation, the original temperature setting value is substituted into the temperature constraint function, and the incineration temperature setting value is output.
7. The method of claim 4, wherein, The fusion feature vector is input into the oxygen supply decision network of the auxiliary channel, time series feature extraction is performed through a convolution-recurrent hybrid structure, local features are extracted through a one-dimensional convolution processing layer, and a long short-term memory network layer is used to model the time series dependency relationship, and an oxygen supply adjustment value is output, including: The fusion feature vector is dimensionally reconstructed and converted into a one-dimensional feature sequence, the sequence length is the dimension of the fusion feature vector, and each time step corresponds to a dimension value of the fusion feature vector; The one-dimensional feature sequence is input into the one-dimensional convolution processing layer of the auxiliary channel oxygen supply decision network, the one-dimensional convolution processing layer uses a preset number of convolution kernels, each convolution kernel has a preset window size, and local feature extraction is performed on the one-dimensional feature sequence through a sliding window to generate a multi-channel convolution feature map; The multi-channel convolution feature map is subjected to maximum pooling processing in the time dimension, the pooling window size is a preset value, the step length is equal to the pooling window size, the maximum value of the local features of each channel is taken, and a reduced dimension convolution feature sequence is generated; The reduced dimension convolution feature sequence is input into the long short-term memory network layer, and the hidden state vector of each time step is output, and the hidden state vector of the last time step is taken as the output of the long short-term memory network layer; The hidden state vector output by the long short-term memory network layer is input into the full connection layer, and the hidden state vector is mapped to a single numerical value through linear transformation as the original oxygen supply adjustment value; The original oxygen supply adjustment value is substituted into the adjustment threshold constraint function, if the absolute value of the adjustment value exceeds the preset single adjustment threshold, it is truncated to the threshold boundary value, otherwise the original value is maintained, and the oxygen supply adjustment value is generated.
8. The method of claim 1, wherein, According to the control action sequence, the performance of the incineration process is evaluated by a reward function, wherein the reward function calculates the weighted sum of the heat efficiency optimization indicator, the pollutant emission control indicator and the equipment life maintenance indicator to generate a multi-objective reward value, including: According to the time stamp order of the control action sequence, each action element is sequentially sent to the incinerator simulation model, the incinerator simulation model is constructed by physical field modeling, the input is the incineration temperature setting value and the oxygen supply adjustment value, and the output is the heat efficiency related data, the pollutant emission data and the equipment running state data within a preset time length; The heat efficiency optimization indicator is calculated based on the heat efficiency related data, the heat efficiency related data includes the input heat measurement value and the output heat measurement value of the incinerator, and is calculated by comparing the input-output heat ratio and the historical optimal heat efficiency; The pollution emission control index is calculated based on the pollution emission data, the pollution emission data including concentration measurement values of multiple pollutants, the emission compliance rate being calculated by comparing each pollution concentration with an emission standard, and then being averaged by weighting; The equipment life maintenance index is calculated based on the equipment operation state data, the equipment operation state data including temperature measurement values, vibration amplitude measurement values and operation time cumulative values of key components, the life loss rate being converted from the equipment aging model, and then being normalized to obtain; The dynamic weight library of the reinforcement learning environment is used to obtain the thermal efficiency optimization index weight, the pollution emission control index weight and the equipment life maintenance index weight, the dynamic weight library automatically adjusting the weight values according to the current working condition of the incineration process, and the emission control index weight being increased when the working condition is abnormal; The thermal efficiency optimization index, the pollution emission control index and the equipment life maintenance index are multiplied by their respective weights, and the three are added to obtain a multi-objective reward value, the value range of the multi-objective reward value being limited in a preset interval through standardization processing.
9. The method of claim 8, wherein, The thermal efficiency optimization index is calculated based on the thermal efficiency related data, the thermal efficiency related data including input heat measurement values and output heat measurement values of the incinerator, the input heat measurement values and the output heat measurement values being calculated by comparing the input-output heat ratio and the historical optimal thermal efficiency, including: The input heat measurement values and the output heat measurement values are extracted from the thermal efficiency related data, the input heat measurement values being the total heat released by garbage combustion, calculated by the product of the total mass of garbage and the heat value per unit mass, and the output heat measurement values being the steam heat generated by the incinerator, calculated by the difference between the steam flow, the steam enthalpy value and the inlet water enthalpy value; The ratio of the output heat measurement values to the input heat measurement values is calculated to obtain an actual thermal efficiency value, the actual thermal efficiency value reflecting the energy conversion efficiency of the incinerator under the current control action; The highest thermal efficiency record value under the same working condition is retrieved from the historical database as a thermal efficiency reference value, the same working condition being determined by matching the garbage composition parameters, the incineration load parameters and the environmental temperature parameters; The ratio of the actual thermal efficiency value to the thermal efficiency reference value is calculated to obtain a thermal efficiency relative value; The thermal efficiency relative value is multiplied by a preset thermal efficiency reference coefficient, and the product is taken as the thermal efficiency optimization index, the larger the thermal efficiency optimization index value, the better the optimization effect of the current control action on the thermal efficiency.
10. A computer system, characterized by The memory stores a computer program; The processor loads the computer program to implement the garbage incineration plant task optimization method using reinforcement learning according to any one of claims 1-9.