A garbage incinerator HCl concentration prediction method based on a mixture model
By using a hybrid model of Transformer feature extractor and XGBoost predictor, the problems of data lag and nonlinearity in HCl concentration prediction in waste incinerators were solved, achieving multi-step advanced and accurate prediction, and improving the accuracy and adaptability of the control system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUTH CHINA UNIV OF TECH
- Filing Date
- 2025-09-18
- Publication Date
- 2026-08-04
AI Technical Summary
Existing methods for predicting HCl concentration in waste incinerators suffer from data lag, nonlinearity, and multivariate coupling characteristics, making it difficult to meet real-time control requirements. Traditional hardware monitoring systems cannot detect future changes in emission concentrations, leading to difficulties in precise control.
A hybrid model is adopted, combining the Transformer feature extractor and the XGBoost predictor. It captures time series dependencies through a self-attention mechanism for feature extraction and auxiliary prediction, and performs multi-step advance prediction through feature fusion.
It enables multi-step, advanced, and accurate prediction of HCl concentration, improves the accuracy and generalization ability of the prediction model, provides sufficient control time margin for the deacidification control system, and ensures safe and stable operation.
Smart Images

Figure CN121328802B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of flue gas pollutant prediction for waste incinerators, specifically relating to a method for predicting HCl concentration in waste incinerators based on a hybrid model. Background Technology
[0002] Waste-to-energy incineration technology has gradually become the main method for urban and rural domestic waste treatment due to its advantages such as significant waste reduction and utilization of waste heat resources. As of October 2024, there were 1,010 incineration enterprises nationwide, with 2,172 incinerators and a total incineration capacity of approximately 1.11 million tons / day. However, the complex and diverse composition of waste inevitably generates a large amount of flue gas pollutants during the incineration process. Among them, hydrogen chloride (HCl) is one of the most significant acidic pollutants. Its large-scale emissions not only cause environmental pollution but also corrode equipment, affecting the safe operation of waste incinerators. With increasingly stringent environmental emission standards, higher requirements have been placed on the emission control of pollutants such as HCl. In some areas, the emission limit for HCl is controlled at 10 mg / Nm³. 3 The following is a summary of the main sources of HCl. HCl mainly comes from chlorides in waste components, such as NaCl and KCl in PVC, rubber, leather, and kitchen waste. However, the types of municipal solid waste are numerous and fluctuate greatly with the seasons, resulting in significant fluctuations in HCl generation concentration and weak regularity, which poses a severe challenge to the desulfurization control effect of waste incinerators. At present, traditional hardware monitoring systems have the following defects: (1) Existing measurement systems have certain data lag problems due to the use of sampling, preprocessing, and re-detection, making it difficult to meet the needs of real-time control. (2) Affected by the high temperature, high humidity, and high dust environment of the flue, the system needs to be backflushed and maintained regularly, during which time it cannot work normally, resulting in data loss. (3) Traditional hardware monitoring systems cannot sense future changes in emission concentration, making it difficult to provide key data support for desulfurization optimization control. These limitations make it difficult to meet the needs of accurate and real-time control by simply relying on hardware monitoring. Therefore, by integrating big data analysis and machine learning modeling, a real-time prediction method for HCl concentration in waste incinerators with high prediction accuracy and strong generalization performance can be developed. This method, combined with a hardware monitoring system, forms a software-hardware collaborative intelligent sensing system, which plays a key role in improving the intelligent, precise, and economical control level of acidic gases.
[0003] An existing method and control system for controlling HCl concentration emissions from a waste incinerator (CN114110616B) includes a database module, a data classification module, a model calculation module, a parameter optimization module, an output compensation module, and a slurry valve PID module. The control method primarily establishes a database based on historical data from the incinerator. After determining the first HCl emission concentration and the first slurry flow rate within a preset range, a calculation model is built, and the model coefficients are optimized. Then, the difference between the HCl emission concentration in the previous hour and a set threshold is used to calculate the slurry flow rate and compensation coefficient. Finally, the target slurry volume value is obtained, and the slurry flow rate is adjusted in real time to achieve HCl emission control from the incinerator. This method does not predict the HCl concentration; the first HCl emission concentration is determined within a preset HCl emission concentration range, and the acquisition of HCl concentration data relies entirely on historical data from the hardware measurement system.
[0004] A method and system for predicting the raw emission concentrations of SO2 and HCl generated from waste incineration (CN115389714A) are disclosed. This method selects operational parameter data related to the raw emission concentrations of SO2 and HCl based on waste incineration theoretical analysis. After standardizing the data, a Copula function is used to reduce the data dimensionality and determine the input parameters. The determined input parameters are used as input, and pre-acquired raw emission concentrations of SO2 and HCl corresponding to the input parameters are used as output to repeatedly train a preset backpropagation (BP) neural network until the model error is lower than a preset training threshold. Training is then stopped, and the BP neural network model from the last training iteration is used as the prediction model for the raw emission concentrations of SO2 and HCl. Real-time input data from the incinerator is input into the prediction model to predict the raw emission concentrations of SO2 and HCl at the current moment. This method predicts the raw emission concentrations of SO2 and HCl before deacidification treatment. However, most power plants lack measurement points for raw emission concentrations of SO2 and HCl due to high operating and maintenance costs of monitoring equipment, resulting in a lack of subsequent calibration data for the prediction model and affecting its long-term prediction performance.
[0005] A method, system, and storage medium for predicting acid gas emission concentrations (CN113947142B) are disclosed. The acid gases primarily refer to HCl and SO2 produced during waste incineration. A machine learning model is built in the model building module, and after training, it serves as the final acid gas concentration prediction model. The specific model used is not specified. The input parameters for this prediction model are the composition of municipal solid waste and the type of incinerator. Collinearity-free preprocessing is performed on the municipal solid waste composition variables. This method relies on information about the composition of municipal solid waste for predicting HCl and SO2 concentrations. However, waste-to-energy plants in actual operation do not have real-time waste composition information; they only have offline monitoring data from third-party institutions collected quarterly. Therefore, the applicability of the model is limited.
[0006] A method and system for collaborative prediction and intelligent control of multiple pollutants in waste incineration flue gas (CN 119289371A) is disclosed. This method constructs a collaborative prediction model for four pollutants—HCl, SO2, NOx, and PM—in waste incineration flue gas based on an LSTM layer structure. The method resamples data with a 1-second sampling interval using a 5-minute mean resampling, which may lead to the loss of some short-term variation details, affecting the accuracy of control. The method selects model input variables based on a Pearson correlation coefficient absolute value greater than 0.3, but the Pearson correlation coefficient can only measure the linear correlation between two variables. The waste incineration process exhibits significant nonlinearity and multivariate coupling characteristics, and feature selection based on Pearson correlation analysis cannot effectively overcome these characteristics, limiting the improvement of model prediction performance. Furthermore, while the method uses an LSTM model for multi-objective prediction, its ability to extract local feature information is weak, and it does not pay sufficient attention to individual pollutants during modeling, affecting the prediction accuracy and generalization performance of HCl emission concentration.
[0007] There are three main ways to obtain HCl concentration in existing waste incinerators: (1) monitoring HCl concentration through a continuous emission monitoring system (CEMS) installed in the flue of the waste incineration plant; (2) predicting HCl concentration by establishing a mechanism model by analyzing the influence of chlorine content and carbon-chlorine ratio in the waste components on HCl concentration generation; and (3) predicting HCl concentration through traditional statistical models (such as PLS) or single machine learning models (such as LSTM).
[0008] The existing technologies for obtaining HCl concentration in waste incinerators have the following limitations: (1) The existing CEMS system, due to the use of sampling, preprocessing and re-detection, results in a certain lag in measurement, which cannot meet the real-time control requirements; (2) The complex and irregular composition of waste leads to the variable HCl generation mechanism, and it is difficult to adapt to actual production conditions to establish a mechanism model by analyzing waste composition to predict HCl concentration; (3) HCl emissions during waste incineration are affected by multiple factors (combustion temperature, waste composition, flue gas treatment parameters, etc.), and have significant nonlinear, multivariate coupling and time series characteristics. Traditional statistical models (such as PLS, PCA) have problems such as weak nonlinear mapping ability, poor adaptability to complex working conditions and insufficient prediction accuracy; while single machine learning models (such as BP, SVM, ELM, LSTM, etc.) are difficult to capture spatiotemporal change characteristics at the same time, and have poor ability to deeply extract operating data and insufficient model generalization ability. At the same time, existing prediction methods are mostly for predicting the concentration value at the next moment, lacking prediction of the concentration value after multiple time steps, and cannot provide sufficient time margin for the advanced regulation of the desulfurization control system. Summary of the Invention
[0009] The purpose of this invention is to provide a method for predicting HCl concentration in waste incinerators based on a hybrid model, in order to solve the problem that the HCl generation concentration fluctuates greatly and has weak regularity due to the wide variety of urban domestic waste and large seasonal fluctuations, resulting in strong nonlinearity and time-varying characteristics that make it difficult to predict.
[0010] The present invention is achieved by at least one of the following technical solutions.
[0011] A method for predicting HCl concentration in waste incinerators based on a hybrid model includes the following steps:
[0012] Step 1: Obtain historical operating data of the waste incinerator, including various operating parameters, and perform data preprocessing on the historical operating data;
[0013] Step 2: Calculate the MI value between each preprocessed operating parameter and the target HCl concentration using the MI method. Select multiple variables as input feature variables based on the MI value and standardize them to serve as the original input features for the feature extractor.
[0014] Step 3: Input the original input features into the Transformer feature extractor, and output the feature extraction results and the HCl concentration-assisted prediction results;
[0015] Step 4: Perform feature fusion on the original input features, feature extraction results, and HCl concentration-assisted prediction results;
[0016] Step 5: Input the fused features into the XGBoost predictor to predict the HCl concentration, and then output the final predicted HCl concentration value after inverse standardization.
[0017] Furthermore, in step 2, operating parameters with an MI value greater than 0.5 relative to the HCl concentration data are selected, and the operating parameters with an MI value greater than 0.5 and the historical values of HCl concentration are standardized by z-score and used as the original input features of the Transformer feature extractor.
[0018] Furthermore, in step 3, the Transformer feature extractor includes an input layer, a fully connected projection layer, a positional encoding layer, an encoder, a global average pooling layer, and a dual output layer. The encoder consists of N identical stacked layers, each with two sub-layers. The first sub-layer uses a multi-head attention mechanism, and the second sub-layer is a fully connected feedforward network. Residual connections and normalization are used to process the output after both sub-layers. The number of layers N in the encoder is one of the hyperparameters of this hybrid model.
[0019] Furthermore, in step 3, the Transformer feature extractor produces two outputs: the feature extraction result F and the auxiliary prediction result P. The formula for calculating the feature extraction result F is as follows:
[0020]
[0021] In the formula, L is the sequence length, and H is... t W represents the feature encoded by multiple transformers at time t. f b is a trainable weight matrix f Here, is the bias vector of the feature extraction layer, and swish is the self-gated activation function;
[0022] The formula for calculating the auxiliary prediction result P is as follows:
[0023]
[0024] In the formula, W p b is the weight of the prediction layer. p It is a scalar bias vector.
[0025] Furthermore, the overall loss function of the Transformer feature extractor is a weighted sum of the loss for feature extraction and the loss for auxiliary prediction.
[0026] Furthermore, the feature fusion method in step 4 is as follows:
[0027]
[0028] In the formula, X comb Let x' represent the fused feature variables, where n is the number of original input features, k is the number of samples, and x'' is the number of samples. kn For the k-th data point of the n-th original input feature variable after standardization, F kn P represents the k-th output of the Transformer feature extractor after extracting features from the n-th original input feature variable. k This is the k-th output of the Transformer feature extractor for auxiliary prediction, where α is the weight of the feature extraction variable and (1-α) is the weight of the auxiliary prediction variable.
[0029] Furthermore, in step 5, the objective function of the XGBoost predictor includes a loss function and a regularization term, wherein the loss function uses mean squared error.
[0030] The system for implementing the aforementioned method for predicting HCl concentration in waste incinerators based on a hybrid model includes:
[0031] The data preprocessing module is used to preprocess historical operating data containing various operating parameters;
[0032] The feature variable selection module is used to select original input feature variables from various operating parameters in historical operating data, calculate the MI value between each operating parameter variable and the target value HCl concentration using the mutual information method, and select multiple variables as input feature variables based on the MI value.
[0033] The standardization module is used to standardize the preprocessed data, including standardization and destandardization. The selected input feature variables are standardized and used as the original input features of the feature extractor. The output of the prediction module is destandardized and used as the final predicted value.
[0034] The feature extraction module uses the Transformer feature extractor to extract features and assist in prediction from the standardized raw input features and output the results.
[0035] The feature fusion module is used to weight and fuse the original input features, the extracted features output by the Transformer feature extractor, and the auxiliary prediction results, and use them as the input to the XGBoost predictor.
[0036] The model prediction module uses the XGBoost predictor to predict the HCl concentration at time t+3.
[0037] The parameter optimization module is used to optimize hyperparameters, utilizing the Optuna hyperparameter optimization framework to find the best values for each parameter.
[0038] A computer device according to the present invention includes a memory and a processor, the memory being electrically connected to the processor, the memory storing a computer program, which, when executed by the processor, causes the processor to implement the method described herein.
[0039] The present invention provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor implements the method described herein.
[0040] Compared with existing technologies, the beneficial effects of the present invention are as follows:
[0041] This invention discloses a hybrid model-based method for predicting HCl concentration in waste incinerators. It utilizes a Transformer feature extractor to capture the dependency relationship between any two time steps within the sequence from the raw data using a self-attention mechanism, extracting features and performing auxiliary prediction. The feature vectors are then fused, with the predictor's input features being a weighted combination of standardized raw input features, extracted features from the Transformer feature extractor, and auxiliary prediction results. Finally, an XGBoost predictor is used to achieve multi-step, accurate, and forward-looking prediction of the HCl concentration at time t+3. This method effectively combines the feature extraction capabilities of deep learning with the powerful predictive capabilities of gradient boosting tree models, significantly improving the accuracy and generalization ability of the prediction model. Furthermore, this invention's prediction method can predict the HCl concentration after three time steps, providing data support and time margin for precise HCl concentration control, thereby ensuring the safe, stable, and environmentally friendly operation of the desulfurization system. Attached Figure Description
[0042] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as limiting the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 This is a block diagram of a waste incinerator HCl concentration prediction system according to an embodiment of the present invention.
[0044] Figure 2 A flowchart illustrating a method for predicting HCl concentration in a waste incinerator based on a hybrid model, provided in this embodiment of the invention.
[0045] Figure 3 This is a prediction result diagram of an embodiment of the present invention. Detailed Implementation
[0046] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0047] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that the specific embodiments described herein are only for explaining this application and are not intended to limit this application.
[0048] like Figure 1 As shown in this embodiment, a waste incinerator HCl concentration prediction system based on a hybrid model includes:
[0049] The data preprocessing module is used to preprocess historical operating data containing various operating parameters, mainly including outlier handling and noise reduction.
[0050] The feature variable selection module is used to select original input feature variables from various operating parameters in historical operating data. It uses the mutual information (MI) method to calculate the MI value between each operating parameter variable and the target value HCl concentration, and selects multiple variables as input feature variables based on the MI value.
[0051] The standardization module is used to standardize the preprocessed data, including standardization and destandardization. The preprocessed data is divided into training, validation, and test sets in a 7:2:1 ratio before standardization. The training set is used for training the hybrid model; the validation set is used for hyperparameter tuning and early stopping strategies; the test set is only used for the final performance evaluation of the hybrid model to test its generalization ability and is not used during training or tuning. The selected input feature variables are standardized and used as the raw input features for the feature extractor. The output of the prediction module is destandardized and used as the final predicted value.
[0052] The feature extraction module is used to extract features and assist in prediction from the standardized raw input features and output the results. It uses a simplified Transformer model as the feature extractor.
[0053] The feature fusion module is used to weight and fuse the original input features, the extracted features output by the Transformer feature extractor, and the auxiliary prediction results, and use them as input to the XGBoost predictor.
[0054] The model prediction module is used to predict the HCl concentration at time t+3, and it uses the XGBoost model as the predictor.
[0055] The parameter optimization module is used to optimize the hyperparameters of the entire hybrid model, using the Optuna hyperparameter optimization framework to find the optimal parameters.
[0056] This method employs a hybrid model to predict the HCl concentration at time t+3. The hybrid model is a cascaded hybrid structure containing a Transformer feature extractor and an XGBoost predictor. The Transformer feature extractor captures the dependency between any two time steps within the sequence through a self-attention mechanism. The input to the XGBoost predictor includes standardized original features, a global temporal pattern representation of the historical sequence obtained by the Transformer feature extractor, and a preliminary prediction of the current sequence, thus achieving enhanced representation of temporal information. This method effectively combines the feature extraction capabilities of deep learning with the powerful predictive capabilities of gradient boosting tree models, significantly improving the accuracy and generalization ability of the prediction model. It achieves multi-step, accurate, and forward-looking prediction of the HCl concentration at time t+3, providing sufficient response time for adjustments in industrial control systems.
[0057] like Figure 2 As shown, the HCl concentration prediction method for waste incinerators based on a hybrid model in this embodiment includes the following steps:
[0058] Step 1: Obtain historical operating data from the waste incinerator's DCS system. This data includes operating parameters such as primary air volume, secondary air volume, grate velocity, slurry pipe flow rate, and average air preheater outlet temperature, as well as historical values of HCl concentration. Data preprocessing is performed on the historical operating data, primarily including outlier handling and noise reduction. Outlier handling includes handling extreme and regular outliers. First, extreme outliers with negative HCl concentrations due to temporary instrument malfunctions under harsh conditions are filtered and removed. Then, after removing extreme outliers, the 3σ principle is used to detect other outliers, and these outliers are replaced with linear interpolation. Gaussian smoothing is used for noise reduction.
[0059] Step 2: Calculate the MI value between each running parameter variable after preprocessing in Step 1 and the target HCl concentration using the Mutual Information (MI) method. Select n feature variables with MI values greater than 0.5 and the historical HCl concentration at time t to form a dataset. Divide the dataset into training set, validation set and test set in a 7:2:1 ratio. After standardization using the z-score standardization method, the dataset is used as the original input features of the Transformer feature extractor.
[0060] Step 3: Input the n standardized raw feature variables obtained in Step 2 into the Transformer feature extractor for feature extraction, obtaining n feature extraction results F and an auxiliary prediction result P of HCl concentration at time t+3. The hyperparameters of the feature extractor are optimized using the Optuna hyperparameter optimization framework. The formula for calculating the feature extraction result F is as follows:
[0061]
[0062] In the formula, L is the sequence length, and H is... t For features encoded by multiple transformers, W f b is a trainable weight matrix f Let x be the bias vector of the feature extraction layer, and swish be the self-gated activation function. linear ·σ(x linear ), where σ(x) linear ) represents the Sigmoid function, x linear This is the result of linear transformation and biasing of the global feature vector.
[0063] The formula for calculating the auxiliary prediction result P is as follows:
[0064]
[0065] In the formula, W p b is the weight of the prediction layer. p It is a scalar bias vector.
[0066] Step 4: Combine the standardized n original input features x′ obtained in Step 2 with the n extracted feature variables F and one auxiliary prediction variable P output in Step 3, and fuse them according to the following formula:
[0067]
[0068] In the formula, X comb Let x' represent the fused feature variables, where n is the number of original input features, k is the number of samples, and x'' is the number of samples. kn For the k-th data point of the n-th original input feature variable after standardization, F kn P represents the k-th output of the Transformer feature extractor after extracting features from the n-th original input feature variable. kThis is the k-th output of the Transformer feature extractor for auxiliary prediction. α represents the weight of the feature extraction variable, and (1-α) represents the weight of the auxiliary prediction variable. The weight α of the feature extraction variable is one of the hyperparameters of this hybrid model. The Optuna hyperparameter optimization framework is used to optimize the weight α. The number of variables after feature fusion is 2n+1.
[0069] Step 5: Combine the 2n+1 feature variables X at time t after fusion. comb The input to the XGBoost predictor consists of a weighted combination of n standardized original input features, n extracted features obtained by the Transformer feature extractor, and one auxiliary prediction result. The target value is the HCl concentration at time t+3. The calculation formula for the prediction result is as follows:
[0070]
[0071] In the formula, The predicted HCl concentration at time t+3 is β. y As the initial prediction baseline, η t Let K be the learning rate, K be the total number of decision trees, and x be the identifier of the z-th decision tree. comb_i f represents the fused feature variables obtained from step 4. z (x comb_i Let be the predicted output of the z-th tree. For the z-th tree, the prediction result is calculated using the following formula:
[0072] f z (x comb_i )=ω q(xcomb_i) (8)
[0073] In the formula, q(x) comb_i Let ω be the decision rule for the tree, mapping samples to leaf node indices, and ω be the leaf node weight vector. The hyperparameters of the XGBoost predictor are optimized using the Optuna hyperparameter optimization framework, and the final predicted HCl concentration at time t+3 is output after inverse standardization. The periodic patterns and temporal features extracted by the Transformer complement the nonlinear fitting ability of XGBoost, effectively improving the accuracy of multi-step advance prediction.
[0074] The Transformer feature extractor in this prediction method is a simplified Transformer architecture, which mainly includes an input layer, a fully connected projection layer, a positional encoding layer, an encoder, a global average pooling layer, and a dual output layer.
[0075] The encoder consists of N identical stacked layers, each with two sub-layers. The first sub-layer uses a multi-head attention mechanism, and the second sub-layer is a fully connected feedforward network. Residual connections and normalization are used to process the output after both sub-layers. The number of layers N of the encoder is one of the hyperparameters of the hybrid model. The number of layers N of the encoder is optimized using the Optuna hyperparameter optimization framework.
[0076] The overall loss function for Transformer feature extraction in this prediction method is:
[0077] loss = λ·loss feat +(1-λ)·loss apre (9)
[0078] In the formula, λ represents the weight of the feature extraction loss, (1-λ) represents the weight of the auxiliary prediction loss, and λ is the hyperparameter of the hybrid model. The weight λ is optimized using the Optuna hyperparameter optimization framework; loss feat The output of the Transformer feature extractor is the loss function for feature extraction, which is the mean absolute error (MAE). apre The output of the Transformer feature extractor is used as the loss for auxiliary prediction, and its loss function is the Huber loss.
[0079] The objective function of the XGBoost predictor in this prediction method is:
[0080]
[0081] In the formula, M is the objective function value; y i For the true value, For predicted values, For loss function, That is, the mean squared error (MSE), where k is the sample size; f z Let Ω(f) be the predicted output of the z-th tree. z ) represents the regularization term of the z-th decision tree, and K represents the total number of decision trees.
[0082] As a specific embodiment, this embodiment presents a method for predicting HCl concentration in a waste incinerator based on a hybrid model. This method utilizes historical operational data from the DCS system of a waste-to-energy plant in Guangdong Province. This historical operational data includes operational parameters such as primary air volume, secondary air volume, grate velocity, slurry pipe flow rate, and average air preheater outlet temperature, as well as historical values of HCl concentration. The sampling interval is 1 minute. Sixty-thousand consecutive data sets representing the samples are selected and divided into training, validation, and test sets in a 7:2:1 ratio. The HCl concentration is then predicted for the next three minutes.
[0083] Data preprocessing: Historical operational data is preprocessed, primarily including outlier handling and noise reduction. Outlier handling includes handling extreme outliers and regular outliers. First, extreme outliers, such as negative HCl concentrations due to temporary instrument malfunctions in harsh environments, are filtered out and directly removed. Then, after removing extreme outliers, the 3σ principle is used to detect other outliers, and these outliers are replaced with linear interpolation. Data noise reduction uses Gaussian smoothing. For the data point y at position i... i The value after Gaussian smoothing and noise reduction for:
[0084]
[0085] In the formula, d is the offset relative to the current position i, and h i G is the window radius. i (d) represents the Gaussian weights, determined by the Gaussian function. In this example, the window size is set to 9, and the correlation between the denoised data and the original data is 0.988.
[0086] Feature variable selection: The mutual information (MI) method was used to calculate the MI value between each operating parameter variable and the target HCl concentration. The calculation formula is as follows:
[0087]
[0088] In the formula, X represents the various operating parameters, Y represents the HCl concentration, p(x,y) is the joint probability distribution of X and Y, and p(x) and p(y) are the marginal probability distributions of X and Y, respectively. Variables with an MI value greater than 0.5 were selected, ultimately determining 17 variables as input variables for the model, including the reaction tower slurry pipe flow rate, primary air flow rate, secondary air flow rate, grate velocity, and the historical value of the current HCl concentration. The z-score standardization method was used to standardize the data, and its calculation formula is as follows:
[0089]
[0090] In the formula, x represents the original data, μ represents the mean of the data, and σ represents the standard deviation of the data.
[0091] Feature extraction: The 17 standardized feature variables are input into the Transformer feature extractor, which uses a position encoder to extract positional information to obtain temporal relationships.
[0092]
[0093] In the formula, pos is the time step position, pos∈[0,L-1], where L is the set sequence length; X pe [pos,2ζ] represents the sinusoidal positional code value of the encoding vector at position pos in the even-numbered dimension 2ζ, X pe [pos,2ζ+1] represents the cosine position code value of the encoding vector at position pos in the odd-numbered dimension 2ζ+1; d model Let ζ be the hidden layer dimension of the transformer model (obtained by Optuna optimization), and ζ be the dimension index, where ζ∈[0,Q]. model / 2-1], in this example d model After optimization using the Optuna hyperparameter optimization framework, the value was determined to be 248.
[0094] This Transformer feature extractor utilizes a multi-head attention layer to calculate the correlation weights between different positions in the input sequence, with the result of the m-th attention head being [head number missing]. m The calculation formula is as follows:
[0095]
[0096] In the formula, Q m Let K be the query matrix with the m-th head. m Let V be the key matrix of the m-th head. m Let be the value matrix of the m-th head, and Q, K, and V be the initial query, key, and value matrices obtained from the input sequence through linear transformation. Let K be the weight matrix. T This represents the transpose of the key matrix. This is the scaling factor. The final multi-head output:
[0097] MultiHead(Q,K,V)=Concat(head1,…,head h W O (18)
[0098] In the formula, h is the number of attention heads, and W O Here is the weight matrix. In this example, h is determined to be 3 after optimization using the Optuna hyperparameter optimization framework.
[0099] The 17 standardized feature variables were input into the Transformer feature extractor, and the hyperparameters of the model were optimized using the Optuna hyperparameter optimization framework. The Transformer feature extractor has two outputs: (1) the result F of feature extraction from the 17 standardized original input features; (2) the result P of auxiliary prediction of HCl concentration after 3 minutes. The formula for calculating the feature extraction result F is as follows:
[0100]
[0101] In the formula, L is the sequence length (10 in this example), H is the feature encoded by multiple transformers, and W... f b is a trainable weight matrix f Let x be the bias vector of the feature extraction layer, and swish be the self-gated activation function. linear ·σ(x linear ), where σ(x) linear ) represents the Sigmoid function, x linear This is the result of linear transformation and biasing of the global feature vector. The formula for calculating the auxiliary prediction result P is as follows:
[0102]
[0103] In the formula, W p b is the weight of the prediction layer. p It is a scalar bias vector.
[0104] The overall loss function of the Transformer feature extractor is:
[0105] loss = λ·loss feat +(1-λ)·loss apre (twenty one)
[0106] In the formula, λ represents the weight of the feature extraction loss, and (1-λ) represents the weight of the auxiliary prediction loss; loss feat The output of the Transformer feature extractor is the loss function for feature extraction, which is the mean absolute error (MAE). apre The output of the Transformer feature extractor serves as the loss for auxiliary prediction, and its loss function is the Huber loss. Here, λ is one of the hyperparameters of this hybrid model; in this example, the weight λ is optimized to 0.14 using the Optuna hyperparameter optimization framework.
[0107] This Transformer feature extractor is a simplified Transformer architecture, mainly consisting of an input layer, a fully connected projection layer, a positional encoding layer, an encoder, global average pooling, and a dual output layer. The encoder comprises N identical stacked layers, each with two sub-layers. The first sub-layer employs a multi-head attention mechanism, and the second sub-layer is a fully connected feedforward network. Residual connections and normalization are applied to the output after both sub-layers. The number of encoder layers, N, is one of the hyperparameters of this hybrid model. In this example, the number of encoder layers N was optimized to 2 using the Optuna hyperparameter optimization framework.
[0108] Feature fusion: The 17 standardized original input feature variables x′, the 17 extracted feature variables F output by the Transformer feature extractor, and 1 auxiliary prediction variable P are fused. The specific fusion method is as follows:
[0109]
[0110] In the formula, X comb Let x' represent the fused feature variables, where n is the number of original input features, k is the number of samples, and x'' is the number of samples. kn For the k-th data point of the n-th original input feature variable after standardization, F kn P is the k-th output result after extracting features from the n-th original input feature variable using the Transformer feature extractor. k This is the k-th output of the Transformer feature extractor for auxiliary prediction. α represents the weight of the feature extraction variable, and (1-α) represents the weight of the auxiliary prediction variable. The weight α of the feature extraction variable is one of the hyperparameters of this hybrid model. After optimization using the Optuna hyperparameter optimization framework, the weight α is determined to be 0.31. In this example, n is 17, and the final features are fused into 25 feature variables.
[0111] Model prediction: The 25 fused feature variables of the current time are input into the XGBoost predictor, with the target value being the HCl concentration 3 minutes later. The input features are a weighted combination of the standardized original input features, the feature extraction results from the Transformer feature extractor, and the auxiliary prediction results. The calculation formula for the prediction result is as follows:
[0112]
[0113] In the formula, The predicted HCl concentration after 3 minutes, β y As the initial prediction baseline, η t Let K be the learning rate, K be the total number of decision trees, z be the identifier of the z-th decision tree, and x be the learning rate.comb_i f is the fused feature vector obtained by the feature fusion module. z (x comb_i Let be the predicted output of the z-th tree. For the z-th tree, the prediction result is calculated using the following formula:
[0114] f z (x comb_i )=ω q(xcomb_i) (twenty four)
[0115] In the formula, q(x) comb_i Let ω be the decision rule for the tree, mapping samples to leaf node indices, and ω be the leaf node weight vector. The hyperparameters of the XGBoost predictor are optimized using the Optuna hyperparameter optimization framework. After inverse standardization, the final predicted HCl concentration value after 3 minutes is output, such as... Figure 3 As shown.
[0116] The objective function of the XGBoost predictor in this prediction method is:
[0117]
[0118] In the formula, M is the objective function value; y i For the true value, For predicted values, For loss function, That is, the mean squared error (MSE), where k is the sample size; f z Let Ω(f) be the predicted output of the z-th tree. z ) represents the regularization term of the z-th decision tree, and K represents the total number of decision trees.
[0119] The evaluation index uses the coefficient of determination R. 2 Mean square error (MSE) and mean absolute error (MAE), coefficient of determination R 2 The calculation formula is as follows:
[0120]
[0121] In the formula, y i ′ represents the true value before standardization. This is the final predicted value after destandardization. R represents the mean of the true values before standardization, and k is the sample size. 2 The closer the value is to 1, the better the model fit. The mean squared error (MSE) is calculated as follows:
[0122] The formula for calculating the Mean Absolute Error (MAE) is as follows:
[0123] Smaller MSE and MAE values indicate smaller prediction errors. This method was used to predict the HCl concentration after 3 minutes. Figure 3 As shown in (a), (b), and (c), the R of the training set 2 The R value for the validation set is 0.9888, the MSE is 0.0478, and the MAE is 0.1459; 2 The R value for the test set was 0.9698, the MSE was 0.1926, and the MAE was 0.2593; 2 The value is 0.9667, the MSE is 0.1899, and the MAE is 0.2867.
[0124] The above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not intended to limit the implementation of the present invention. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively describe all possible implementations here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A method for predicting HCl concentration in waste incinerators based on a hybrid model, characterized in that, Includes the following steps: Step 1: Obtain historical operating data of the waste incinerator, including various operating parameters, and perform data preprocessing on the historical operating data; Step 2: Calculate the MI value between each preprocessed operating parameter and the target HCl concentration using the MI method. Select multiple variables as input feature variables based on the MI value and standardize them to serve as the original input features for the feature extractor. Step 3: Input the original input features into the Transformer feature extractor, and output the feature extraction results and the HCl concentration-assisted prediction results; Step 4: Perform feature fusion on the original input features, feature extraction results, and HCl concentration-assisted prediction results; Step 5: Input the fused features into the XGBoost predictor to predict the HCl concentration, and then output the final predicted HCl concentration value after inverse standardization.
2. The method for predicting HCl concentration in a waste incinerator based on a hybrid model according to claim 1, characterized in that, In step 2, operating parameters with an MI value greater than 0.5 relative to the HCl concentration data are selected. The operating parameters with an MI value greater than 0.5 and the historical values of HCl concentration are then standardized using the z-score method and used as the original input features of the Transformer feature extractor.
3. The method for predicting HCl concentration in a waste incinerator based on a hybrid model according to claim 1, characterized in that, In step 3, the Transformer feature extractor includes an input layer, a fully connected projection layer, a positional encoding layer, an encoder, global average pooling, and a dual output layer. The encoder consists of N identical stacked layers, each with two sub-layers. The first sub-layer uses a multi-head attention mechanism, and the second sub-layer is a fully connected feedforward network. Residual connections and normalization are used to process the output after both sub-layers. The number of layers N in the encoder is one of the hyperparameters of this hybrid model.
4. The method for predicting HCl concentration in a waste incinerator based on a hybrid model according to claim 1, characterized in that, In step 3, the Transformer feature extractor outputs two results: the feature extraction result and the feature extraction result. and auxiliary prediction results The feature extraction results The calculation formula is as follows: (1) In the formula, For sequence length, The features at time t are encoded by multiple Transformer layers. For trainable weight matrix, This is the bias vector of the feature extraction layer. It is a self-gated activation function; Results of auxiliary prediction The calculation formula is as follows: (2) In the formula, For the prediction layer weights, It is a scalar bias vector.
5. The method for predicting HCl concentration in a waste incinerator based on a hybrid model according to claim 1, characterized in that, The overall loss function of the Transformer feature extractor is a weighted sum of the loss for feature extraction and the loss for auxiliary prediction, with the loss weights being one of the hyperparameters of the hybrid model.
6. The method for predicting HCl concentration in a waste incinerator based on a hybrid model according to claim 1, characterized in that, The feature fusion method in step 4 is as follows: (3) In the formula, Represents the feature variables after fusion. The number of original input features. For the sample size, For the standardized first The first original input feature variable One data point, For the Transformer feature extractor to the 1st The first feature extraction is performed on the original input feature variable. Each output result The first step to assist prediction for the Transformer feature extractor Each output result Weights for feature extraction variables, Weights are used to assist in predicting variables. It is one of the hyperparameters of the hybrid model.
7. The method for predicting HCl concentration in a waste incinerator based on a hybrid model according to claim 1, characterized in that, In step 5, the objective function of the XGBoost predictor includes a loss function and a regularization term, where the loss function uses mean squared error.
8. A system for implementing the method for predicting HCl concentration in a waste incinerator based on a hybrid model as described in claim 1, characterized in that, include: The data preprocessing module is used to preprocess historical operating data containing various operating parameters; The feature variable selection module is used to select original input feature variables from various operating parameters in historical operating data, calculate the MI value between each operating parameter and the target value HCl concentration using the MI method, and select multiple variables as input feature variables based on the MI value. The standardization module is used to standardize the preprocessed data, including standardization and destandardization. The selected input feature variables are standardized and used as the original input features of the feature extractor. The results of the prediction module are destandardized and used as the final predicted values. The feature extraction module uses the Transformer feature extractor to extract features and assist in prediction from the standardized raw input features and output the results. The feature fusion module is used to weight and fuse the original input features, the extracted features output by the Transformer feature extractor, and the auxiliary prediction results, and use them as the input to the XGBoost predictor. The model prediction module uses the XGBoost predictor to predict the HCl concentration after three time steps; The parameter optimization module is used to optimize hyperparameters, utilizing the Optuna hyperparameter optimization framework to find the best values for each parameter.
9. A computer device comprising a memory and a processor, the memory being electrically connected to the processor, the memory storing a computer program, characterized in that: When the computer program is executed by the processor, it causes the processor to implement the method as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor implements the method as described in any one of claims 1 to 8.