Urban solid waste incinerator temperature sensing method based on self-attention hybrid ensemble network

By using the AHENN-SA method, combined with MLE and self-attention mechanisms, the problem of modeling dynamic changes in the FT model during MSWI was solved, improving prediction accuracy and robustness, adapting to changes in multiple operating conditions, and achieving precise control and optimization of the FT.

CN122287283APending Publication Date: 2026-06-26BEIJING UNIV OF TECH +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING UNIV OF TECH
Filing Date
2025-06-18
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In existing MSWI processes, incinerator temperature (FT) models struggle to accurately capture dynamic changes, especially in complex nonlinear, multi-condition, and time-varying environments. Traditional methods are insufficient to meet modeling requirements, and conventional neural networks are susceptible to overfitting and lack the ability to model time-series features.

Method used

We employ a self-attention hybrid ensemble network (AHENN-SA) approach, combining AdaBoost and Bagging ensemble strategies. By using the moving block bootstrapping sampling method of MLE to preserve the temporal dependency features of the data, and introducing a self-attention mechanism, we improve the robustness and generalization performance of the model.

Benefits of technology

It significantly improves the accuracy and robustness of FT predictions, can adapt to dynamic changes in complex industrial environments, and provides better modeling accuracy and engineering adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122287283A_ABST
    Figure CN122287283A_ABST
Patent Text Reader

Abstract

The method for sensing the temperature of urban solid waste incinerators based on a self-attention hybrid ensemble network belongs to the field of industrial process modeling and intelligent sensing. This method addresses the highly nonlinear, multi-condition variations, and complex physical-chemical reaction mechanisms in the MSWI process. Combining an ensemble learning framework and improved neural network structure, it introduces moving block bootstrap (MBB) sampling based on maximum likelihood estimation (MLE), a self-attention mechanism, and a hybrid ensemble strategy fusing AdaBoost and Bagging to achieve accurate modeling and dynamic prediction of the Fourier transform (FT). This results in good modeling accuracy, generalization ability, and engineering adaptability. Experimental verification under multiple benchmark problems and actual MSWI conditions demonstrates that the proposed FT sensing method exhibits superior performance in both prediction accuracy and robustness, possessing significant engineering application potential and widespread value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial process modeling and intelligent sensing in municipal solid waste incineration (MSWI), specifically to a furnace temperature (FT) sensing method based on an adaptive hybrid ensemble neural network (AHENN-SA) with a self-attention mechanism. This method addresses the highly nonlinear, multi-condition variations, and complex physicochemical reaction mechanisms in MSWI processes. By combining an ensemble learning framework with improved neural network structures, it introduces moving block bootstrap (MBB) sampling based on maximum likelihood estimation (MLE), a self-attention mechanism, and a hybrid ensemble strategy fusing AdaBoost and Bagging. This achieves accurate modeling and dynamic prediction of FT, exhibiting good modeling accuracy, generalization ability, and engineering adaptability. Experimental verification under multiple benchmark problems and actual MSWI conditions demonstrates that the proposed FT sensing method exhibits superior performance in both prediction accuracy and robustness, possessing significant engineering application potential and widespread value. Background Technology

[0002] With the acceleration of global urbanization and the continuous growth of population, the production of municipal solid waste (MSW) is rising steadily, becoming a significant factor affecting urban environmental quality and public health. Statistics show that the global annual MSW production exceeds 2.1 billion tons, and is projected to increase to 3.8 billion tons by 2050, with an average annual growth rate of approximately 3.4%. Asia accounts for about half of the global MSW production. Efficient and safe disposal of MSW has become a key task for sustainable urban development.

[0003] Currently, MSWI technology, due to its high volume reduction efficiency, short processing cycle, small footprint, and large energy recovery potential, is gradually replacing traditional landfill and composting technologies, becoming one of the mainstream treatment methods widely used globally. According to data from the International Solid Waste Association, approximately 15% of MSW globally is treated by incineration, with the proportion exceeding 45% in European countries and 30%-35% in Asia and North America. To promote the development of MSWI technology, many countries have successively introduced relevant laws and regulations, supplemented by financial subsidies and tax incentives, to encourage the construction and intelligent transformation of efficient and environmentally friendly incineration facilities.

[0004] In the MSWI process, the fuel flow factor (FT) is a key control variable directly affecting combustion efficiency, energy recovery level, and pollutant emission concentration. The variation of FT is influenced by multiple operating conditions, such as primary and secondary air flow parameters and the speed control of multiple grate sections. Furthermore, the complex and variable composition of the MSW, with significant fluctuations in moisture content and calorific value, leads to significant nonlinearity, time-varying nature, and stochasticity in the FT dynamics, making modeling challenging. Traditional linear modeling methods struggle to accurately capture the evolution of FT, while conventional neural network models are susceptible to overfitting and lack effective modeling capabilities for time-series characteristics, failing to meet the modeling needs of complex industrial environments.

[0005] To address this, this invention proposes a Fourier Transform (FT) perception method based on AHENN-SA. This method integrates AdaBoost and Bagging ensemble strategies to construct a hybrid ensemble neural network structure, improving robustness and generalization performance while maintaining the diversity of base learners (BLs) and the overall model stability. To enhance the temporal structure and representativeness of the training data, a moving block bootstrapping sampling method based on maximum likelihood estimation is designed to effectively preserve the temporal dependency features in the data. Furthermore, a self-attention mechanism is introduced, and the attention score of the suboptimal learner (SL) is trained through a feedforward neural network (FNN) to improve the SL's responsiveness to important information, thereby further enhancing the modeling accuracy and adaptability for the dynamic changes in the Fourier transform. Summary of the Invention

[0006] This invention designs an AHENN-SA-based Fourier Transform (FT) model for the FT-sensing prediction problem in MSWI processes. First, a MLE-based MBB sampling method is used to select block nodes while maintaining the temporal continuity of the data, reducing disruption to the time series and improving the model's generalization performance on real-world data. Second, by combining Bagging and AdaBoost ensemble strategies, excessive attention to difficult samples is balanced, and the diversity of BL (Blank-Based Learning) is increased by fully utilizing different subsets of datasets, thus ensuring the model's robustness under various data distributions. Finally, an FT neural network is introduced into the training process of the Service Level (SL). The dynamic changes of SL features are captured by training the attention scores of each SL, and these features are systematically integrated using the attention scores and weight distributions obtained after training, improving the model's dynamic performance. This invention significantly improves the prediction accuracy of the FT model, laying the model foundation for precise control and optimization of FT, and providing a reliable solution for environmental protection in incineration processes.

[0007] The present invention adopts the following technical solution and implementation steps:

[0008] FT model based on AHENN-SA

[0009] (1) A Radial Basis Function (RBF) neural network is used as the base learner in the hybrid ensemble neural network. It has a typical three-layer structure, including an input layer, a hidden layer, and an output layer. The input layer is responsible for receiving external feature signals and passing them to the hidden layer. The hidden layer consists of multiple RBF neurons, using a Gaussian function as the activation function. By calculating the Euclidean distance between the input vector and each center point, a nonlinear mapping from the input space to the feature space is achieved, thereby enhancing the model's ability to express complex patterns. The output layer performs weighted integration of the response results from the hidden layer to generate the final output of the network. This structure can effectively capture the nonlinear relationship between input features and target variables, providing a basic unit with strong learning capabilities for the ensemble model.

[0010] Choose the Gaussian function θ j (t) is the activation function of the hidden layer of BL.

[0011]

[0012] The output y of each BL out (t) is as follows:

[0013]

[0014] Where x(t)=[x1(t),x2(t),...x n [t] is the input vector of the neural network, carrying a set of samples containing all input features at the current moment. (n=4, indicating that the number of selected input features is 4. The inputs to the neural network are primary total air, secondary total air, grate velocity in the drying section, and grate velocity in the incineration section.) c j (t) represents the center vector of the hidden layer; ||x(t)-c j (t)|| represents the Euclidean distance from the input sample at the current time to the center of the j-th hidden layer neuron; σ j w represents the width of the j-th hidden layer neuron; J is the total number of hidden layer neurons; j This represents the connection weight from the j-th hidden layer node to the output layer.

[0015] (2) The MBB method based on MLE is adopted to preserve the temporal dependency structure of the data in the subsample set, thereby reducing the destruction of temporal characteristics by traditional Bootstrap sampling and improving the accuracy and reliability of modeling.

[0016] The algorithm flow of MBB is as follows:

[0017] Table 1. MBB Algorithm Flow

[0018]

[0019]

[0020] ① The size of each data block is determined by MLE.

[0021] The size of each block is determined by MLE, and each block is numbered. The size of each block should be large enough to contain the important temporal dependency structure in the time series data, but not too large to avoid losing the diversity of the samples.

[0022] ②Bootstrap Sampling Extraction Block

[0023] Blocks are randomly drawn from the original dataset with replacement until a Bootstrap sample of the same length as the original dataset is constructed. Because sampling is done with replacement, some blocks may be drawn multiple times, while others may not be drawn at all.

[0024] ③ Connecting block

[0025] The extracted blocks are concatenated in chronological order within the original data sequence to form a Bootstrap subset.

[0026] ④ Handling boundary effects

[0027] Since the block size may not be an integer multiple of the original data sequence length, some extra data points may be generated when concatenating blocks. To fill these gaps, some extra data points are removed from the beginning or end of the Bootstrap sample.

[0028] The MLE algorithm process is as follows:

[0029] Table 2. MLE Algorithm Flow

[0030]

[0031]

[0032] The algorithm flow and more detailed formula and variable explanations in Table 2 are as follows:

[0033] Let D = {(x1,y1),(x2,y2),...,(x N ,y N )}, y=[y1,...y N ] TMLE is used to detect mutation points in y, and these mutation points are used to break the data, thus dividing D into θ sub-blocks. y represents the desired output, N represents the total number of samples, n represents the sequence number of the mutation point, and D is the set recording each set of input and output samples. ε is the pre-set minimum sub-block length, which is set to 15 in this invention. θ is the maximum number of mutation points to be found. To round down, [] T Represents the transpose of a matrix or vector.

[0034] ① Define the model

[0035] We use the normal distribution function to approximate the local statistical characteristics before and after the mutation point, that is, we expect the output y to follow a normal distribution with different mean and variance before and after the mutation point t. Let the data before the mutation point t follow a normal distribution. After the mutation point t, it obeys ε+1 and Ω-ε+1 represent the data points that need to be traversed in each data segment, and Ω represents the total number of observed data points. Here, μ1 represents the mean of the expected output y before the mutation point t; μ1 represents the variance of the expected output y before the mutation point t; μ2 represents the mean of the expected output y after the mutation point t. This represents the mean of the expected output y after the mutation point t;

[0036] ② Calculate the likelihood function

[0037] For a given mutation point t, the likelihood function L(t) of the data is:

[0038]

[0039] L(t) represents the observed data {x1, x2, ..., x} when the mutation point is located at time t. Ω The total likelihood function of x i This represents the observed data value at time point i, where i is the index variable of the observed sample point, used to traverse the data sequence. f(x) i ;μ;σ 2 ) represents the variable x i With a mean of μ and a variance of σ 2 The probability density function of Ω under a normal distribution, where t is the hypothetical location of the mutation point, and its value ranges from 1 to t < Ω.

[0040]

[0041] ③ Find the log-likelihood function

[0042] Taking the logarithm of the likelihood function to simplify the calculation, we obtain the log-likelihood function:

[0043]

[0044] logL(t) represents the log-likelihood function value given the position t of the mutation point.

[0045] Further simplification yields:

[0046]

[0047] ④ Parameter estimation

[0048] For each possible mutation point t, estimate the parameters μ1, μ2,

[0049]

[0050] Let represent the sample mean estimate of the observed data up to and including the t-th point; This represents an estimate of the sample mean of the observed data after the mutation point t. This represents the sample variance estimate of the observation data before the mutation point t; This represents the sample variance estimate of the observed data after the mutation point t. They are respectively The abbreviation of .

[0051] ⑤ Maximize the log-likelihood function

[0052] For each possible mutation point t, the estimated parameters are substituted into the log-likelihood function:

[0053]

[0054] By iterating through all possible mutation points t, calculating the log-likelihood function value for each t, ​​and selecting the t with the largest log-likelihood function value as the optimal mutation point:

[0055]

[0056] This represents an estimate of the location of the optimal mutation point. This represents the parameter values ​​that maximize the objective function logL(t) on variable t. Each saved optimal mutation point is recorded as t = t1, t2, ... t. n , t n Let θ represent the nth mutation point, where 1 ≤ n ≤ θ. Taking the case where the number of optimal mutation points found equals the maximum number of mutation points required (i.e., n = θ) as an example...

[0057]

[0058] At this point, the dataset has been divided into θ sub-blocks. Further MBB sampling yields a set of subsamples Φ = {D1, D2, ..., D...}. T}. T represents T sub-training datasets used to train SL; D represents one of the partitioned training datasets in set Φ, containing θ subsets of samples; Let t represent the i-th subset, where i = 1, ..., θ; i Represents the split point (t1, t2, ..., t) of each subset. θ-1 (Only reflects the index of the input feature vector), x i y represents the input feature vector of the i-th sample; N is the total number of original training samples; i Represents the relationship with input feature x i The corresponding target output value.

[0059] (3) The AdaBoost algorithm and the Bagging algorithm are combined to design a hybrid ensemble algorithm to solve the problem of AdaBoost algorithm's over-focus on difficult samples and improve the diversity of base learners.

[0060] For each subset D m Train an SL in parallel and independently, and denote its output as SL. m Let m = 1, ..., T, and let AdaBoost be used to train each SL. For the first subset D corresponding to each BL... m,j j=1, equal distribution D m,j The weight of each sample.

[0061] Set the threshold matrix after each training session: Δ j =[Δ j,1 ,Δ j,2 ,...Δ j,N This is used to determine the learning effectiveness of each BL (Broadcast Learning). j Let Δ represent the threshold matrix of the j-th BL. j,k This represents the threshold corresponding to the k-th sample in the threshold matrix of the j-th BL. c is a constant used to adjust the prediction accuracy. T is the preset number of SLs, and U is the number of BLs in each SL. In this invention, we set c = 0.03, T = 4, and U = 3.

[0062] Table 3. Bagging-AdaBoost Hybrid Integration Algorithm Flow

[0063]

[0064]

[0065] The algorithm flow and more detailed explanations of the formula variables in Table 3 are as follows:

[0066] Initialize the weights of each sample in the subsample set

[0067] W j =[w j,1 ,...,w j,N ], w j,k =1N, j=1,...,U, k=1,...,N(14)

[0068] W j Let w be the sample weight matrix for the j-th BL. j,k Let F represent the weight corresponding to the k-th sample of the j-th BL. Training the j-th BL yields F. j (x):x→y,F j (x) represents each trained BL.

[0069] Calculate the regression error for each sample.

[0070]

[0071] y k This represents the expected output of the k-th sample. Let e ​​represent the predicted output of the k-th sample learned by the j-th BL. j,k The prediction error of the k-th sample learned by the j-th BL is represented by

[0072] Design logic matrix

[0073] L j =[L j,1 ,...,L j,N (16)

[0074] This is used to record the training performance of each sample in the j-th BL, where L j,k The position where =1 represents a poorly trained sample, L j,k The sample at position 0 is the best-trained sample, and k = 1, ..., N.

[0075] Calculate the regression error rate E of the j-th BL. j

[0076]

[0077] Calculate the weight coefficient α of the j-th BL. j

[0078]

[0079] Design feature matrix H j

[0080] H j =[h j,1 ,...,h j,N (19)

[0081] H j The elements in the only numbers are -1 and 1. If e j,k >Δ j,k Then h j,k Assigning a value of -1 increases the weight of the k-th sample in the j-th base learner; conversely, it increases the weight of h. j,k The value is set to 1, thereby reducing the sample weight.

[0082] Calculate the normalization factor

[0083]

[0084] Υ j This is a normalization factor used to ensure that the sum of the weights of all samples in the sample set is 1.

[0085] W j Updated to W j+1 and through W j+1 D m,j Updated to D m,j+1

[0086]

[0087] W j+1 Consider W as the probability distribution of each sample being drawn from the new sample set. j+1 The sum of all elements in D is 1. m,j According to the weight matrix W j+1 Weighted sampling yields a new sample set D m,j+1 This is used to train the next BL.

[0088] Calculate the output of the j-th BL.

[0089] BL j =α j F j (x) (22) The output of each SL is:

[0090]

[0091] SL m Let be the output of the m-th service flow (SL). Finally, calculate the attention score for each SL. mThe attention score is calculated for the m-th SL, and the attention weights w are computed using a softmax probabilistic network. m By weighting and combining all learning samples (SLs), a more powerful ensemble learner is output, i.e. Figure 4 The strong learner in the network is m = 1, ..., T. The core of the softmax probabilistic network is to use the softmax function to convert the model's output into a probability distribution. Its formula is shown in Equation (26).

[0092] (4) A self-attention (SA) mechanism is designed to replace the traditional simple averaging method. This mechanism comprehensively considers the training performance of each training sample (SL) to allocate weights, thus more rationally integrating the modules. The output of each SL is treated as sequence data. The attention score and weight of each SL to the training samples are calculated through the SA mechanism, and these prediction results are linearly combined. The definitions of Q, K, and V in the proposed SA mechanism are as follows:

[0093] The feature representation of the training samples is used to evaluate the attention of each SL to the training sample set.

[0094] The output of each SL is compared with Q to determine the importance of each SL.

[0095] The prediction result of each SL for the training samples, i.e., the output of the SL, V = K, K = [k1,...,k T V = [v1,...,v] T ], k m =v m =SL m =[y1,...,y N ], where the subscript m represents the index of the m-th suboptimal learner.

[0096] Step 1: Calculate the attention score AS:

[0097]

[0098] In this invention, i = 1, ..., n, n = 4 represents the dimension of the sample, which are the four inputs of the neural network. AS is the comprehensive score of each SL across all dimensions of the training samples, and a m For the attention score of the m-th SL, w m Let K be the self-attention weight of the m-th SL. i and Q i Let K and Q represent the i-th dimensions of matrices K and Q, respectively, corresponding to the i-th input feature. i T Representation matrix K i The transpose of .

[0099] Step 2: Calculate the attention weight AW:

[0100] AW = [w1,...,w T = softmax(AS) (25)

[0101] The softmax probabilistic network transforms the attention score into a probability distribution, or attention weights, where each element has a value between 0 and 1, and the sum of all elements is 1.

[0102]

[0103] in K represents i The dimension of K is used as a scaling factor. In high-dimensional space, the dot product can become extremely large, causing gradient vanishing or exploding. Therefore, to ensure K... i With a stable range of attention scores, scaling is essential.

[0104] Step 3: Calculate the weighted output SL' for each SL. m :

[0105] SL' m =w m v m (27)

[0106] Step 4: Sum the weighted outputs of all SLs to obtain an integrated output X, and use it as the input of the FNN:

[0107]

[0108] (5) Training self-attention scores using FNN. The FNN introduced in this paper is a three-layer network structure, including an input layer, a hidden layer, and an output layer. The ReLU function is used as the activation function of the hidden layer. The ReLU function is explained in Equation (39). w1 and b1 represent the weights and biases of the input layer neurons, respectively; Z1 is the output of the input layer neurons; A1 is the ReLU activation function of the hidden layer; w2 represents the weights and biases of the hidden layer neurons, respectively; Z2 is the output of the hidden layer neurons.

[0109] The forward propagation process is as follows:

[0110] Input layer to hidden layer:

[0111] Z1=w1X+b1 (29)

[0112] Hidden layer:

[0113] A1 = ReLU(Z1)(30)

[0114] Hidden layer to output layer:

[0115] Z2=w2A1+b2(31)

[0116] Output layer:

[0117] y pred =Z2(32)

[0118] Define the loss function:

[0119]

[0120] To predict the output, y i This is the expected output.

[0121] The backpropagation process is as follows: (Note: T (Represents the transpose of a vector or matrix)

[0122] Output layer to hidden layer:

[0123]

[0124] Hidden layer to input layer:

[0125]

[0126]

[0127] ⊙ represents the Hadamard product of a matrix.

[0128] From input layer to attention layer:

[0129]

[0130] During each training iteration, the five parameters w1, b1, w2, b2, and K are updated along the negative gradient direction of the loss function.

[0131]

[0132] Where η∈(0,1) is the learning rate of the FNN, and the value of K at the end of training is denoted as K'. Steps 1 to 4 are repeated with V'=K' to obtain the training attention weights w′1,...,w'. T And the final integrated output: ELO

[0133]

[0134] The inventiveness of this invention is mainly reflected in:

[0135] (1) In view of the complex background of MSWI process data with nonlinearity, multiple operating conditions and significant temporal characteristics, this invention proposes an FT perception method based on AHENN-SA, which effectively improves the accuracy and robustness of FT prediction. This method introduces a self-attention mechanism to dynamically mine the feature differences of sub-models, realizes the difference enhancement and information fusion among multiple learners, and thus significantly improves the modeling performance and generalization ability.

[0136] (2) This invention designs a training sample construction method that combines the MLE algorithm with block bootstrap, which can adaptively identify sequence mutation points and complete reasonable data segmentation accordingly, effectively maintaining the structural continuity and physical consistency of the time series. While maintaining the original sequence information structure, this method improves the model's adaptability to non-stationary data and solves the problem that traditional sample construction methods severely damage time series features.

[0137] (3) This invention innovatively integrates two ensemble learning strategies, Bagging and AdaBoost. On the one hand, it enhances the stability of the model under different data distributions by resampling samples, and on the other hand, it suppresses excessive attention to abnormal or difficult samples, thereby improving the overall robustness and anti-interference ability of the model.

[0138] (4) This invention introduces FNN training SL and combines it with an attention mechanism for dynamic feature representation and learner weighting, which effectively improves the adaptability of the ensemble model to complex dynamic processes. By calculating attention scores, differentiated modeling of the contribution of different sub-models is achieved, enabling the overall network to have stronger dynamic feature extraction and decision-making capabilities. Attached Figure Description

[0139] Figure 1 This is a structural diagram of the BL of the present invention.

[0140] Figure 2 This is a structural diagram of the AHENN-SA FT model of the present invention.

[0141] Figure 3 This is the RMSE curve of each training level during the training phase of this invention.

[0142] Figure 4 This is a graph showing the fitting results of the FT model output and the expected output during the testing phase of this invention.

[0143] Figure 5 This is a graph showing the fitting error between the FT model output and the expected output during the testing phase of this invention. Detailed Implementation

[0144] This invention designs an AHENN-SA-based Fourier Transmission (FT) method to address the problems of low accuracy, complex data distribution, and significant dynamic features in the MSWI process. By constructing a multi-level, dynamically responsive modeling framework, this invention effectively improves the modeling capability for complex time-series data and enhances the model's stability and robustness. First, a moving block bootstrapping algorithm based on maximum likelihood estimation is employed to identify abrupt changes in the sequence and perform reasonable data segmentation, constructing a representative subset of training samples while maintaining temporal continuity. This method ensures that the model can extract key structural features from non-stationary, multi-condition data. Second, Bagging and AdaBoost are integrated, enhancing the diversity and synergistic performance of BL through dual enhancements in both the sample and error dimensions. Bagging enhances model stability through resampling, while AdaBoost guides the model to focus on unpredictable samples, thereby improving overall model performance. Next, a self-attention mechanism is introduced into the Fourier Transmission Neural Network (FNN) to construct a Service Level Filter (SL), enabling the modeling and differentiation of the SL's dynamic features. The self-attention mechanism dynamically adjusts the output weights of the SL based on its performance on different sub-data sets, thereby achieving fine characterization and ensemble optimization of the differences between SLs. The AHENN-SA model constructed using the above method possesses strong dynamic modeling capabilities, adapts to data characteristics under multiple operating conditions, and can effectively cope with the non-stationary fluctuations of the Fourier Transmission (FT). Experimental results show that the model exhibits significant advantages in both baseline time series prediction and FT sensing prediction tasks under actual MSWI operating conditions, providing technical support for intelligent control and environmental protection of MSWI processes.

[0145] The present invention adopts the following technical solution and implementation steps:

[0146] (1) A Radial Basis Function (RBF) neural network is used as the base learner in the hybrid ensemble neural network. It has a typical three-layer structure, including an input layer, a hidden layer, and an output layer. The input layer is responsible for receiving external feature signals and passing them to the hidden layer. The hidden layer consists of multiple RBF neurons, using a Gaussian function as the activation function. By calculating the Euclidean distance between the input vector and each center point, a nonlinear mapping from the input space to the feature space is achieved, thereby enhancing the model's ability to express complex patterns. The output layer performs weighted integration of the response results from the hidden layer to generate the final output of the network. This structure can effectively capture the nonlinear relationship between input features and target variables, providing a basic unit with strong learning capabilities for the ensemble model.

[0147] Choose the Gaussian function θ j (t) is the activation function of the hidden layer of BL.

[0148]

[0149] The output y of each BL out (t) is as follows:

[0150]

[0151] Where x(t)=[x1(t),x2(t),...x n [t] is the input vector of the neural network, carrying a set of samples containing all input features at the current moment. (n=4, indicating that the number of selected input features is 4. The inputs to the neural network are primary total air, secondary total air, grate velocity in the drying section, and grate velocity in the incineration section.) c j (t) represents the center vector of the hidden layer; ||x(t)-c j (t)|| represents the Euclidean distance from the input sample at the current time to the center of the j-th hidden layer neuron; σ j w represents the width of the j-th hidden layer neuron; J is the total number of hidden layer neurons; j This represents the connection weight from the j-th hidden layer node to the output layer.

[0152] (2) The MBB method based on MLE is adopted to preserve the temporal dependency structure of the data in the subsample set, thereby reducing the destruction of temporal characteristics by traditional Bootstrap sampling and improving the accuracy and reliability of modeling.

[0153] The algorithm flow of MBB is as follows:

[0154] Table 1. MBB Algorithm Flow

[0155]

[0156] ① The size of each data block is determined by MLE.

[0157] The size of each block is determined by MLE, and each block is numbered. The size of each block should be large enough to contain the important temporal dependency structure in the time series data, but not too large to avoid losing the diversity of the samples.

[0158] ②Bootstrap Sampling Extraction Block

[0159] Blocks are randomly drawn from the original dataset with replacement until a Bootstrap sample of the same length as the original dataset is constructed. Because sampling is done with replacement, some blocks may be drawn multiple times, while others may not be drawn at all.

[0160] ③ Connecting block

[0161] The extracted blocks are concatenated in chronological order within the original data sequence to form a Bootstrap subset.

[0162] ④ Handling boundary effects

[0163] Since the block size may not be an integer multiple of the original data sequence length, some extra data points may be generated when concatenating blocks. To fill these gaps, some extra data points are removed from the beginning or end of the Bootstrap sample.

[0164] The MLE algorithm process is as follows:

[0165] Table 2. MLE Algorithm Flow

[0166]

[0167] Let D = {(x1,y1),(x2,y2),...,(x N ,y N )}, y=[y1,...y N ] T MLE is used to detect mutation points in y, and these mutation points are used to break the data, thus dividing D into θ sub-blocks. y represents the desired output, N represents the total number of samples, n represents the sequence number of the mutation point, and D is the set recording each set of input and output samples. ε is the pre-set minimum sub-block length, which is set to 15 in this invention. θ is the maximum number of mutation points to be found. To round down, [] T Represents the transpose of a matrix or vector.

[0168] ① Define the model

[0169] We use the normal distribution function to approximate the local statistical characteristics before and after the mutation point, that is, we expect the output y to follow a normal distribution with different mean and variance before and after the mutation point t. Let the data before the mutation point t follow a normal distribution. After the mutation point t, it obeys ε+1 and Ω-ε+1 represent the data points that need to be traversed in each data segment, and Ω represents the total number of observed data points. Here, μ1 represents the mean of the expected output y before the mutation point t; μ1 represents the variance of the expected output y before the mutation point t; μ2 represents the mean of the expected output y after the mutation point t. This represents the mean of the expected output y after the mutation point t;

[0170] ② Calculate the likelihood function

[0171] For a given mutation point t, the likelihood function L(t) of the data is:

[0172]

[0173] L(t) represents the observed data {x1, x2, ..., x} when the mutation point is located at time t. Ω The total likelihood function of x i This represents the observed data value at time point i, where i is the index variable of the observed sample point, used to traverse the data sequence. f(x) i ;μ;σ 2 ) represents the variable x i With a mean of μ and a variance of σ 2 The probability density function of Ω under a normal distribution, where t is the hypothetical location of the mutation point, and its value ranges from 1 to t < Ω.

[0174]

[0175] ③ Find the log-likelihood function

[0176] Taking the logarithm of the likelihood function to simplify the calculation, we obtain the log-likelihood function:

[0177]

[0178] logL(t) represents the log-likelihood function value given the position t of the mutation point.

[0179] Further simplification yields:

[0180]

[0181] ④ Parameter estimation

[0182] For each possible mutation point t, estimate the parameters μ1, μ2,

[0183]

[0184] Let represent the sample mean estimate of the observed data up to and including the t-th point; This represents an estimate of the sample mean of the observed data after the mutation point t. This represents the sample variance estimate of the observation data before the mutation point t; This represents the sample variance estimate of the observed data after the mutation point t. They are respectively The abbreviation of .

[0185] ⑤ Maximize the log-likelihood function

[0186] For each possible mutation point t, the estimated parameters are substituted into the log-likelihood function:

[0187]

[0188] By iterating through all possible mutation points t, calculating the log-likelihood function value for each t, ​​and selecting the t with the largest log-likelihood function value as the optimal mutation point:

[0189]

[0190] This represents an estimate of the location of the optimal mutation point. This represents the parameter values ​​that maximize the objective function logL(t) on variable t. Each saved optimal mutation point is recorded as t = t1, t2, ... t. n , t n Let θ represent the nth mutation point, where 1 ≤ n ≤ θ. Taking the case where the number of optimal mutation points found equals the maximum number of mutation points required (i.e., n = θ) as an example...

[0191]

[0192] At this point, the dataset has been divided into Δ sub-blocks. Further MBB sampling yields a set of subsamples φ = {D1, D2, ..., D...}. T}. T represents T sub-training datasets used to train SL; D represents the partitioned total training data set, containing θ subsets of samples; Let t represent the i-th subset, where i = 1, ..., θ; i This represents the split point of each subset, where N is the total number of original training samples; x i y represents the input feature vector of the i-th sample; i Represents the relationship with input feature x i The corresponding target output value.

[0193] (3) The AdaBoost algorithm and the Bagging algorithm are combined to design a hybrid ensemble algorithm to solve the problem of AdaBoost algorithm's over-focus on difficult samples and improve the diversity of base learners.

[0194] For each subset D m Train an SL in parallel and independently, and denote its output as SL. m Let m = 1, ..., T, and let AdaBoost be used to train each SL. For the first subset D corresponding to each BL... m,j j = 1 , equally distributed D m,j The weight of each sample.

[0195] Set the threshold matrix after each training session: Δ j =[Δ j,1 ,Δ j,2 ,...Δ j,N [], used to determine the learning effectiveness of each BL. c is a constant used to adjust the prediction accuracy. In this invention, c = 2.

[0196] Table 3. Bagging-AdaBoost Hybrid Integration Algorithm Flow

[0197]

[0198] Initialize the weights of each sample in the subsample set

[0199] W j =[w j,1 , ... ,w j,N ], W j,k =1N, j=1, ... U, k=1 ... ,N(14)

[0200] Training the j-th BL yields F j (x):x→y, F j (x) represents each trained BL.

[0201] Calculate the regression error for each sample.

[0202]

[0203] Design logic matrix

[0204] L j =[L j,1 , ... ,L j,N (16)

[0205] Calculate the regression error rate of the j-th BL.

[0206]

[0207] Calculate the weight coefficient of the j-th BL.

[0208]

[0209] Design feature matrix H j

[0210] H j=[h j,1 ,...,h j,N (19)

[0211] Calculate the normalization factor

[0212]

[0213] Update W j →W j+1 ,

[0214]

[0215] W j+1 Consider W as the probability distribution of each sample being drawn from the new sample set. j+1 The sum of all elements in D is 1. m,j Weighted sampling yields a new sample set D m,j+1 This is used to train the next BL.

[0216] Calculate the output of the j-th BL.

[0217] BL j =α j F j (x) (22)

[0218] The output of each SL is:

[0219]

[0220] Finally, the attention score a for each SL is calculated. m Attention weights w are calculated using a softmax probabilistic network. m By weighting and combining all learning samples (SLs), a more powerful ensemble learner is output, i.e. Figure 4 The strong learner in the algorithm. U is the number of BLs in each SL, m = 1,...,T

[0221] (4) A self-attention (SA) mechanism is designed to replace the traditional simple averaging method. This mechanism comprehensively considers the training performance of each training sample (SL) to allocate weights, thus more rationally integrating the modules. The output of each SL is treated as sequence data. The attention score and weight of each SL to the training samples are calculated through the SA mechanism, and these prediction results are linearly combined. The definitions of Q, K, and V in the proposed SA mechanism are as follows:

[0222] The feature representation of the training samples is used to evaluate the attention of each SL to the training sample set.

[0223] The output of each SL is compared with Q to determine the importance of each SL.

[0224] The prediction result of each SL for the training samples, i.e., the output of the SL, V = K, K = [k1,...,k T V = [v1,...,v] T ], k m =v m =SL m =[y1,...,y N ]

[0225] Step 1: Calculate the attention score AS:

[0226]

[0227] i = 1, ..., n, where n is the dimension of the sample, and AS is the comprehensive score of each SL across all dimensions of the training samples.

[0228] Step 2: Calculate the attention weight AW:

[0229] AW = [w1,...,w T = softmax(AS) (25)

[0230] The softmax probabilistic network transforms the attention score into a probability distribution, or attention weights, where each element has a value between 0 and 1, and the sum of all elements is 1.

[0231]

[0232] in K represents i The dimension of K is used as a scaling factor. In high-dimensional space, the dot product can become extremely large, causing gradient vanishing or exploding. Therefore, to ensure K... i With a stable range of attention scores, scaling is essential.

[0233] Step 3: Calculate the weighted output for each SL:

[0234] SL' m =w m v m (27)

[0235] Step 4: Sum the weighted outputs of all SLs to obtain an integrated output X, and use it as the input of the FNN:

[0236]

[0237] (5) Training self-attention scores using FNN. The FNN introduced in this paper is a three-layer network structure, including an input layer, a hidden layer, and an output layer. The ReLU function is used as the activation function for the hidden layer.

[0238] The forward propagation process is as follows:

[0239] Input layer to hidden layer:

[0240] Z1=w1X+b1(29)

[0241] Hidden layer:

[0242] A1 = ReLU(Z1)(30)

[0243] Hidden layer to output layer:

[0244] Z2=w2A1+b2(31)

[0245] Output layer:

[0246] y pred =Z2(32)

[0247] Define the loss function:

[0248]

[0249] The backpropagation process is as follows:

[0250] Output layer to hidden layer:

[0251]

[0252] Hidden layer to input layer:

[0253]

[0254] From input layer to attention layer:

[0255]

[0256]

[0257] During each training iteration, the five parameters w1, b1, w2, b2, and K are updated along the negative gradient direction of the loss function.

[0258]

[0259] Where η∈(0,1) is the learning rate of the FNN, and the value of K at the end of training is denoted as K'. Steps 1 to 4 are repeated with V'=K' to obtain the training attention weights w′1,...,w'. T And the final integrated output: ELO

[0260]

Claims

1. A method for sensing furnace temperature in urban solid waste incineration processes based on an adaptive hybrid ensemble neural network with a self-attention mechanism, characterized in that, Includes the following steps: (1) Determine the manipulated variables: Select some process variables that are most closely related to the furnace temperature (FT) during the MSWI process of urban solid waste incineration as input variables of the adaptive hybrid ensemble neural network (AHENN-SA) based on the self-attention mechanism; these variables include the primary air volume of the drying section, the primary air volume of the combustion section, the primary average velocity of the grate in the drying section, and the primary average velocity of the combustion section. (2) Radial basis function neural networks are selected as the base learners (BLs) in the adaptive hybrid ensemble neural network; wherein, the main structure of each BL consists of three layers, including an input layer, a hidden layer and an output layer; wherein, the input layer is used to receive external feature information, the hidden layer processes the input data through nonlinear transformation, and the output layer is used to generate the final prediction result; wherein, the hidden layer uses a Gaussian function as the activation function; Choose the Gaussian function θ j (t) as the activation function of BL The output y of each BL out (t) is as follows: Where x(t)=[x1(t),x2(t),...x n [t] is the input vector of the neural network, carrying a set of samples containing all input features at the current moment; n=4, indicating that the number of selected input features is 4; the inputs of the neural network are primary total air, secondary total air, grate velocity in the drying section, and grate velocity in the incineration section; c j (t) represents the center vector of the hidden layer; ||x(t)-c j (t)|| represents the Euclidean distance from the input sample at the current time to the center of the j-th hidden layer neuron; σ j w represents the width of the j-th hidden layer neuron; J is the total number of hidden layer neurons; j This represents the connection weight from the j-th hidden layer node to the output layer; (3) Introduce the moving block bootstrapping (MBB) method based on maximum likelihood estimation (MLE) to preserve the temporal dependency structure of the data in the subsample set, thereby reducing the destruction of temporal characteristics by traditional bootstrap sampling and improving the accuracy and reliability of modeling. The algorithm flow of MBB is as follows: ① The size of each data block is determined by the MLE algorithm. The size of each block is determined by MLE, and each block is numbered; the size of each block should be greater than 5 to include important temporal dependency structures in the time series data, while the size of each block should be less than N / 5 to avoid losing sample diversity; N represents the total number of samples; ②Bootstrap Sampling Extraction Block Blocks are randomly drawn from the original dataset with replacement until a Bootstrap sample of the same length as the original dataset is constructed; concatenated blocks are formed by linking the drawn blocks in chronological order in the original data sequence to create a Bootstrap subset. ④ Handling boundary effects Since the block size may not be an integer multiple of the original data sequence length, some extra data points may be generated when concatenating blocks; some extra data points are removed from the beginning or end of the Bootstrap sample to fill the gaps; parfor represents parallel loop; The MLE algorithm is as follows: Let D = {(x1,y1),(x2,y2),...,(x N ,y N )}, y=[y1,...y N ] T MLE is used to detect mutation points in y, and these mutation points are used to break the data, thus dividing D into θ sub-blocks; y represents the desired output, N represents the total number of samples, n represents the sequence number of the mutation point, and D is the set recording each set of input and output samples; where, ε is the preset minimum sub-block length, set to ε = 15, and θ is the maximum number of mutation points to be found. To round down, [ ] T Represents the transpose of a matrix or vector; ① Define the model The normal distribution function is used to approximate the local statistical characteristics before and after the mutation point, that is, the expected output y follows a normal distribution with different mean and variance before and after the mutation point t; let the data before the mutation point t follow a normal distribution. After the mutation point t, it obeys ε+1 and Ω-ε+1 represent the data points that need to be traversed in each data segment, and Ω represents the total number of observed data points; where μ1 represents the mean of the expected output y before the mutation point t; μ1 represents the variance of the expected output y before the mutation point t; μ2 represents the mean of the expected output y after the mutation point t. This represents the mean of the expected output y after the mutation point t; ② Calculate the likelihood function For a given mutation point t, the likelihood function L(t) of the data is: L(t) represents the observed data {x1, x2, ..., x} when the mutation point is located at time t. Ω The total likelihood function of x i f(x) represents the observed data value at time point i, where i is the index variable of the observed sample point, used to traverse the data sequence; i ;μ;σ 2 ) represents the variable x i With a mean of μ and a variance of σ 2 The probability density function of Ω under a normal distribution, where t is the hypothetical location of the mutation point, and its value ranges from 1 to t < Ω. ③ Find the log-likelihood function Taking the logarithm of the likelihood function to simplify the calculation, we obtain the log-likelihood function: logL(t) represents the log-likelihood function value given the position t of the mutation point. Further simplification yields: ④ Parameter estimation For each possible mutation point t, estimate the parameters μ1, μ2, Let represent the sample mean estimate of the observation data before the mutation point t, including the t-th point; This represents an estimate of the sample mean of the observed data after the mutation point t. This represents the sample variance estimate of the observation data before the mutation point t; This represents the sample variance estimate of the observed data after the mutation point t; They are respectively abbreviation; ⑤ Maximize the log-likelihood function For each possible mutation point t, the estimated parameters are substituted into the log-likelihood function: By iterating through all possible mutation points t, calculating the log-likelihood function value for each t, ​​and selecting the t with the largest log-likelihood function value as the optimal mutation point: This represents an estimate of the location of the optimal mutation point. This represents the parameter values ​​that maximize the objective function logL(t) on variable t. Each saved optimal mutation point is recorded as t = t1, t2, ... t. n , t n Let θ represent the nth mutation point, where 1 ≤ n ≤ θ. Taking the case where the number of optimal mutation points found equals the maximum number of mutation points required (i.e., n = θ) as an example... At this point, the dataset has been divided into θ sub-blocks. Further MBB sampling yields a set of subsamples Φ = {D1, D2, ..., D...}. T }; T represents T sub-training datasets used to train SL; D represents one of the partitioned training datasets in set Φ, containing θ subsets; Let t represent the i-th subset, where i = 1, ..., θ; i Let t1, t2, ..., t represent the split points of each subset of samples. θ-1 Only the index of the input feature vector, x i y represents the input feature vector of the i-th sample; N is the total number of original training samples; i Represents the relationship with input feature x i The corresponding target output value; (4) Integrating the AdaBoost algorithm with the Bagging algorithm: For each subset D m Train an SL in parallel and independently, and denote its output as SL. m Let m = 1, ..., T, and let AdaBoost be used to train each SL; for each BL, the first subset D m,j j=1, equal distribution D m,j The weight of each sample in the sample; Set the threshold matrix after each training session: Δ j =[Δ j,1 ,Δ j,2 ,...Δ j,N ], used to determine the learning effectiveness of each BL; Δ j Let Δ represent the threshold matrix of the j-th BL. j,k This represents the threshold corresponding to the k-th sample in the threshold matrix of the j-th BL. c is a constant used to adjust the prediction accuracy; T is the preset number of SLs, and U is the number of BLs in each SL. Let c = 0.03, T = 4, and U = 3. Initialize the weights of each sample in the subsample set W j =[w j,1 ,...,w j,N ],w j,k =1 / N,j=1,...,U,k=1,...,N (14) W j Let w be the sample weight matrix for the j-th BL. j,k This represents the weight corresponding to the k-th sample in the j-th BL; training the j-th BL yields F. j (x):x→y,F j (x) represents each trained BL; Calculate the regression error for each sample. y k This represents the expected output of the k-th sample. Let e ​​represent the predicted output of the k-th sample learned by the j-th BL. j,k The prediction error of the k-th sample learned by the j-th BL is represented by Design logic matrix L j =[L j,1 ,...,L j,N ](16) This is used to record the training performance of each sample in the j-th BL, where L j,k The position where =1 represents a poorly trained sample, L j,k The sample at position 0 represents a well-trained sample, and k = 1, ..., N; Calculate the regression error rate E of the j-th BL. j Calculate the weight coefficient α of the j-th BL. j Design feature matrix H j H j =[h j,1 ,...,h j,N ] (19) H j The elements in the only numbers are -1 and 1. If e j,k >Δ j,k Then h j,k Assigning a value of -1 increases the weight of the k-th sample in the j-th base learner; conversely, it increases the weight of h. j,k The value is assigned to 1, thereby reducing the sample weight; Calculate the normalization factor γ j This is a normalization factor used to ensure that the sum of the weights of all samples in the sample set is 1. W j Updated to W j+1 and through W j+1 D m,j Updated to D m,j+1 W j+1 Consider W as the probability distribution of each sample being drawn from the new sample set. j+1 The sum of all elements in D is 1; m,j According to the weight matrix W j+1 Weighted sampling yields a new sample set D m,j+1 Used to train the next BL; Calculate the output of the j-th BL. BL j =α j F j (x) (22) The output of each SL is: SL m For the output of the m-th SL; finally, calculate the attention score for each SL; a m The attention score is calculated for the m-th SL, and the attention weights w are computed using a softmax probabilistic network. m All SLs are weighted and combined to output a stronger ensemble learner, namely the strong learner in Figure 2, m = 1, ..., T; The core of the softmax probabilistic network is to use the softmax function to convert the model output into a probability distribution; its formula is shown in formula (26); (5) A self-attention (SA) mechanism is designed to replace the traditional simple averaging method. The weights are allocated by comprehensively considering the training performance of each SL, so as to integrate the modules more reasonably. The output of the SL is regarded as sequence data. The attention score and weight of each SL to the training samples are calculated by the SA mechanism, and these prediction results are linearly combined. The Q, K, and V in the proposed SA mechanism are defined as follows: The feature representation of the training samples is used to evaluate the attention of each SL to the training sample set; The output of each SL is compared with Q to determine the importance of each SL; The prediction result of each SL for the training samples, i.e., the output of the SL, V = K, K = [k1,...,k T V = [v1,...,v] T ], k m =v m =SL m =[y1,...,y N ] Step 1: Calculate the attention score AS: i = 1, ..., n, where n is the dimension of the sample, and AS is the comprehensive score of each SL across all dimensions of the training sample; Step 2: Calculate the attention weight AW: AW=[w1,...,w T ]=softmax(AS) (25) The softmax probabilistic network transforms the attention score into a probability distribution, which is the attention weight, where the value of each element is between 0 and 1, and the sum of all elements is 1. in K represents i The dimension of K is used as a scaling factor; in high-dimensional space, the dot product can become very large, causing gradient vanishing or exploding. Therefore, to ensure K... i With a stable range of attention scores, scaling is essential. Step 3: Calculate the weighted output for each SL: SL' m =w m v m (27) Step 4: Sum the weighted outputs of all SLs to obtain an integrated output X, and use it as the input of the FNN: (6) Self-attention scores are trained using a feedforward neural network (FNN); the FNN introduced in this paper is a three-layer network structure, including an input layer, a hidden layer and an output layer; the ReLU function is used as the activation function of the hidden layer; the ReLU function is explained as shown in formula (39); w1 and b1 represent the weights and biases of the input layer neurons, respectively; Z1 is the output of the input layer neurons; A1 is the ReLU activation function of the hidden layer; w2 represents the weights and biases of the hidden layer neurons, respectively; Z2 is the output of the hidden layer neurons; The FNN introduced in this paper is a three-layer network structure, including an input layer, a hidden layer, and an output layer; the ReLU function is used as the activation function of the hidden layer. The forward propagation process is as follows: Input layer to hidden layer: Z1 = w1X + b1 (29) Hidden layer: A1 = ReLU(Z1) (30) Hidden layer to output layer: Z2=w2A1+b2 (31) Output layer: yes pred =Z2 (32) Define the loss function: y predi To predict the output, y i The expected output; The backpropagation process is as follows: T Represents the transpose of a vector or matrix; Output layer to hidden layer: Hidden layer to input layer: ⊙ represents the Hadamard product of a matrix. From input layer to attention layer: During each training iteration, the five parameters w1, b1, w2, b2, and K are updated along the negative gradient direction of the loss function. Where η∈(0,1) is the learning rate of the FNN, and the value of K at the end of training is denoted as K'. Steps 1 to 4 are repeated with V'=K' to obtain the training attention weights w1',...,w'. T And the final integrated output: ELO