Hearth temperature prediction method for urban solid waste incineration process based on ESORNN-IFG
By adopting the ESORNN-IFG model during the incineration of urban solid waste, combined with Adaboost integrated learning and recursive algorithm, the problem of low furnace temperature prediction accuracy is solved, and high accuracy FT control and environmental protection effects are achieved.
Patent Information
- Application Number
- CN202510030908.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-30
AI Technical Summary
During the incineration of urban solid waste, the prediction accuracy of furnace temperature is low, which makes it difficult to control combustion efficiency and pollutant emissions.
The furnace temperature prediction method based on integrated self-organized recursive neural network and information fusion gain algorithm (ESORNN-IFG) is adopted, and combined with Adaboost integrated learning algorithm and recursive algorithm, the contribution of hidden neurons is evaluated through the information fusion gain index, and the basic learner structure is dynamically adjusted through the self-organizing strategy.
It significantly improves the accuracy and robustness of furnace temperature prediction, overcomes the highly nonlinear and dynamic challenges in the MSWI process, and achieves the precise control and environmental protection effects of FT.
Smart Images

Figure CN120068595A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a furnace temperature (FT) prediction method based on an Ensemble Self-Organizing Recursive Neural Network and Information Fusion Gain algorithm (ESORNN-IFG), which is applicable to the FT control and optimization in the Municipal Solid Waste Incineration (MSWI) process. This method utilizes an ensemble learning mechanism, combined with the weighted combination of multiple base learners, significantly improving the accuracy and robustness of the model, and overcoming the challenges of high nonlinearity, dynamics, and complex reaction mechanisms in the MSWI process. By introducing the Information Fusion Gain (IFG) index, it can effectively evaluate the contributions of hidden neurons and their mutual relationships. In addition, a self-organizing strategy is designed, combined with the IFG index, to dynamically adjust the structure of the base learners during the model training process. Through experimental tests on multiple benchmark problems and comparisons with actual application cases, the superiority and effectiveness of the method proposed in the present invention in FT prediction are verified, and it has good application prospects. Background Art
[0002] The accelerating global economic growth and urbanization process have led to a continuous increase in the generation of Municipal Solid Waste (MSW), posing challenges to the environment, health, and urban aesthetics. Improper treatment of MSW may cause soil and water pollution, endanger the health of residents, and damage the urban landscape. The current main treatment methods of MSW include sanitary landfill, composting, and incineration. The sanitary landfill process is simple and has a large treatment capacity, but it may pose hazards to groundwater and surface water, and there are potential risks of biogas and heavy metal pollution. Composting uses microorganisms to convert biodegradable organic matter into humus, with obvious environmental protection advantages, but the treatment time is relatively long, and the residues may still cause pollution. In contrast, the MSWI technology has gradually become the main solution for treating MSW due to its advantages such as small land occupation, flexible site selection, significant reduction effect, and stable operation, and has received extensive attention.
[0003] In the MSWI process, as a key control variable, FT directly affects the combustion efficiency and pollutant emission concentration. Therefore, precise control of FT is crucial for environmental protection. Before achieving precise control, an accurate FT model must be established. FT is affected by various factors, including the speeds of multiple grate sections and the airflows in multiple air supply pipes. Adjusting the grate speed not only directly affects the fuel supply efficiency but also has an important impact on the smoothness of the combustion process, thus significantly affecting the change of FT. At the same time, adjusting the supply of combustion air is also a key step because it is directly related to the activity of the combustion reaction and further affects the FT fluctuation. Therefore, establishing an accurate FT model is a challenging task. The FT model based on ESORNN-IFG takes the primary and secondary air volumes in the furnace and the speeds of each grate section as the neural network inputs, and FT as the neural network output. FT prediction experiments are carried out based on the real data provided by a waste incineration plant in Beijing.
[0004] The present invention relates to the design and research of an integrated improved FT model for the MSWI process. This model mainly uses ensemble learning techniques, especially introducing the Adaboost algorithm into FT prediction. First, a recursive algorithm is designed in each base learner to enhance the time memory ability of the network. Second, the IFG algorithm is used to evaluate the contributions of hidden neurons and their mutual relationships. At the same time, a self-organization strategy is introduced, using the IFG algorithm, enabling each base learner to selectively optimize neurons, thereby adaptively improving its structure. Finally, the gradient descent method is used to optimize the network parameters, thus improving the accuracy and learning efficiency of the network. Through the above innovations, the present invention solves the problems of poor accuracy and practicability of traditional FT models, and improves the combustion efficiency and environmental protection effect. Summary of the Invention
[0005] The present invention designs an FT model based on ESORNN-IFG to solve the problem of FT prediction in the MSWI process; integrates the Adaboost ensemble learning algorithm to solve the problems of low prediction accuracy and poor generalization performance of a single model; configures a recursive algorithm for each base learner to enhance the learning ability for time-series data; designs a self-organization strategy based on the IFG algorithm to optimize the network structure and improve the dynamic performance of the model. Through this model, the accuracy of the FT model can reach the best, solving the problem of high uncertainty in FT changes in MSWI and difficult to accurately predict; this invention significantly improves the accuracy of the FT model and provides a reliable solution for environmental protection in the incineration process.
[0006] The present invention adopts the following technical solutions and implementation steps:
[0007] FT Model Based on ESORNN-IFG
[0008] (1) Design of the base learner structure. A three-layer architecture is adopted. Neurons in the input layer receive external signals, neurons in the hidden layer process the input combinations from the input layer and the output layer, and neurons in the output layer integrate the information in the hidden layer to generate the final network output. The Gaussian function is used as the activation function in the hidden layer to quantify the distance between the input and their respective centers, thereby improving the sensitivity of each hidden neuron and realizing the non-linear mapping of the input data. In addition, to enhance the dynamic performance and memory ability of the model, an internal recursive structure is added to each base learner, and the specific implementation is as follows:
[0009] Input layer: x(t) = [x 1 (t), x 2 (t),... x n (t)] is a set of samples containing all input features at the current moment. (n = 4, indicating that the number of selected input features is 4. The inputs of the neural network are the primary total air volume, secondary total air volume, grate speed in the drying section, and grate speed in the incineration section)
[0010] Hidden layer:
[0011]
[0012] Column vector c .j (t) represents the j-th column of matrix c
[0013]
[0014] x j (t) is the input vector of the j-th neuron at time t, c ij (t) represents the center of the j-th hidden layer neuron of the i-th dimensional feature of the input sample, c .j (t) represents the center of the j-th hidden layer neuron at the current time, p represents the number of hidden layer nodes, ||χ j (t) - c .j (t)|| represents the Euclidean distance from the input sample to the center of the j-th hidden layer neuron at this moment, b(t) = [b 1 (t), b 2 (t),..., b p (t)] T , b j (t) is the radius of the j-th hidden layer neuron.
[0015] Note: [ ] T represents the transpose of a matrix or vector.
[0016] χ j (t) = [x(t), z j (t)] is the input vector of the self-organizing recurrent neural network, where
[0017]
[0018] represents the output of the self-organizing recurrent neural network at time t - 1,
[0019] ν(t) = [v 1 (t), v 2 (t),..., v p (t)] T , ν j (t) is the bias vector connecting the output layer and the hidden layer neurons.
[0020] Output layer:
[0021]
[0022] w(t) = [w 1 (t), w2 (t),...,w p (t)] T ,w j (t) is the weight of the connection between the output layer neuron and each hidden layer neuron. h(t) = [h 1 (t), h 2 (t),..., h p (t)] T ,h j (t) is the output vector of the j-th hidden layer neuron at time t.
[0023] Using the training of the base learner K times as a sliding window, record the output Φ l (t) of the l-th hidden layer neuron in each sliding window, and the normalized true output Y(t). The matrix Φ l (t) represents the output of the l-th hidden layer neuron in a sliding window at time t, and Φ(t) represents the set of outputs of all hidden layer neurons in a sliding window. y(t) represents the true value of FT at time t, and Y(t) is the output vector recording the true value of FT
[0024] Y(t) = [y(t - N + 1),..., y(t - 1), y(t)] T (5)
[0025]
[0026] Φ(t) = [Φ 1 (t), Φ 2 (t),..., Φ p (t)], l = 1,..., p (7)
[0027] Where N is the number of samples, T is the current training iteration, h T-K+1 (t - N + 1) is the output vector of the hidden layer at time t - N + 1 during the (T - K + 1)-th training of the base learner
[0028] Calculate the correlation coefficient R l→y between the l-th hidden layer neuron and the output y at time t, and the independent contribution I l→y of the l-th hidden layer neuron with respect to the output y, l = 1,..., p.
[0029]
[0030] Φ l (i, j) and P l (i, j) represent the i-th row and j-th column of the matrix Φ l (t) and the matrix P l (t) respectively
[0031] Q(j) and Y(j) represent the j-th elements of vectors Q(t) and Y(t) respectively.
[0032] P l (t) and Q(t) are the output matrix of the l-th hidden layer neurons and the probability distribution vector of the output y respectively. ∑ i,j Φ l (i, j) represents the sum of all rows and columns of matrix Φ l ∑ represents the sum of all rows and columns. j Y(j) represents the sum of the column vector Y(t).
[0033] Define the correlation coefficient between the l-th hidden layer neurons and the output layer neurons as:
[0034]
[0035] represents the i-th row and j-th column of the matrix represents the sum of all rows and columns of the matrix ∑ represents the sum of all rows and columns.
[0036] The matrix represents the output probability distribution of the l-th hidden layer neurons at the k-th training
[0037]
[0038] R y is the correlation coefficient matrix between each hidden layer neuron and the output y within a window. is the correlation coefficient between the l-th hidden layer neurons and the output y at the k-th training within a window.
[0039] Centering process is performed on Φ l (t) and Y(t) respectively to obtain and
[0040] and represent the mean of each column vector of matrix Φ l (t) and the mean of Y(t) respectively.
[0041] Denote matrix Q(t) = [Q 1 (t), Q 2 (t),..., Q K (t)], represents the k-th column vector of the matrix, symbolizing the k-th training matrix Φ within a sliding window l k(t) represents the output of the l-th hidden layer neuron in the k-th training within a sliding window at time t.
[0042]
[0043] B k (t) is composed of the autocovariance and cross-covariance of Φ l k (t) and Y(t). Q k T (t) represents the transpose of Q k (t), and Cov is the covariance symbol. Cov(Φ l k (t), Y(t)) means to obtain the cross-covariance of Φ l k (t) and Y(t), and Cov(Φ l k (t), Φ l k (t)) means to obtain the autocovariance of Φ l k (t).
[0044] Perform de-correlation processing on Φ l k (t) and Y(t) respectively to obtain and
[0045]
[0046] represents the symbol of the Hadamard product operation of matrices. Y(t) is the normalized true output at time t, and are the results of de-correlation processing on Φ l k (t) and Y(t) respectively, and the result is denoted as Ω l (t) = [Η l 1 (t), Η l 2 (t),..., Η l K (t)],
[0047]
[0048] Ω(t) = [Ω 1 (t),..., Ω l (t),..., Ω p (t)] (16)
[0049] Define the independent contribution \(I\) of the \(l\)-th hidden layer neuron to the output l→y as
[0050]
[0051] Denote the matrix \(I = [I 1→y , I 2→y , \cdots, I p→y \), where \(I l→y is the independent contribution of the \(l\)-th hidden layer neuron to the output layer neuron.
[0052] \(I\) is the independent contribution matrix of the hidden layer neurons to the output layer neurons.
[0053] represents the independent contribution of the \(l\)-th hidden layer neuron at the \(k\)-th training in a sliding window at the input time of the \(n\)-th sample.
[0054] The self-organization mechanism designed by the present invention includes a total of three rules: neuron splitting rule, neuron merging rule, and neuron retention rule.
[0055] Case 1. Neuron splitting rule
[0056] When the \(l\)-th hidden layer neuron simultaneously satisfies \(I l→y =\max\{I\}\) and \(R l→y =\min\{R y \}\), (\(R y represents the correlation of the \(l\)-th hidden layer neuron with respect to the output layer neuron), that is, when the contribution degree and correlation of the \(l\)-th hidden layer neuron to the output layer neuron are both the largest, a new hidden layer neuron is added to the base learner. The parameter settings of the new hidden layer neuron are as follows:
[0057]
[0058] where \(e(t)\) is the prediction error at time \(t\), and \(b new (t)\), \(c new (t)\), \(w new (t)\) and \(h new (t)\) represent the radius, center, weight, and output of the newly added hidden layer neuron at time \(t\) respectively; \(v new (t)\) represents the connection weight between the newly added hidden layer neuron and the \(l\)-th hidden layer neuron at time \(t\); \(b l (t)\), \(v l (t)\) and \(e(t)\) represent the radius, connection weight with the output, and the sum of the prediction error of the base learner at time \(t\) of the \(l\)-th hidden layer neuron respectively; \(c l (t)\) and \(h l(t) represents the center and output of the l-th hidden layer neuron at time t respectively.
[0059] Case 2. Neuron merging rule
[0060] When the u-th hidden layer neuron simultaneously satisfies I u→y = min{I} and R u→y = max{R y}, that is, when the contribution degree and correlation of the u-th hidden layer neuron to the output layer neuron are both the smallest, calculate the correlation coefficient R uo .
[0061]
[0062] Φ u (i,j) and Φ o (i,j) represent the i-th row and j-th column of matrix Φ u and matrix Φ o respectively, symbolizing the outputs of the u-th and o-th hidden layer neurons in a sliding window. S(i,j) and T(i,j) represent the i-th row and j-th column of matrix S and matrix T respectively. S and T are the probability distribution matrices of the outputs of the u-th and o-th hidden layer neurons. ∑ i,j Φ u (i,j) and ∑ i,j Φ o (i,j) represent the sum of all rows and columns of matrix Φ u and matrix Φ o respectively. u, o = 1,..., p, where p is the number of hidden layer neurons. On this basis, define the correlation coefficient R u→o between the u-th and o-th hidden layer neurons as:
[0063]
[0064] ∑ i,j S(i,j) represents the sum of all rows and columns of matrix S
[0065] Denote matrix
[0066] R uo = [R u→1 ,…R u→u-1 ,R u→u+1 ,...R u→p (22)
[0067] When R u→o = min{R uo}, that is, when the correlation between the $u$-th and $o$-th hidden layer neurons is the largest, the $u$-th and $o$-th hidden layer neurons are merged into a new neuron. The parameters of the new hidden layer neuron are set as follows:
[0068]
[0069] where $e(t)$ is the prediction error at time $t$, and $b$ new (t), $c$ new (t), $w$ new (t) and $h$ new (t) represent the radius, center, weight, and output of the newly added hidden layer neuron at time $t$ respectively; $v$ new (t) represents the connection weight between the newly added hidden layer neuron and the $l$-th hidden layer neuron at time $t$; $b$ u (t) and $b$ o (t) represent the radii of the $u$-th and $o$-th hidden layer neurons at time $t$ respectively; $c$ u (t) and $c$ o (t) represent the centers of the $u$-th and $o$-th hidden layer neurons at time $t$ respectively; $w$ u (t) and $w$ o (t) represent the weights of the $u$-th and $o$-th hidden layer neurons at time $t$ respectively, and $h$ u (t) and $h$ o (t) represent the outputs of the $u$-th and $o$-th hidden layer neurons at time $t$ respectively. $v$ u (t) represents the connection weight between the $u$-th hidden layer neuron and the output at time $t$.
[0070] Case 3. Neuron retention rule
[0071] When neither Case 1 nor Case 2 occurs or both Case 1 and Case 2 are triggered simultaneously, all hidden layer neurons are retained and the structure of the base learner does not change.
[0072] Using the self-organization mechanism, hidden layer neurons are added or pruned, so that the structure of the base learner changes dynamically during training, simulating the characteristics of the dynamic change of neuron activity, which is more conducive to adaptive learning in a dynamic environment.
[0073] (3) Training the base learner. There are four types of parameters to be learned in the training stage of the self-organizing recurrent neural network, namely $b(t)$, $c(t)$, $w(t)$ and $v(t)$
[0074] Define the loss function:
[0075]
[0076] The update processes of these four parameters are as follows:
[0077]
[0078] η and θ are the learning rate and momentum factor respectively, and the value ranges of η and θ are both between 0 and 1
[0079]
[0080]
[0081] (4) Integrate all base learners. The training process of a single base learner in the Adaboost algorithm based on the regression error rate is as follows:
[0082] Input: For a training sample set S with multiple inputs and a single output m
[0083] S m ={{X 1 ,Y 1},{X 2 ,Y 2},...,{X N ,Y N}} (33)
[0084] X i =[x 1 ,x 2 ,…x n ,i = 1,...,,N (34)
[0085] N is the number of samples; n is the dimension of the neural network input features (n = 4); M is the number of base learners (set to 10); X i represents the neural network input, which are (total primary air volume, total secondary air volume, grate speed in the drying section, grate speed in the combustion section) respectively, and Y i represents the true value of FT corresponding to the i-th group of inputs
[0086] Integrated output:
[0087]
[0088] α m represents the weight of the m-th base learner, f m (X) represents the output of the m-th base learner, and H(X) represents the final integrated output
[0089] Define the matrix
[0090] delta = [delta 1 ,delta 2 ,...delta N (36)
[0091]
[0092] c is a constant used to adjust the prediction accuracy. The value range of c is generally between 0.1 and 5 (c = 3). N is the number of samples, and Y i represents the i-th output value, and represents the mean value of the output.
[0093] delta i is used to compare with the absolute value of the prediction error If then assign 1 to the element at the corresponding position in the logical matrix I m and assign 0 otherwise. y mi represents the true output corresponding to the i-th input, i and represents the neural network prediction output corresponding to the i-th input of the m-th base learner.
[0094] ① Initialize the sample weight distribution matrix D 1
[0095] D 1 = [d 11 , d 12 ,... d 1N , where d 1i = 1 / N, i = 1,..., N (38)
[0096] d 1i represents the weight corresponding to the i-th sample in the first base learner, and N is the number of samples
[0097] ② for m = 1 to M
[0098] ③ Train the base learner according to the sample distribution weight matrix
[0099] ④ Calculate the regression error rate
[0100]
[0101] delta i is determined by formula (37)
[0102]
[0103] where d mi is the weight corresponding to the i-th sample in the m-th base learner, and I m = [I m1 , I m2 ,..., I mN is a logical matrix used to record the training quality of each sample in the m-th base learner. The matrix I mAmong them, the sample at position 1 indicates a poorly trained sample, and the sample at position 0 is a well-trained sample. e m represents the regression error rate of the m-th base learner.
[0104] ⑤ Calculate the weights of the base learners
[0105]
[0106] ⑥ Update the sample weights
[0107]
[0108] Z m is a normalization factor used to ensure that the sum of the weights of all samples in the sample set is 1. Z m The calculation formula of is shown in formula (45)
[0109] D m (i)=[d m1 ,d m2 ,...d mN and D m+1 (i) represent the sample weight distribution matrices of the m-th and the (m + 1)-th base learners respectively. d mi is the weight corresponding to the i-th sample in the m-th base learner and is also the probability that this sample is sampled when generating the sample set S m+1 . By adopting a weighted sampling strategy on S m , an updated sample set S m+1 is generated for the training and update of the (m + 1)-th base learner.
[0110] S m ={{X 1 ,Y 1},{X 2 ,Y 2},...,{X N ,Y N}} represents the multi-input single-output training sample set corresponding to the m-th base learner
[0111]
[0112] The elements in are only -1 and 1. If then is assigned -1, calculated by formula (42), thereby increasing d mi , that is, increasing the weight corresponding to the i-th sample in the m-th base learner. Conversely, is assigned 1, thereby reducing d mi .
[0113] y idenotes the true output corresponding to the i-th input, denotes the neural network prediction output corresponding to the i-th input of the m-th base learner.
[0114] where is the normalization factor, which is used to ensure that the sum of the weights of all samples in the sample set is 1
[0115]
[0116] H m (X) = α m f m (X) (46)
[0117] H m (X) is the output of the m-th base learner.
[0118] The creativity of the present invention is mainly reflected in:
[0119] (1) In view of the complexity of the MSWI process, especially the characteristics that the FT is affected by various factors, the present invention proposes a comprehensively improved FT model. By introducing the Adaboost algorithm and the RRBFNN structure, the precise control of the FT is realized by using the ensemble learning technology, which has strong adaptability and self-learning ability, can effectively cope with the FT fluctuation, and improve the combustion efficiency and environmental protection effect.
[0120] (2) The FT model designed by the present invention integrates the recursive algorithm and the IFG algorithm, solves the deficiencies of the traditional model in dealing with nonlinear and time-varying systems, and realizes the real-time monitoring and adjustment of the FT. This innovative method significantly reduces the control complexity, has the advantages of low energy consumption and simple structure, and provides a reliable solution for the stability of the incineration process and environmental protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0121] Figure 1 is the flowchart of the Adaboost ensemble algorithm of the present invention
[0122] Figure 2 is the structure diagram of the FT model of ESORNN-IFG of the present invention
[0123] Figure 3 is the structure diagram of the base learner of the present invention
[0124] Figure 4 is the RMSE curve diagram of each base learner in the training stage of the present invention
[0125] Figure 5 is the fitting result diagram of the output and the expected output of the FT model in the test stage of the present invention
[0126] Figure 6It is the fitting error graph between the output of the FT model and the expected output during the test phase of the present invention. Detailed implementation manners
[0127] The present invention designs an FT model based on ESORNN-IFG to solve the FT prediction problem in the MSWI process. By integrating the Adaboost ensemble learning algorithm, the present invention effectively improves the prediction accuracy and overcomes the insufficient generalization performance of a single model. At the same time, a recursive algorithm is configured for each base learner to enhance the learning ability for time-series data. The self-organization strategy based on the IFG algorithm optimizes the network structure and improves the dynamic performance of the model. Through this model, the prediction accuracy of FT is significantly improved, effectively solving the uncertainty of FT changes in MSWI and ensuring the realization of accurate prediction. The invention provides a reliable solution for environmental protection in the incineration process.
[0128] The present invention adopts the following technical solutions and implementation steps:
[0129] Input layer: x(t) = [x 1 (t), x 2 (t),... x n (t)] is a set of samples containing all input features at the current moment.
[0130] Hidden layer:
[0131]
[0132] x j (t) is the input vector of the j-th neuron at time t, c ij (t) is the center of the j-th hidden layer neuron, ||χ j (t) - c .j (t)|| represents the Euclidean distance from the input sample at this moment to the center of the j-th hidden layer neuron, b(t) = [b 1 (t), b 2 (t),..., b p (t)] T , b j (t) is the radius of the j-th hidden layer neuron.
[0133] Initialize the internal feedback to 0. χ j (t) = [x(t), z j (t)] is the input vector of the self-organizing recurrent neural network, where
[0134]
[0135] ν(t) = [v 1 (t), v 2(t),...,v p (t)] T ,ν j (t) is the bias vector connecting the output layer and the hidden layer neurons.
[0136] Output layer:
[0137]
[0138] w(t) = [w 1 (t), w 2 (t),..., w p (t)] T ,w j (t) are the weights connecting the output layer neurons and each hidden layer neuron. h(t) = [h 1 (t), h 2 (t),..., h p (t)] T ,h j (t) is the output vector of the j-th hidden layer neuron at time t.
[0139] Taking the training of the base learner K times as a sliding window, record the output Φ of the l-th hidden layer neuron in each window l (t), and the normalized true output Y(t)
[0140] Y(t) = [y(t - N + 1),..., y(t - 1), y(t)] T (5)
[0141]
[0142] Φ(t) = [Φ 1 (t), Φ 2 (t),..., Φ p (t)], l = 1,..., p (7)
[0143] where N is the number of samples, T is the current training number, h T-K+1 (t - N + 1) is the output vector of the hidden layer at time t - N + 1 of the base learner in the (T - K + 1)-th training. Calculate the correlation coefficient R between the l-th hidden layer neuron and the output y at time t l→y and the independent contribution I of the l-th hidden layer neuron with respect to the output y l→y , l = 1,..., m.
[0144]
[0145] P l(i, j) and Q(j) are the probability distribution matrices of the outputs of the neurons in the l-th hidden layer and the output y, respectively. Σ i,j Φ(i, j) represents the sum of all rows and columns of the matrix Φ(i, j), Σ j Y(j) represents the sum of all columns of the column vector Y(j).
[0146] Define the correlation coefficient between the neurons in the l-th hidden layer and the neurons in the output layer as:
[0147]
[0148] P k (i, j) represents the output probability distribution of the neurons in the l-th hidden layer during the k-th training
[0149]
[0150] R y is the correlation coefficient matrix between each neuron in the hidden layer and the neurons in the output layer within a window. is the correlation coefficient between the neurons in the l-th hidden layer and the output during the k-th training within a window.
[0151] Centrally process Φ l (t) and Y(t) to obtain and
[0152] and represent the mean of the k-th column vector of the matrix Φ l (t) and the mean of Y(t), respectively.
[0153] Denote the matrix Q(t) = [Q 1 (t), Q 2 (t), …, Q K (t)],
[0154]
[0155] Decorrelate Φ l k (t) and Y(t) to obtain and
[0156]
[0157] Denote Ω l (t) = [Η l 1 (t), Η l 2(t),...,Η l K (t)],
[0158]
[0159] Ω(t) = [Ω 1 (t), …, Ω l (t), …, Ω p (t)] (16)
[0160] Define the independent contribution I of the l-th hidden layer neuron to the output as: l→y as:
[0161]
[0162] Denote the matrix I = [I 1→y , I 2→y ,..., I p→y , I l→y is the independent contribution of the l-th hidden layer neuron to the output layer neuron. The self-organization mechanism designed in the present invention includes three rules in total: neuron splitting rule, neuron merging rule, and neuron retention rule.
[0163] Case 1. Neuron splitting rule
[0164] When the l-th hidden layer neuron simultaneously satisfies I l→y = max{I} and R l→y = min{R y}, that is, when the contribution degree and correlation of the l-th hidden layer neuron to the output layer neuron are both the largest, a new hidden layer neuron is added to the base learner. The parameter settings of the new hidden layer neuron are as follows:
[0165]
[0166] where b new (t), c new (t), w new (t) and h new (t) represent the radius, center, weight, and output of the newly added hidden layer neuron at time t respectively; v new (t) represents the connection weight between the newly added hidden layer neuron and the l-th hidden layer neuron at time t; b l (t), v l (t) and e(t) represent the radius, connection weight between the l-th hidden layer neuron and the output, and the sum of the prediction errors of the base learner at time t respectively; c l (t) and h l(t) represent the center and output of the l-th hidden layer neuron at time t respectively.
[0167] Case 2. Neuron merging rule
[0168] When the k-th hidden layer neuron simultaneously satisfies I k→y = min{I} and R k→y = max{R y}, that is, when the contribution degree and correlation of the k-th hidden layer neuron to the output layer neuron are both the smallest, calculate the correlation coefficient R ko .
[0169]
[0170] S(i, j) and T(i, j) are the probability distribution matrices of the outputs of the k-th and o-th hidden layer neurons respectively, and ∑ i,j Φ(i, j) represents the sum of all rows and columns of the matrix Φ(i, j), where k, o = 1,..., m.
[0171] On this basis, define the correlation coefficient between the k-th and o-th hidden layer neurons as:
[0172]
[0173] Denote the matrix
[0174] R ko = [R k→1 ,...R k→k-1 , R k→k+1 ,...R k→p (22)
[0175] When R k→o = min{R ko}, that is, when the correlation between the k-th and o-th hidden layer neurons is the largest, merge the k-th and o-th hidden layer neurons into a new neuron. The parameter settings of the new hidden layer neuron are as follows:
[0176]
[0177] b k (t) and b o (t) represent the radii of the k-th and o-th hidden layer neurons at time t respectively; c k (t) and c o (t) represent the centers of the k-th and o-th hidden layer neurons at time t respectively; w k (t) and w o (t) represent the weights of the k-th and o-th hidden layer neurons at time t respectively, h k(t) and h o (t) represents the outputs of the k-th and o-th hidden layer neurons at time t. v k (t) represents the connection weight between the k-th hidden layer neuron and the output at time t.
[0178] Case 3. Neuron retention rule
[0179] When neither Case 1 nor Case 2 occurs or when both Case 1 and Case 2 are triggered simultaneously, all hidden layer neurons are retained and the structure of the base learner does not change.
[0180] Using the self-organization mechanism, hidden layer neurons are added or pruned, so that the structure of the base learner changes dynamically during training, simulating the characteristics of the dynamic change of neuron activity, which is more conducive to adaptive learning in a dynamic environment.
[0181] (3) Train the base learner. There are four parameters to be learned in the training phase of the self-organizing recurrent neural network, namely b(t), c(t), w(t) and v(t)
[0182] Define the loss function:
[0183]
[0184] The update processes of these four parameters are as follows:
[0185]
[0186]
[0187] η and θ are the learning rate and momentum factor respectively, η, θ ∈ (0, 1)
[0188]
[0189] (4) Integrate all base learners. The training process of a single base learner in the Adaboost algorithm based on the regression error rate is as follows:
[0190] Input: For a training sample set with multiple inputs and a single output
[0191] S m ={{X 1 ,Y 1},{X 2 ,Y 2},...,{X N ,Y N}} (33)
[0192] X i =[x 1 ,x 2 ,...xn , i = 1, ..., N (34)
[0193] N is the number of samples; n is the input dimension; M is the number of base learners
[0194] Ensemble output:
[0195]
[0196] Threshold matrix:
[0197] delta = [delta 1 , delta 2 ,... delta N (36)
[0198]
[0199] c is a constant used to adjust the prediction accuracy
[0200] ① Initialize the sample distribution weight matrix
[0201] D 1 = [d 11 , d 12 ,... d 1N , d 1i = 1 / N (38)
[0202] ② for m = 1 to M
[0203] ③ According to the sample distribution weight matrix D m Train the base learner f m : X → Y
[0204] ④ Calculate the regression error rate
[0205]
[0206]
[0207] where I m = [I m1 , I m2 ,..., I mN is the logical matrix
[0208] ⑤ Calculate the weight of the base learner
[0209]
[0210] ⑥ Update the sample weights
[0211]
[0212] D m+1 represents a probability distribution, where the sum of the sampling probabilities of each component is 1, facilitating the training of subsequent base learners. By adopting a weighted sampling strategy on S m , an updated sample set S m+1 is generated for further analysis and model update.
[0213]
[0214] where is the normalization factor
[0215]
[0216] H m (X) = α m f m (X) (46).
Claims
1. The furnace temperature prediction method of municipal solid waste incineration process based on ESORNN-IFG is characterized by: The following steps are involved: (1) The base learner structure is designed with a three-layer architecture. The neurons in the input layer receive external signals, the neurons in the hidden layer process the input combination from the input layer and the output layer, and the neurons in the output layer integrate the information of the hidden layer to generate the final network output. The hidden layer uses a Gaussian function as the activation function to quantify the distance between the input and the respective center. The specific implementation is as follows: Input layer: x(t) = [x1(t), x2(t), ...x n (t)] is a set of samples containing all input features at the current moment; n = 4, indicating that the number of selected input features is 4; the inputs of the neural network are primary total air, secondary total air, drying section grate speed, and incineration section grate speed; Hidden Layer: Column vector c .j (t) represents the jth column of matrix c x j (t) is the input vector of the jth neuron at time t, c ij (t) represents the center of the jth hidden layer neuron of the i-th dimension feature of the input sample, c .j (t) represents the center of the jth hidden layer neuron at the current moment, p represents the number of hidden layer nodes, ||χ j (t)-c .j (t)|| represents the Euclidean distance from the input sample to the center of the jth hidden layer neuron at this moment, b(t) = [b1(t), b2(t), ..., b p (t)] T , b j (t) is the radius of the jth hidden layer neuron; [] T Represents the transpose of a matrix or vector; χ j (t)=[x(t),z j (t)] is the input vector of the self-organizing recurrent neural network, where represents the output of the self-organizing recurrent neural network at time t-1, ν(t)=[v1(t),v2(t),...,v p (t)] T , ν j (t) is the bias vector connecting the output layer and the hidden layer neurons; Output layer: w(t)=[w1(t),w2(t),...,w p (t)] T , w j (t) is the weight of the connection between the output layer neuron and each hidden layer neuron; h(t) = [h1(t),h2(t),…,h p (t)] T ,h j (t) is the output vector of the jth hidden layer neuron at time t; Take the base learner training K times as a sliding window and record the output Φ of the lth hidden layer neuron in each sliding window l (t), and the normalized true output Y(t); the matrix Φ l (t) represents the output of the lth hidden layer neuron in a sliding window at time t, Φ(t) represents the set of outputs of all hidden layer neurons in a sliding window; y(t) represents the true value of FT at time t, and Y(t) is the output vector recording the true value of FT Y(t)=[y(t-N+1),…,y(t-1),y(t)] T (5) Φ(t)=[Φ1(t),Φ2(t),...,Φ p (t)],l=1,...,p (7) Among them, N is the number of samples, T is the number of training times currently performed, and h T-K+1 (t-N+1) is the hidden layer output vector of the base learner at time t-N+1 after T-K+1 training. Calculate the correlation coefficient R between the lth hidden layer neuron and the output y at time t l→y and the independent contribution of the lth hidden layer neuron to the output y l→y , l=1,…,p; Φ l (i,j) and P l (i,j) represent the matrix Φ l (t) and the matrix P l (t) is the i-th row and j-th column Q(j) and Y(j) represent the jth element of vectors Q(t) and Y(t) respectively. P l (t) and Q(t) are the output matrix of the lth hidden layer neuron and the probability distribution vector of the output y, respectively; ∑ i,j Φ l (i,j) represents the pair matrix Φ l Sum all rows and columns, Σ j Y(j) represents the sum of the column vector Y(t); The correlation coefficient between the lth hidden layer neuron and the output layer neuron is defined as: Representation Matrix The i-th row and j-th column of Represents the matrix Sum all rows and columns; matrix Represents the output probability distribution of the lth hidden layer neuron at the kth training R y is the correlation coefficient matrix between each hidden layer neuron and output y in a window; is the correlation coefficient between the k-th training and output y of the l-th hidden layer neuron in a window; For Φ l (t) and Y(t) are centrally processed to obtain and and Represents the matrix Φ l (t) The mean of each column vector, and the mean of Y(t); Let the matrix Q(t) = [Q1(t),Q2(t),...,Q K (t)], Representative Matrix The kth column vector represents the kth training in a sliding window. matrix Represents the output of the lth hidden layer neuron in a sliding window at time t during the kth training; B k (t) is given by Φ l k The matrix consisting of the variance and covariance of (t) and Y(t); Q k T (t) represents Q k (t), Cov is the covariance symbol, Cov(Φ l k (t),Y(t)) represents the solution of Φ l k The covariance of (t) and Y(t), Cov(Φ l k (t),Φ l k (t)) represents the calculation of Φ l k The variance of (t); For Φ l k (t) and Y(t) are treated as irrelevant. and represents the Hadamard product operator of the matrix, Y(t) is the normalized true output at time t, and They are Φ l k The result after making (t) and Y(t) irrelevant Denote Ω l (t) = [Η l 1 (t), Η l 2 (t),..., Η l K (t)], Ω(t)=[Ω1(t),...,Ω l (t),...,Ω p (t)] (16) Define the independent contribution of the lth hidden layer neuron to the output I l→y for: Remember matrix I=[I 1→y ,I 2→y ,...,I p→y ], I l→y is the independent contribution of the lth hidden layer neuron to the output layer neuron; I is the independent contribution matrix of hidden layer neurons to output layer neurons; represents the independent contribution of the lth hidden layer neuron at the kth training time within a sliding window at the nth sample input moment; The designed self-organizing mechanism includes three rules: neuron splitting rule, neuron merging rule and neuron retention rule; Case 1. Neuron division rules When the lth hidden layer neuron satisfies I l→y = max{I} and R l→y =min{R y }, R y Indicates the relevance of the lth hidden layer neuron to the output layer neuron, that is, when the contribution and relevance of the lth hidden layer neuron to the output layer neuron are both the largest, a new hidden layer neuron is added to the base learner; the parameters of the new hidden layer neuron are set as follows: Where e(t) is the prediction error at time t, b new (t), c new (t), w new (t) and h new (t) represent the radius, center, weight and output of the newly added hidden layer neuron at time t; v new (t) represents the connection weight between the newly added hidden layer neuron and the lth hidden layer neuron at time t; b l (t), v l (t) and e(t) represent the radius of the lth hidden layer neuron at time t, the connection weight between it and the output, and the prediction error of the base learner at time t; c l (t) and h l (t) represents the center and output of the lth hidden layer neuron at time t respectively; Case 2. Neuron merging rules When the u-th hidden layer neuron satisfies I u→y =min{I}andR u→y =max{R y }, that is, when the contribution and correlation of the u-th hidden layer neuron to the output layer neuron are both the smallest, calculate the correlation coefficient R between the u-th hidden layer neuron and all other hidden layer neurons uo ; Φ u (i,j) and Φ o (i,j) represent the matrix Φ u and the matrix Φ o The i-th row and j-th column of represent the outputs of the u-th and o-th hidden layer neurons in a sliding window respectively; S(i,j) and T(i,j) represent the i-th row and j-th column of the matrix S and the matrix T respectively, and S and T are the probability distribution matrices of the outputs of the u-th and o-th hidden layer neurons respectively; ∑ i,j Φ u (i,j) and ∑ i,j Φ o (i,j) represent the matrix Φ u and the matrix Φ o Sum all rows and columns, u,o=1,...,p, p is the number of neurons in the hidden layer On this basis, define the correlation coefficient R between the uth and oth hidden layer neurons u→o for: ∑ i,j S(i,j) means summing all rows and columns of matrix S Remember the matrix R uo =[R u→1 ,...R u→u-1 ,R u→u+1 ,...R u→p ] (22) When R u→o =min{R uo }, that is, when the correlation between the uth and oth hidden layer neurons is the largest, the uth and oth hidden layer neurons are merged into a new neuron; the parameters of the new hidden layer neuron are set as follows: Where e(t) is the prediction error at time t, b new (t), c new (t), w new (t) and h new (t) represent the radius, center, weight and output of the newly added hidden layer neuron at time t; v new (t) represents the connection weight between the newly added hidden layer neuron and the lth hidden layer neuron at time t; b u (t) and b o (t) represents the radius of the uth and oth hidden layer neurons at time t respectively; c u (t) and c o (t) represents the center of the uth and oth hidden layer neurons at time t respectively; w u (t) and w o (t) represent the weights of the uth and oth hidden layer neurons at time t, respectively, h u (t) and h o (t) represents the output of the uth and oth hidden layer neurons at time t respectively; v u (t) represents the connection weight between the u-th hidden layer neuron and the output at time t; Case 3. Neurons keep the rules When neither case 1 nor case 2 occurs or case 1 and case 2 are triggered at the same time, all hidden layer neurons are retained and the base learner structure does not change; By using the self-organizing mechanism, hidden layer neurons are added or pruned, so that the base learner structure changes dynamically during the training process, simulating the characteristics of dynamic changes in neuron activity, which is more conducive to adaptive learning in a dynamic environment; (3) Training base learner: There are four parameters that need to be learned in the self-organizing recursive neural network during the training phase, namely b(t), c(t), w(t) and v(t) Define the loss function: The four parameter update processes are as follows: η and θ are the learning rate and momentum factor respectively, and the value range of η and θ is between 0 and 1 (4) Integrate all base learners; the training process of a single base learner of the Adaboost algorithm based on the regression error rate is as follows: Input: For a training sample set S with multiple inputs and a single output m S m ={{X1,Y1},{X2,Y2},...,{X N ,Y N }} (33) X i =[x1,x2,...x n ],i=1,…,N (34) N is the number of samples; n is the dimension of the neural network input feature, n = 4; M is the number of base learners, set to 10; X i represents the neural network input, which are the total amount of primary air, the total amount of secondary air, the grate speed of the drying section, and the grate speed of the combustion section. i Represents the true value of FT corresponding to the i-th group of input Integrated output: α m represents the weight of the mth base learner, f m (X) represents the output of the mth base learner, and H(X) represents the final integrated output Defining the Matrix delta=[delta1,delta2,...delta N ] (36) c is a constant, ranging from 0.1 to 5; N is the number of samples, Y i represents the i-th output value, represents the mean of the output; delta i The absolute value of the prediction error Compare, if Then the logic matrix I m The corresponding element I mi Assign 1, otherwise assign 0; y i represents the true output corresponding to the i-th input, represents the neural network prediction output corresponding to the i-th input of the m-th base learner; ① Initialize the sample weight distribution matrix D1 D1=[d 11 ,d 12 ,...d 1N ],d 1i =1 / N,i=1,...,N (38) d 1i Represents the weight corresponding to the i-th sample in the first base learner, and N is the number of samples ②for m=1to M ③Train the base learner according to the sample distribution weight matrix ④Calculate the regression error rate delta i Determined by formula (37) where d mi is the weight corresponding to the i-th sample in the m-th base learner, I m =[I m1 ,I m2 ,...,I mN ] is a logical matrix used to record the training quality of each sample in the mth base learner. The matrix I m In the example, the 1 position indicates a poorly trained sample, and the 0 position indicates a well-trained sample; e m Represents the regression error rate of the mth base learner; ⑤Calculate the weight of the base learner ⑥Update sample weights Z m is a normalization factor, which is used to ensure that the sum of the weights of all samples in the sample set is 1; Z m The calculation formula of is shown in formula (45) D m (i) = [d m1 ,d m2 ,...d mN ] and D m+1 (i) represents the sample weight distribution matrix of the mth and m+1th base learners respectively; d mi is the weight corresponding to the i-th sample in the m-th base learner, which is also used in generating the sample set S m+1 When , the probability of the sample being sampled; by S m A weighted sampling strategy is adopted to generate an updated sample set S m+1 , used for training and updating the m+1th base learner; S m ={{X1,Y1},{X2,Y2},...,{X N ,Y N }} represents the multi-input and single-output training sample set corresponding to the mth base learner The elements in are only -1 and 1. Then Assign a value of -1 and calculate it by formula (42), thereby increasing d mi , that is, increase the weight corresponding to the i-th sample in the m-th base learner, otherwise, Assign a value of 1, thereby reducing d mi ; y i represents the true output corresponding to the i-th input, represents the neural network prediction output corresponding to the i-th input of the m-th base learner; in is a normalization factor used to ensure that the sum of the weights of all samples in the sample set is 1 H m (X)=α m f m (X) (46) H m (X) is the output of the mth base learner.