Transform and EWC-based boiler combustion system modeling method

The boiler combustion system model is constructed through Transformer and EWC algorithms, which solves the problem of high-precision multi-step prediction of the boiler combustion system, and realizes stable prediction in a time-varying environment, which is suitable for boiler combustion optimization control.

CN120373112APending Publication Date: 2025-07-25SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510469356.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

It is difficult to establish a high-precision multi-step prediction model suitable for boiler combustion systems in the prior art, and the offline model cannot adapt to changes in coal quality and equipment characteristics when applied online, resulting in performance deterioration.

Method used

The Transformer model is combined with the EWC algorithm, and the boiler combustion system model is built through offline training and online update methods, and the cache queue and self-attention mechanism are used to make multi-step predictions, and the changes in important parameters during the network update process are restricted through the EWC algorithm.

Benefits of technology

It realizes high-precision multi-step prediction in boiler combustion system, can adapt to time-varying environments, maintain long-term and stable prediction performance, and provides a foundation for closed-loop combustion optimization control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373112A_ABST
    Figure CN120373112A_ABST
Patent Text Reader

Abstract

The invention discloses a boiler combustion system modeling method based on a Transform model and EWC, and the method comprises the following steps: collecting offline historical data of a boiler combustion system, and organizing to form a training sample; training the Transform multi-step prediction model to obtain a second derivative of a loss function to each parameter of the model, and taking the second derivative as an initialized importance weight corresponding to the model parameters; constructing model input by utilizing a cache queue, wherein the model input is divided into an encoder input part and a decoder input part; the Transform multi-step prediction model carries out recursive multi-step prediction according to the input of the codec; and the model parameters and the importance weights of the model parameters are updated online through the elastic weight consolidation EWC algorithm. According to the method, the accurate boiler combustion system model can be established, the problem of performance degradation of an off-line model during online application is effectively solved, the model can keep high-precision multi-step prediction performance in the environment of real-time change of coal quality and equipment characteristics, and a stable and accurate dynamic model can be provided for realizing closed-loop combustion optimization control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of thermal automatic control, and particularly relates to a modeling method for a boiler combustion system based on Transformer and EWC. Background Art

[0002] Combustion optimization systems that are mainly data-driven for boiler combustion process modeling and combine heuristic search algorithms to optimize combustion-related operation quantities have been widely studied. As a complex high-dimensional input-output dynamic system, the boiler combustion system also has characteristics such as nonlinearity and large lag. For such objects, the traditional methods consider too few factors, so the traditional modeling methods are not suitable for the overall analysis and modeling of the boiler combustion system.

[0003] Establishing an accurate and reliable dynamic model of the boiler combustion system is the basis for realizing combustion optimization. To cope with the large delay and time-varying characteristics of the boiler combustion system and achieve closed-loop optimization control, this model needs to have the following characteristics: 1) Good multi-step prediction performance; 2) Online self-learning ability. However, due to the large number of control variables in the boiler object and the large delay and inertia between the control variables and the output. Conventional steady-state modeling methods, such as support vector machines and BP neural networks, are difficult to capture the temporal dependence relationship between variables, so the multi-step prediction performance is poor. In addition, the boiler combustion system shows obvious time-varying characteristics. The current modeling research work does not update the model after establishing the initial model. This approach will cause the model deviation to increase with time and cannot cope with disturbances such as coal quality changes and load changes, making it difficult to meet the requirements of on-line operation in industrial fields. Summary of the Invention

[0004] Object of the Invention: In order to overcome the deficiencies in the prior art, a modeling method for a boiler combustion system based on Transformer and Elastic Weight Consolidation (EWC) is provided. It is applicable to the non-stationary data stream scenario of the boiler combustion system, and realizes stable and high-precision multi-step prediction of the combustion system target through fewer update times, laying a foundation for realizing closed-loop combustion optimization control.

[0005] Technical Solution: To achieve the above object, the present invention provides a modeling method for a boiler combustion system based on the Transformer model and EWC, including the following steps:

[0006] S1: Collect historical data of the offline boiler combustion system and organize it into training samples;

[0007] S2: Select the sample at the last moment of the training set to train the Transformer multi-step prediction model, obtain the second-order derivative of the loss function with respect to each parameter of the model, and use it as the model parameter θ iThe corresponding initialized importance weights;

[0008] S3: Collect real-time data and save it in a buffer queue with a fixed length; construct the model input at time t using the buffer queue, which is divided into two parts: the encoder input and the decoder input;

[0009] S4: The Transformer multi-step prediction model performs recursive multi-step prediction based on the encoder-decoder input to obtain the multi-step prediction results of a certain output variable at time t d is the encoding length, that is, the number of steps of multi-step prediction;

[0010] S5: At time t+1, calculate the single-step prediction error of the prediction results at time t If the prediction error is greater than the set threshold th, online update the model parameters and the importance weights of the model parameters through the Elastic Weight Consolidation (EWC) algorithm.

[0011] Furthermore, in step S1, considering the encoder-decoder time span setting and the input variable dimension, the encoder input at time t during the training process is represented as [x t-e:t-1 e×17 , the decoder input is represented as [x' t:t+d d×15 , where the unit load and the target quantity are unknown, so x' represents the set of input variables after removing the unit load and the target value, the true output is [y t+1:t+d , and the predicted output is The iterator generates training samples in batch form from the historical data (X, Y) of the boiler combustion system according to the input-output format.

[0012] Furthermore, in step S2, the Transformer multi-step prediction model is completely based on the self-attention mechanism and consists of an encoder and a decoder. The encoder encodes the known historical information sequence into a fixed-dimensional tensor containing context information, and the decoder combines the context and additional information to generate the multi-step prediction results of the target variable; the core computational unit of the Transformer multi-step prediction model is the self-attention mechanism, which converts the input matrix into a query matrix Q, a key matrix K, and a value matrix V through three groups of different linear transformations respectively, and its calculation formula is as follows:

[0013] Q = Linear(X) = XW Q (1)

[0014] K = Linear(X) = XW K (2)

[0015] V = Linear(X) = XW V (3)

[0016] Among them,​​ d k is the dimension of each row of the weight matrix; the multi-head self-attention mechanism combines different sample information learned based on the same attention mechanism. The specific calculation process is as follows: when given the same query, key, and value, Q, K, and V are transformed through multiple groups of different linear transformations and are fed into the scaled dot-product attention operation in parallel. Finally, the outputs of multiple groups of attention calculations are concatenated together to produce the final output. The calculation formula is as follows:

[0017] MutiHead(Q, K, V) = Concat(head1, head2, …, head h ,)W 0 (4)

[0018]

[0019] where W 0 is the linear transformation matrix after concatenating the multi-head self-attention matrices, are the linear transformation weight matrices of the i-th head of Q, K, and V respectively.

[0020] Furthermore, in step S2, the Transformer multi-step prediction model omits the embedding step and fits the predicted value through a linear fully connected layer.

[0021] Furthermore, in step S2, the training sample data is divided into a training set, a validation set, and a test set; using the grid search technique, with the mean absolute error (MAE) of the validation set prediction results as an indicator, the network hyperparameters are optimized, and an offline Transformer multi-step prediction model is established for each output indicator respectively;

[0022] The definition of the mean absolute error is:

[0023]

[0024] In the formula, h is the number of forward prediction steps, and N is the total number of samples in the evaluation time domain; the average value of MAE at each prediction step will be used as the standard for selecting the optimal model parameters;

[0025] The loss function used in the process of training the model is defined as:

[0026]

[0027] In the formula, θ is the model parameter; Loss is the loss function.

[0028] Furthermore, the process of updating the model parameters in step S3 includes:

[0029] A1: Construct a training set (X t , Y t ) for the model using the data in the cache queue L;

[0030] A2: Use the constructed (X t , Y t ) to perform retraining for epochs times; epochs is recommended to be set to 1 to reduce computational overhead;

[0031] A3: Calculate the value of the loss function when the model is updated online. The loss function consists of two parts: the prediction error MAE and the penalty term for model parameter changes;

[0032] A4: Perform error backpropagation to calculate the gradient of the loss function with respect to each parameter of the model;

[0033] A5: Use the Adam algorithm to update the model parameters.

[0034] Furthermore, the process of updating the importance weights of the model parameters in step S3 includes:

[0035] B1: Construct the model input from t to d using the data in the cache queue L, including the encoder input x t-d-e:t-d-1 and the decoder input x t-d:t-1 ;

[0036] B2: Use the backpropagation algorithm to calculate the second derivative of the loss function with respect to each parameter of the model and evaluate the importance weights of each model;

[0037] B3: Update the importance weights of each parameter.

[0038] Furthermore, the loss function in step A3 is expressed as:

[0039]

[0040] In the formula, (X t , Y t ) is the training set constructed from the data in the current cache queue; is the parameter before model update, including the weights and biases of the network; λ is the regularization coefficient; Ω i is the importance weight of the current model parameter θ i .

[0041] Furthermore, for the model input at time x - d in step B2, the second derivative of the EWC loss function with respect to each parameter of the model is used to estimate the importance weights of each parameter of the model:

[0042]

[0043] In the formula, Ω iDenote the model parameter as θ i with the importance weight; Loss(θ) is the error loss between the multi-step prediction value and the actual value of the model at time t-1.

[0044] The present invention provides a modeling method for a boiler combustion system based on Transformer and EWC, which has the following two characteristics: (1) A multi-step prediction model of the boiler combustion system is established by using Transformer, which can accurately predict the change trend of the output index of the boiler combustion system in the future for a period of time; (2) An online learning method applicable to the time-varying scenario of the object is introduced, and the EWC regularization technology is adopted to achieve multi-step prediction of stable, continuous and high-precision output indexes through fewer model update times, laying a foundation for realizing a closed-loop combustion optimization control system for long-term online operation.

[0045] The method of the present invention proposes an online update framework for the combustion system model aiming at the boiler combustion system with time-varying characteristics. First, an offline model based on Transformer is established, and the unit load, the set value of the oxygen content at the economizer outlet, the coal feeding amount of each layer of coal feeder, the opening degree of each layer of secondary air damper, the opening degree of the burnout air damper, and the set value of the primary air pressure are selected as input features, and the flue gas temperature, the CO concentration at the SCR inlet, the NOx concentration at the SCR inlet, and the reheat steam temperature are selected as output features. In the online operation link, when it is detected that the model prediction performance declines, the training set is constructed by using the samples in the buffer area to complete the model retraining, so that the model can track the dynamic characteristics of the boiler in real time. It should be noted that in the retraining link, by introducing the regularization technology, the change of important parameters in the network update process is restricted, and the stability of the network prediction performance is maintained. The method of the present invention can provide a stable and accurate dynamic model for realizing closed-loop combustion optimization control, and can be applied to the dynamic modeling and optimization control of other complex and time-varying industrial processes.

[0046] The present invention can establish an accurate boiler combustion system model, effectively solve the problem of performance degradation of the offline model in online applications, enable the model to maintain high-precision multi-step prediction performance in an environment where the coal quality and equipment characteristics change in real time, and can provide a stable and accurate dynamic model for realizing closed-loop combustion optimization control.

[0047] Beneficial effects: Compared with the prior art, first, an offline model is established by using Transformer, and the network hyperparameters are optimized by using the grid search technology, so that the model has better global generalization ability. At the same time, the self-attention mechanism is more proficient in capturing time series dependencies and processing high-dimensional data than traditional machine learning models, and has significant advantages in the modeling of boiler combustion systems. Secondly, the proposed EWC online learning algorithm enables the model to still maintain high-precision and stable prediction performance when the object characteristics change, laying a foundation for realizing the long-term online operation of the combustion optimization control system. Description of the Drawings

[0048] Figure 1 It is a flowchart of the EWC online learning method for the Transformer model of the present invention;

[0049] Figure 2 It is a curve graph showing the change of training, validation, and offline / online test sample metrics in the present invention;

[0050] Figure 3 It is a curve graph of the multi-step prediction results of each offline modeling method in the present invention for each task of the boiler combustion system;

[0051] Figure 4 It is a curve graph of the multi-step prediction results of the offline Transformer model and the EWC-Transformer model in the present invention for each task of the boiler combustion system. Detailed Embodiments

[0052] The present invention will be further illustrated below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications made by those skilled in the art fall within the scope defined by the appended claims of this application.

[0053] Embodiment 1:

[0054] As Figure 1 shown, this embodiment provides a modeling method for a boiler combustion system based on a Transformer model and EWC, including the following steps:

[0055] S1: Collect historical data of the offline boiler combustion system and organize it into training samples;

[0056] In this embodiment, according to engineering experience, the unit load, the set value of the oxygen content at the economizer outlet, the coal feeding amount of each coal feeder, the opening degree of each secondary air damper, the opening degree of the burnout air damper, and the set value of the primary air pressure are selected as input variables. The flue gas temperature, the CO concentration at the SCR inlet, the NOx concentration at the SCR inlet, and the reheat steam temperature are selected as the output variables of the model. Transformer models based on the encoder-decoder architecture and the self-attention mechanism are established for each output variable respectively;

[0057] According to the response transition time of the combustion system output to each operation amount, the input time span e of the encoder is determined, and according to the multi-step prediction length, the input time span d of the decoder is determined. The input variables and the autoregressive term are used as the encoder input, and its feature dimension is 17. In the actual multi-step prediction process, it is set that the unit load remains unchanged and the future target value y is unknown, so the decoder input feature dimension is 15, and the training samples are organized according to the above settings;

[0058] Considering the encoder-decoder time span setting and the dimension of input variables, the encoder input at time t during the training process is represented as [x t-e:t-1 e×17 , and the decoder input is represented as [x' t:t+d d×15 , where the unit load and the target value are unknown. Therefore, x' represents the set of input variables after removing the unit load and the target value. The true output is [y t+1:t+d , and the predicted output is The iterator generates training samples in batch form from the historical data (X, Y) of the boiler combustion system according to the input-output format;

[0059] The future operation amount [x' t:t+d d×15 included in the decoder input can be obtained by online solving the optimization problem regarding the prediction target. These variables adopt actual observed values during the multi-step prediction process of the target value to evaluate the multi-step prediction performance of the model.

[0060] S2: Select the samples at the last moment of the training set to train the Transformer multi-step prediction model, obtain the second-order derivatives of the loss function with respect to each parameter of the model, and use them as the initial importance weights corresponding to the model parameters θ i ;

[0061] The Transformer multi-step prediction model is completely based on the self-attention mechanism and consists of an encoder and a decoder. The encoder encodes a sequence of known historical information into a fixed-dimensional tensor containing context information, and the decoder combines the context and additional information to generate multi-step prediction results of the target variable. The core computational unit of the Transformer multi-step prediction model is the self-attention mechanism, which converts the input matrix into a query matrix Q, a key matrix K, and a value matrix V through three groups of different linear transformations respectively. The calculation formulas are as follows:

[0062] Q = Linear(X) = XW Q (1)

[0063] K = Linear(X) = XW K (2)

[0064] V = Linear(X) = XW V (3) where, d k ​​​is the dimension of each row of the weight matrix; different sample information learned based on the same attention mechanism is combined through the multi-head self-attention mechanism. The specific calculation process is as follows: When given the same query, key, and value, Q, K, and V are transformed through multiple groups of different linear transformations and are fed into the scaled dot-product attention operation in parallel. Finally, the outputs of multiple groups of attention calculations are concatenated together to produce the final output. The calculation formula is as follows:

[0065] MutiHead(Q, K, V) = Concat(head1, head2, …, head h ,)W 0 (4)

[0066]

[0067] where W 0 is the linear transformation matrix after concatenating the multi-head self-attention matrices, are the linear transformation weight matrices of the i-th head of Q, K, and V respectively.

[0068] To enable the Transformer model to achieve the goal of predicting the relevant indicators of the boiler combustion system, this study made certain changes to the original Transformer model architecture. Without changing the overall encoder-decoder structure, since the sample inputs of the encoder and decoder are in the form of feature vectors, the embedding step is omitted in this model. Secondly, since this model is used for regression tasks, the softmax is not used to map the output values into probability form at the output of the model. Instead, the predicted values are fitted through a linear fully connected layer.

[0069] The training sample data is divided into a training set, a validation set, and a test set; using the grid search technique, with the mean absolute error (MAE) of the prediction results of the validation set as the indicator, the network hyperparameters are optimized, and an offline Transformer multi-step prediction model is established for each output indicator;

[0070] The definition of the mean absolute error is:

[0071]

[0072] In the formula, h is the number of forward prediction steps, and N is the total number of samples in the evaluation time domain; the average value of MAE at each prediction step will be used as the criterion for selecting the optimal model parameters;

[0073] The loss function used in the process of training the model is defined as:

[0074]

[0075] In the formula, θ is the model parameter; Loss is the loss function.

[0076] The hyperparameters of the Transformer model are optimized using offline samples and grid search techniques. The optimized hyperparameters include: the number of encoder layers, the number of decoder layers, and the number of heads in the multi-head self-attention mechanism. The model is retrained 10 times at each grid point, and the optimal model parameters are selected based on the accuracy of the prediction results on the validation set. The remaining common parameter settings of the model are as follows: the activation function is set to ReLU, the batch size is set to 128, the maximum number of epochs is set to 100, the optimization algorithm is set to Adam, the learning rate is set to 0.001, and the loss function is set to L1Loss (i.e., MAE).

[0077] S3: Assume that the current time is t. Collect real-time data and save it in a fixed-length buffer queue L. Use the buffer queue L to construct the model input at time t, which is divided into two parts: the encoder input and the decoder input.

[0078] S4: The Transformer multi-step prediction model performs recursive multi-step prediction based on the encoder-decoder input to obtain the multi-step prediction results of a certain output variable at time t. d is the encoding length, that is, the number of steps for multi-step prediction.

[0079] Different from traditional recursive multi-step prediction methods, in this embodiment, although the multi-step prediction of the Transformer model is in a recursive form, the previous prediction results are not used during multi-step prediction, but only the decoder input at the time before the current prediction step is considered. Therefore, the Transformer model avoids the occurrence of error accumulation during multi-step prediction.

[0080] S5: At time t + 1, calculate the single-step prediction error of the prediction results at time t. If the prediction error is greater than the set threshold th, online update the model parameters and the importance weights of the model parameters through the Elastic Weight Consolidation (EWC) algorithm.

[0081] The process of updating the model parameters includes:

[0082] A1: Use the data in the buffer queue L to construct the training set (X t , Y t ) of the model;

[0083] A2: Use the constructed (X t , Y t ) to perform retraining for epochs times; epochs is recommended to be set to 1 to reduce the computational overhead;

[0084] A3: Calculate the value of the loss function during online model update. The loss function consists of two parts: the prediction error MAE and the penalty term for changes in model parameters;

[0085] In this embodiment, the EWC algorithm is used to retrain the model to maintain the stability of the important parameters of the network. In the scenario of non-stationary data streams, directly retraining the model with the recently cached samples will cause the model to only focus on a certain local process of the system, resulting in the loss of the global generalization ability of the model. The EWC algorithm adds a penalty term to the loss function during the online update of the model to reduce the change of the important parameters in the network structure. The loss function of the model update process can be expressed as:

[0086]

[0087] In the formula, (X t , Y t ) is the training set constructed from the data in the current cache queue; is the parameter before model update, including the weights and biases of the network; λ is the regularization coefficient; Ω i is the importance weight of the current model parameter θ i .

[0088] A4: Error backpropagation, calculating the gradient of the loss function with respect to each parameter of the model;

[0089] A5: Using the Adam algorithm to update the model parameters.

[0090] The process of updating the importance weights of the model parameters includes:

[0091] B1: Using the data in the cache queue L to construct the model input at time t - d, including the encoder input x t-d-e:t-d-1 and the decoder input x t-d:t-1 ;

[0092] B2: Using the backpropagation algorithm to calculate the second derivative of the loss function with respect to each parameter of the model, and evaluating the importance weights of each model;

[0093] For the model input, EWC uses the second derivative of the loss function with respect to each parameter of the model to estimate the importance weights of each parameter of the model:

[0094]

[0095] In the formula, Ω i represents the importance weight of the model parameter θ i ; Loss(θ) is the error loss between the multi-step prediction value and the actual value of the model at time t - d.

[0096] B3: Updating the importance weights of each parameter.

[0097] Embodiment 2:

[0098] In this embodiment, the method of the present invention is applied to establish a dynamic model of a 1000MW supercritical coal-fired boiler combustion system and perform online updates to achieve real-time high-precision multi-step prediction. The basic process of the online learning method refers to Figure 1 , and the specific implementation process is as follows:

[0099] 1) Select the inputs and outputs of the prediction model and perform data preprocessing. In this embodiment, the boiler object adopts a tangentially fired combustion method with four corners and is equipped with six medium-speed coal mills. The modeling data comes from the distributed control system of the power plant, with a sampling period of 40s and a total of 20,000 groups. The inputs of the model are considered as the main operating variables of the boiler combustion system, unit load, and autoregressive terms, totaling 17 dimensions, and the outputs are the SCR inlet NOx concentration, SCR inlet CO concentration, flue gas temperature, and reheat steam temperature.

[0100] Table 1 Modeling variables and variation ranges

[0101]

[0102] 2) Data partitioning and training set construction. The sample data is divided into two parts. The first part contains the first 12,000 samples. The first 60% of the data is set as the training set for offline model building of the model; the subsequent 10% of the data is set as the validation set for optimizing the model hyperparameters, and the last 30% of the data is set as the test set for evaluating the advantages and disadvantages among different offline models. This partitioning ensures that both the training set and the test set contain different operating conditions, avoiding the influence of different sample distributions on the model accuracy and generalization. The second part contains the subsequent 8,000 samples for evaluating the online update effect of the Transformer model. As Figure 2 shown, the selected samples cover the range of stable operation of the boiler unit and frequent load changes.

[0103] Considering that the response transition time of the combustion system output to each operation amount is within 10 minutes, the input sample dimension enc in of the encoder is set to 15. Since introducing historical autoregressive terms can significantly improve the prediction level, the state variables, control variables, and autoregressive terms in Table 2-4 are used as the decoder input, and the feature dimension enc feature is 17. For forward 10-minute multi-step prediction with an interval of 40s, the decoder input sample dimension dec in is also set to 15. Since the boiler load is set to remain unchanged during the actual multi-step prediction process and the future target value y is unknown, the decoder input feature dimension dec feature is set to 15. Therefore, during the training process, the encoder input x enc is represented as [x t-15:t-1 15×18 , and the decoder input x dec is represented as [x' t:t+14 15×16 ​​, where x' represents the set of input variables after removing the LOAD and the target value, and the true label is [y t+1:t+15 , and the predicted output is

[0104] 3) Parameter setting and hyperparameter optimization. Table 2 shows the hyperparameters and search ranges of the Transformer model. enclayers and dec layers are the number of encoder layers and decoder layers respectively, d model is the dimension after the sample features are dimensionally elevated, and num heads is the number of heads in the multi-head self-attention mechanism.

[0105] Table 2 Hyperparameter Search Ranges of Transformer

[0106]

[0107] The optimal hyperparameter combination is searched by minimizing the losses of the training set and the validation set. The optimization results are as follows: the number of encoder layers enc_num_layers = 3, the number of decoder layers dec_num_layers = 2, the feature dimension after the sample features are dimensionally elevated is d model = 64, and the number of heads in the multi-head self-attention mechanism is num_heads = 8.

[0108] 4) Comparison methods. To evaluate the effectiveness of the method of the present invention, the following methods are compared and analyzed.

[0109] LSSVM: Use the incremental least squares support vector machine algorithm. It adopts the radial basis kernel function, the sample size is set to 1000, and the penalty coefficient c and the kernel parameter g are also optimized by using the offline samples and the grid search technique. The search ranges are both [1e-3, 1e-2,..., 1e2, 1e3], and the optimization results are c = 1000 and g = 0.01.

[0110] LSTM-Seq2Seq: An encoder-decoder architecture based on the long short-term memory network. Similar to the Transformer model, enc in is set to 15, enc feature is 17, dec in is set to 15, and dec feature is 15. The number of LSTM layers is an important hyperparameter that determines the neural network structure. Similarly, grid search optimization is performed on the number of encoder hidden layers enc layers, the number of decoder hidden layers declayers, the number of encoder hidden layer nodes enc hides, and the number of decoder hidden layer nodes dec hides. The optimization results are enc layers = 2, dec layers = 1, enc hides = 64, and dec hides = 64.

[0111] Transformer: The Transformer model established offline is not updated.

[0112] EWC-Transformer (the method of the present invention): The online Transformer model with the EWC update method added.

[0113] Based on the above scheme, in order to verify the effectiveness of the Transformer modeling method, the LSSVM and LSTM-Seq2Seq models are compared in the offline test set. There are a total of 2,000 groups of test samples, which are used to evaluate the multi-step prediction performance of each offline model, and the prediction step length is 15 steps. Table 3 statistically shows the MAE indicators of the multi-step prediction results of each model. In the table, h = 1:15 represents the average value between the prediction steps from 1 to 15.

[0114] Table 3 Comparison of multi-step prediction performance of different modeling methods for each prediction task

[0115]

[0116] As can be seen from Table 3, compared with the other two models, the Transformer model has no advantage in the single-step prediction accuracy, but the overall accuracy of the multi-step prediction of this model in the four prediction tasks is much higher than that of the other two models. In the prediction time domain of 2 to 5 steps, except for the reheat steam temperature, the MAE is less than that of the other two models. In the prediction time domain of 6 - 15 steps, the MAE of the model in the four tasks is much lower than that of the other models. It can be seen that the Transformer model introducing the self-attention mechanism can more easily learn the long-distance dependencies in the sequence and has excellent performance in the multi-step prediction of the time series of each index of the boiler combustion system.

[0117] Figure 3 Shows the multi-step prediction result curves of the LSSVM, LSTM-Seq2Seq, and Transformer modeling methods for each task of the boiler combustion system. The black solid line in the figure represents 500 groups of continuous historical observations, and the rest are the forward 15-step prediction curves of each comparison model. From Figure 3 it can be seen that in each task, the Transformer has the best performance in dynamically capturing the future trend of the curve.

[0118] To verify the effectiveness of the EWC online model updating method, the Transformer and EWC-Transformer models are compared in the online test set. There are a total of 8,000 groups of test samples, which are used to evaluate the multi-step prediction performance of each offline model and online model. As can be seen from Table 4, for the flue gas temperature prediction task, the single-step prediction MAE of the EWC-Transformer model reaches 0.088 °C, and the average MAE within the prediction time domain reaches 0.411 °C. Although it is weaker than the performance of the model on the offline test set, it is significantly better than the case where the model is not updated. For the CO, NOx, and reheated steam temperature prediction tasks, the EWC update algorithm can also improve the single-step and multi-step prediction accuracies. In terms of the mean absolute error and goodness-of-fit index, it is better than the offline model, indicating that the Transformer model combined with the EWC update method can timely adjust the model parameters to adapt to the actual changes of the boiler combustion system during long-term operation, ensuring that the Transformer model has the ability of long-term high-precision dynamic multi-step prediction. Therefore, it can be used as the model update method for online application of the combustion optimization system.

[0119] Table 4 Multi-step prediction accuracies of the Transformer model in offline and online forms in the online test set

[0120]

[0121] Figure 4 The multi-step prediction curves of the Transformer and EWC-Transformer models and the real curves at corresponding moments are plotted for each task. In the prediction curves of flue gas temperature, CO, NOx, and reheated steam temperature, the prediction trend of the EWC-Transformer method basically coincides with the observed trend and is better than the prediction effect of the non-updated form of the Transformer model, indicating that the EWC algorithm effectively solves the model mismatch problem caused by the change of boiler state through retraining the model.

[0122] The above results all prove that the boiler combustion system modeling method based on Transformer and EWC of the present invention can achieve multi-step prediction of high-precision combustion system output indicators, can quickly adapt to the non-stationary environment, and shows long-term stable multi-step prediction performance.

Claims

1. A modeling method for a boiler combustion system based on the Transformer model and EWC, characterized in that, It includes the following steps: S1: Collect historical data of the offline boiler combustion system and organize it to form training samples; S2: Select the samples at the last moment of the training set to train the Transformer multi-step prediction model, obtain the second-order derivatives of the loss function with respect to each parameter of the model, and use them as the model parameters θ i The corresponding initialized importance weights; S3: Collect real-time data and save it in a buffer queue with a fixed length; Use the buffer queue to construct the model input at time t, which is divided into two parts: encoder input and decoder input; S4: The Transformer multi-step prediction model performs recursive multi-step prediction based on the encoder-decoder input to obtain the multi-step prediction results of a certain output variable at time t. d is the encoding length, that is, the number of steps for multi-step prediction. S5: At time t + 1, calculate the single-step prediction error of the prediction result at time t If the prediction error is greater than the set threshold th, online update the model parameters and the importance weights of the model parameters through the Elastic Weight Consolidation (EWC) algorithm 2. The modeling method of a boiler combustion system based on the Transformer model and EWC according to claim 1, characterized in that, In step S1, considering the encoder-decoder time span setting and the input variable dimension, the encoder input at time t during the training process is represented as [x t-e:t-1 e×17 , and the decoder input is represented as [x' t:t+d d×15 . Since the unit load and the target quantity are unknown, x' represents the set of input variables after removing the unit load and the target value. The true output is [y t+1:t+d , and the predicted output is The iterator generates batch-form training samples from the historical data (X, Y) of the boiler combustion system in the input-output format.​​ 3. A method for modeling a boiler combustion system based on the Transformer model and EWC according to claim 1, characterized in that, In step S2, the Transformer multi-step prediction model is completely based on the self-attention mechanism and consists of an encoder and a decoder. The encoder encodes the known historical information sequence into a fixed-dimensional tensor containing context information, and the decoder combines the context and additional information to generate multi-step prediction results of the target variable; The core computational unit of the Transformer multi-step prediction model is the self-attention mechanism, which converts the input matrix into a query matrix Q, a key matrix K, and a value matrix V through three groups of different linear transformations respectively. The calculation formula is as follows: Q = Linear(X) = XW Q (1) K = Linear(X) = XW K (2) V = Linear(X) = XW V (3) Among them, d k is the dimension of each row of the weight matrix; different sample information learned based on the same attention mechanism is combined through the multi-head self-attention mechanism. The specific calculation process is as follows: when given the same query, key, and value, Q, K, and V are transformed through multiple groups of different linear transformations and are fed into the scaled dot-product attention operation in parallel. Finally, the outputs of multiple groups of attention calculations are concatenated together to produce the final output. The calculation formula is as follows: MutiHead(Q,K,V)=Concat(head1,head2,…,head h ,)W 0 (4) Among them, W 0 is the linear transformation matrix after the concatenation of the multi-head self-attention matrices, are the linear transformation weight matrices of the i-th heads of Q, K, and V respectively.

4. A modeling method for a boiler combustion system based on the Transformer model and EWC according to claim 3, characterized in that, In step S2, the Transformer multi-step prediction model omits the embedding step and fits the predicted value through a linear fully connected layer.

5. A modeling method for a boiler combustion system based on the Transformer model and EWC according to claim 4, characterized in that, In step S2, the training sample data is divided into a training set, a validation set, and a test set; Use the grid search technique to optimize the network hyperparameters with the mean absolute error (MAE) of the prediction results of the validation set as the index, and establish an offline Transformer multi-step prediction model for each output index respectively; The definition of the mean absolute error is: In the formula, h is the number of forward prediction steps, and N is the total number of samples in the evaluation time domain; The average value of MAE at each prediction step will be used as the standard for selecting the optimal model parameters; The loss function used in the process of training the model is defined as: In the formula, θ is the model parameter; Loss is the loss function.

6. A method for modeling a boiler combustion system based on the Transformer model and EWC according to claim 5, characterized in that, The process of updating the model parameters in step S3 includes: A1: Construct a training set (X t , Y t ) of the model using the data in the cache queue L; A2: Use the constructed (X t , Y t ) to perform retraining for epochs times; A3: Calculate the value of the loss function during the online update of the model. The loss function consists of two parts: the prediction error MAE and the penalty term for the change of model parameters; A4: Perform error backpropagation to calculate the gradient of the loss function with respect to each parameter of the model; A5: Use the Adam algorithm to update the model parameters.

7. A modeling method for a boiler combustion system based on the Transformer model and EWC according to claim 6, characterized in that, The process of updating the importance weights of the model parameters in step S3 includes: B1: Construct the model input at time t-d using the data in the cache queue L, including the encoder input x t-d-e:t-d-1 and the decoder input x t-d:t-1 ; B2: Use the backpropagation algorithm to calculate the second derivative of the loss function with respect to each parameter of the model and evaluate the importance weights of each model; B3: Update the importance weights of each parameter.

8. A modeling method for a boiler combustion system based on the Transformer model and EWC according to claim 6, characterized in that, The loss function in step A3 is expressed as: Where, (X t , Y t ) is the training set constructed from the data in the current cache queue; is the parameter before model update, including the weights and biases of the network; λ is the regularization coefficient; Ω i is the importance weight of the current model parameter θ i .

9. A modeling method for a boiler combustion system based on the Transformer model and EWC according to claim 7, characterized in that, In step B2, for the model input at time x - d, the second derivative of the EWC loss function with respect to each parameter of the model is used to estimate the importance weights of each parameter of the model: where, Ω i represents the importance weight of the model parameter θ i ; Loss(θ) is the error loss between the multi-step prediction value and the actual value of the model at time t - 1.