A method for predicting the amount of flue gas pollutants generated

By adopting an improved multi-level bidirectional multi-head attention module and an adaptive gating mechanism in the flue gas pollution collection generation prediction method, combining the smooth loss function and the Elastic-Net regularization term, the problem of difficult to predict the flue gas pollution collection generation in the prior art is solved, and higher prediction accuracy and stability are achieved.

CN118866172BActive Publication Date: 2025-06-06QINGDAO ZHONGYUANBOXIN ENVIRONMENTAL PROTECTION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410872153.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-01
Publication Date
2025-06-06
Estimated Expiration
2044-07-01

AI Technical Summary

Technical Problem

The prior art is difficult to predict the generation of flue gas pollutant by learning existing data, thereby achieving refined management of pollutant emissions and improving the effectiveness of pollution control.

Method used

A method for predicting the generation of flue gas waste collection is proposed, using the threshold number neural network (T3N) model. The model hidden layer is composed of a multi-level bidirectional multi-head attention module, and a residual connection, layer normalization and adaptive gating mechanism are introduced, combining the smooth loss function and Elastic-Net regularization term to reduce model overfitting.

Benefits of technology

Through the improved multi-level bidirectional multi-head attention module and adaptive gating mechanism, the model can more comprehensively capture the dependencies in the data and improve prediction accuracy; the smooth loss function and Elastic-Net regularization terms effectively reduce overfitting, improve the stability and generalization ability of the model, and significantly improve the accuracy and stability of the prediction of the generation of smoke pollution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118866172B_ABST
    Figure CN118866172B_ABST
Patent Text Reader

Abstract

Flue gas pollutants are particulate matter, gaseous pollutants, and other harmful substances generated during the combustion of fossil fuels (such as coal, oil, natural gas). These pollutants include PM2.5, PM10, sulfur dioxide (SO2), nitrogen dioxide (NO2), carbon monoxide (CO), and volatile organic compounds (VOCs), which pose serious hazards to the environment and human health. The present invention proposes a method for predicting the generation amount of flue gas pollutants, which is predicted by the constructed T3N model and smooth loss function. This method can accurately master the data of flue gas generation amount, thereby realizing the real-time regulation and optimization operation of pollutant sources, and providing more effective means for environmental protection and pollution control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of smoke pollutant prevention and control, and in particular relates to a method for predicting the amount of smoke pollutant generated. Background Art

[0002] Flue gas pollutants refer to particulate matter, gaseous pollutants and other harmful substances produced during the combustion of fossil fuels (such as coal, oil, and natural gas). These pollutants include PM2.5, PM10, sulfur dioxide (SO 2 ), nitrogen dioxide (NO 2 ), carbon monoxide (CO) and volatile organic compounds (VOCs), which are seriously harmful to the environment and human health. The generation of pollutants is closely related to the type of fuel, combustion temperature and method. During high-temperature combustion, the chemical components in the fuel generate various pollutants through complex chemical reactions. By predicting the amount of flue gas pollutants generated, combined with real-time regulation and optimization operations, such as physical methods (electrostatic precipitator), chemical methods (wet desulfurization) and biological methods (selective catalytic reduction), it is possible to achieve refined management of pollutant emissions, improve the effect of pollution control, help to more effectively reduce the emission of flue gas pollutants, protect the environment and human health, and promote the development and application of environmental protection technologies. Summary of the invention

[0003] The main purpose of the present invention is to provide a method for predicting the amount of flue gas pollutants generated, aiming to solve the problem of how the existing technology can predict the amount of flue gas pollutants generated by learning from existing data, thereby achieving refined management of pollutant emissions and improving the effect of pollution control.

[0004] To achieve the above object, the present invention provides a method for predicting the amount of flue gas pollutant generation, the method comprising the following steps:

[0005] S1. Obtain the data information required for the flue gas pollutant generation prediction method to form a flue gas pollutant generation data set.

[0006] S2. Preprocessing the current data set, including filling in missing values ​​and constructing a training set and a test set according to the filling results.

[0007] S3. A Threshold Number Neural Network model is proposed, which is referred to as T3N. The hidden layer of the model consists of a multi-level bidirectional multi-head attention module, and the module is improved by introducing residual connections, using layer normalization, and enhancing the adaptive gating mechanism. The module can pay attention to the previous and next information of the input sequence at the same time, so as to capture the dependencies in the data more comprehensively.

[0008] S4. A smooth loss function is proposed. The loss function is related to the set threshold. The weight parameters of the T3N model are adjusted according to the size of the predicted value and the threshold, and the Elastic-Net regularization term is used to reduce the problem of model overfitting. The Elastic-Net regularization term combines the regularization techniques of L1 and L2 regularization, aiming to have both the feature selection ability of L1 regularization and the stability of L2 regularization.

[0009] S5. The output value of the improved multi-level bidirectional multi-head attention module in the T3N model after passing through the fully connected layer is used as the final prediction value of the flue gas pollutants.

[0010] Further, in the step S1, it specifically includes: the data related to the amount of flue gas pollutants generated includes four features, namely: fuel characteristics, combustion process parameters, environmental conditions and operating conditions, wherein the fuel characteristics include fuel type, sulfur content, nitrogen content, ash content, volatile matter content information, different types of fuels (such as coal, oil, natural gas, etc.) will produce different pollutants, combustion process parameters include combustion temperature (the higher the temperature, the more nitrogen oxides are generated), air excess coefficient (excess air will affect the completeness of combustion and carbon monoxide generation), combustion mode (different combustion technologies will affect the amount of pollutants generated) and combustion time (the length of combustion time will affect the completeness of fuel combustion) information, environmental conditions include temperature (ambient temperature affects combustion efficiency and diffusion of pollutants), pressure (ambient pressure may affect the chemical reaction rate during combustion) information, operating conditions include fuel mixing ratio and pollution control equipment (performance and operating status of electrostatic precipitator, desulfurization equipment, etc.), the current features are sent to the prediction model as attributes during the training process; in addition, it also includes flue gas pollutants PM2.5, PM10, sulfur dioxide (SO 2 ), nitrogen dioxide (NO 2 ), carbon monoxide (CO) and volatile organic compounds (VOCs) data values, and during the training process, the harmful gases in the flue gas pollutants are used as the predicted values ​​of the prediction model.

[0011] Furthermore, in step S2, the current data set is preprocessed, specifically including: during the data collection process, the smoke monitoring equipment may be unable to record data due to technical failure, maintenance or damage. In order to improve the data quality, the present invention uses linear interpolation to process the missing values. Given x 0 ,y 0 and x 1 ,y 1 , in x 0 and x 1 The difference y at a point x between is calculated by the following formula:

[0012]

[0013] In the formula, (x 0 ,y 0 ) and (x 1 ,y 1 ) is a known flue gas pollutant data point, x is the point to be interpolated, and x satisfies 0 ≤x≤x 1 , y is the data interpolation of the flue gas pollutants at x;

[0014] The data set after the current preprocessing is segmented and divided into a training set and a test set according to a ratio of 7:3 of the total number of flue gas pollutant data.

[0015] Furthermore, in step S3, a threshold number neural network (T3N) is constructed, and the training set is input into the T3N model. The specific steps are as follows:

[0016] S31. Construct a T3N model. The input layer receives preprocessed feature data, such as fuel characteristics, combustion process parameters, environmental conditions, and operating conditions. The hidden layer uses a multi-bidirectional attention module, and each layer uses a ReLU activation function. The output layer is a plurality of neurons, including a plurality of output nodes, and each output node corresponds to a predicted value of a target threshold number.

[0017] S32, the multi-level bidirectional multi-head attention module is composed of improved multi-head attention and bidirectional attention modules. First, the data output by the input layer is linearly transformed. The formula is as follows:

[0018] Q=XW Q ;

[0019] K=XW K ;

[0020] V=XW V ;

[0021] Wherein, X is the smoke pollutant data after the pretreatment, W Q , W K , W V , is the weight matrix, Q, K, V are the values ​​after linear transformation;

[0022] The attention score calculation formula is as follows:

[0023]

[0024] In the formula, Q, K, and V are the values ​​after linear transformation. is the scaling factor;

[0025] Multi-head attention is achieved by computing h independent attention heads in parallel. The multi-head attention formula is as follows:

[0026] H i =Attention(Q i , K i , V i ), i=1, 2, 3..., h;

[0027] In the formula, h is the number of attention heads, H i is the score of the i-th attention head, Q i is the Q matrix of the i-th attention head, K i is the K matrix of the i-th attention head, V i is the V matrix of the i-th attention head;

[0028] The multi-head attention is concatenated and the final output is obtained through linear transformation:

[0029] MultiHead(Q,K,V)=Concat(H 1 H 2 ......, H h )·W O ;

[0030] Where MultiHead(Q, K, V) is the total output of the multi-head attention. For the bidirectional attention module, the positive multi-head attention is also called the forward attention, that is, H forward =MultiHead(Q, K, V), W O is the weight matrix; residual connection and layer normalization operations are added, and forward and backward attention are calculated for each layer respectively. The forward attention formula of the lth layer is as follows:

[0031] In the formula, is the forward attention weight parameter of the l-1th layer, is the l-th forward multi-head attention, and LayerNorm is the normalization operation on the layer weight parameters;

[0032] The formula for backward attention at layer l is as follows:

[0033]

[0034] in:

[0035] X backward =Reverse(X);

[0036] Q backward =X backward W Q ;

[0037] Kbackward =X backward W K ;

[0038] V backward =X backward W V ;

[0039] In the formula, Reverse is the reverse operation. is the backward attention weight parameter of the l-1th layer.

[0040] S33, the adaptive gating mechanism is used to dynamically adjust the weights of forward and backward attention, so as to flexibly capture important information at different time steps and features. In order to enhance this mechanism, the present invention adopts a complex long short-term memory network (LSTM) to calculate the gating value, which can more effectively process sequence data and introduce more context information in the gating process. The specific steps are as follows:

[0041] S331, the output H of the improved forward attention forward And the output H of the backward attention backward Concatenate as the subsequent input of LSTM:

[0042] H combined =Concat(H forward , H backward );

[0043] Hc ombined(t-1) =H combined ·h t-1 ;

[0044] In the formula, H combined(t-1) It is the output of the t-1 layer network and the input of the t-th layer LSTM network;

[0045] LSTM effectively captures long-term dependencies by introducing three gate mechanisms (input gate, forget gate, and output gate) and a cell state to maintain and adjust hidden states and cell states;

[0046] S332, forget gate determines the previous cell state C t-1 Which information will be discarded? The calculation formula is as follows:

[0047] f t =σ(W f H combined(t-1) +b f );

[0048] In the formula, f t is the activation value of the forget gate, which is used to control which information in the cell state needs to be forgotten, σ is the Sigmoid activation function, Wf is the weight matrix of the forget gate, b f The bias term of the forget gate, H combined (t-1) is the output of the t-1 layer network and the input of the t-th layer LSTM network;

[0049] S333, input gate determines the current input H combined What information in will be written into the cell state? It includes two parts: the input gate activation i t and candidate memory cells The calculation formulas are as follows:

[0050] i t =σ(W i H combined(t-1 )+b i );

[0051] In the formula, i t is the activation value of the input gate, which controls what new information will be written into the cell state, W i is the weight of the input gate, b i is the bias term of the input gate, H combined(t-1) It is the output of the t-1 layer network and the input of the t-th layer LSTM network;

[0052]

[0053] In the formula, is a candidate memory unit, representing the new potential cell state information, W C is the weight matrix of the candidate memory unit, b C is the bias term of the candidate memory unit, H combined(t-1 ) is the output of the t-1 layer network and the input of the t-th layer LSTM network;

[0054] S334, update the cell state C according to the output of the forget gate and the input gate t , the calculation formula is as follows:

[0055]

[0056] In the formula, C t is the updated cell state, C t-1 is the cell state at the previous time step, f t is the activation value of the forget gate, i t is the activation value of the input gate, is a candidate memory unit;

[0057] S335, the final output gate calculation formula is as follows:

[0058] o t =σ(Wo H commbined(t-1) +b o );

[0059] o t The activation value of the output gate controls which cell state information will affect the hidden state, W o is the weight matrix of the output gate, b o is the bias term of the input gate, H commbined(t-1) It is the output of the t-1 layer network and the input of the t-th layer LSTM network;

[0060] The new hidden state calculation formula is as follows:

[0061] h t =o t tanh(C t );

[0062] In the formula, o t The activation value of the output gate, C t is the updated cell state, h t is the new hidden state, used as the input of the t+1 layer LSTM network;

[0063] S336, the present invention will o t As the gate value to dynamically adjust the weights of forward and backward attention, the output formula of the final improved multi-level bidirectional multi-head attention module is as follows:

[0064] H final =o t y·H forward +(1-o t )·H backward ;

[0065] In the formula, o t The activation value of the output gate, H forward is the output of the forward attention, H backward It is the output of backward attention, so that the network can better capture the important information in the smoke pollution data.

[0066] Furthermore, the step S4 proposes a smooth loss function, which specifically includes defining a smooth loss function, and the smooth loss function is as follows:

[0067]

[0068] In the formula, Loss is the output value of the total loss function, the target value to be minimized during the training process, M is the number of thresholds, N is the number of samples in the smoke pollutant dataset, is the predicted value of the nth sample predicted by the T3N model under the mth threshold number, R nis the actual value of the nth sample, ∈ is a constant, is the target threshold number T m The threshold loss function is , where u is the difference between the true sample value and the predicted sample value, and L is the regularization term. The formula is as follows:

[0069]

[0070] Where γ is the regularization coefficient, α is the mixed ratio parameter of L1 and L2 regularization (0≤α≤1), and β is the T3N model parameter. The weight of the T3N model is adjusted according to the set threshold value. The size of the threshold value not only determines the difference in the predicted value of the T3N model, but also affects the different ways of real-time regulation and optimization operation of artificial pollution control.

[0071] Furthermore, in step S5, the output of the improved multi-level bidirectional multi-head attention module is converted into the final flue gas pollutant generation amount prediction value after passing through the fully connected layer, which specifically includes: the fully connected layer performs a linear transformation on each input feature, adds a bias term, and then performs a nonlinear transformation through an activation function. The calculation formula of the fully connected layer is as follows:

[0072] y=σ(H final ·W final +b final );

[0073] In the formula, H final is the output value of the multi-level bidirectional multi-head attention module, W final is the weight parameter of the fully connected layer, b final is the bias vector of the fully connected layer, and σ is the ReLU activation function.

[0074] In summary, the present invention is a method for predicting the amount of flue gas pollutant generation, and the beneficial effects of the present invention are: the multi-level bidirectional multi-head attention module combined with residual connection and layer normalization can more comprehensively capture the dependencies in the data and improve the prediction accuracy of the model; the output gate of LSTM is used to dynamically adjust the attention weight, making the model more flexible and robust; the smooth loss function is combined with Elastic-Net regularization to effectively reduce the overfitting problem and improve the stability and generalization ability of the model, wherein the threshold setting of the smooth loss function plays an important role in the training of the model; the method has made innovations and improvements in data preprocessing, model design, loss function optimization, regularization technology and other aspects, which significantly improves the accuracy and stability of the prediction of the amount of flue gas pollutant generation, and has strong practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Figure 1 The present invention is a flow chart of the steps of a method for predicting the amount of flue gas pollutant generation.

[0076] Figure 2 This is a T3N model structure diagram for a flue gas pollutant generation prediction method.

[0077] Figure 3 This is a structural diagram of the LSTM adaptive gate mechanism in a flue gas pollutant generation prediction method.

[0078] Figure 4 This is a comparison chart between the predicted values ​​and actual values ​​of nitrogen dioxide and carbon monoxide concentrations of the T3N model in a flue gas pollutant generation prediction method. DETAILED DESCRIPTION

[0079] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0080] See also Figure 1-Figure 4 The present invention provides a technical solution: a method for predicting the amount of flue gas pollutant generation, which combines the improved multi-head attention with the bidirectional attention and dynamically adjusts the weights of the forward and backward attention through the proposed adaptive gating mechanism, so as to more comprehensively capture information, dynamically adjust the attention weights, alleviate the gradient disappearance and gradient explosion problems, improve the expression ability of the model, and enhance the generalization ability of the model. The specific steps of the present invention are as follows: Figure 1 shown.

[0081] The present invention provides a method for predicting the amount of flue gas pollutant generated, the method comprising the following steps:

[0082] S1. Obtain the data information required for the prediction method of flue gas pollutant generation to form a flue gas pollutant generation data set, including: the flue gas pollutant generation related data contains four characteristics, namely: fuel characteristics, combustion process parameters, environmental conditions and operating conditions, where the fuel characteristics include fuel type, sulfur content, nitrogen content, ash content, volatile matter content information, different types of fuel (such as coal, oil, natural gas, etc.) will produce different pollutants, combustion process parameters include combustion temperature (the higher the temperature, the more nitrogen oxides are generated), excess air coefficient (excess air will affect the completeness of combustion and the amount of monoxide The features include carbon generation), combustion mode (different combustion technologies will affect the amount of pollutants generated) and combustion time (the length of combustion time will affect the degree of complete combustion of the fuel), environmental conditions include temperature (ambient temperature affects combustion efficiency and the diffusion of pollutants), pressure (ambient pressure may affect the chemical reaction rate during combustion), operating conditions include fuel mixing ratio and pollution control equipment (performance and operating status of electrostatic precipitator, desulfurization equipment, etc.). The above features are currently sent to the prediction model as attributes during the training process; in addition, it also includes flue gas pollutants PM2.5, PM10, sulfur dioxide (SO 2 ), nitrogen dioxide (NO 2 ), carbon monoxide (CO) and volatile organic compounds (VOCs) data values, and during the training process, the harmful gases in the flue gas pollutants are used as the predicted values ​​of the prediction model.

[0083] S2. Preprocessing the current data set, including filling missing values ​​and constructing training sets and test sets according to the filling results, specifically including: during the data collection process, the smoke monitoring equipment may be unable to record data due to technical failure, maintenance or damage. In order to improve the data quality, the present invention uses linear interpolation to process the missing values. Given x 0 ,y 0 and x 1 ,y 1 , in x 0 and x 1 The difference y at a point x between is calculated by the following formula:

[0084]

[0085] In the formula, (x 0 ,y 0 ) and (x 1 ,y 1 ) is a known flue gas pollutant data point, x is the point to be interpolated, and x satisfies 0 ≤x≤x 1 , y is the data interpolation of the flue gas pollutants at x;

[0086] The data set after the current preprocessing is segmented and divided into a training set and a test set according to a ratio of 7:3 of the total number of flue gas pollutant data.

[0087] S3. A Threshold Number Neural Network model is proposed, which is referred to as T3N. The hidden layer of the model consists of a multi-level bidirectional multi-head attention module, and the module is improved by introducing residual connections, using layer normalization, and enhancing the adaptive gating mechanism. The module can pay attention to the previous and next information of the input sequence at the same time, so as to capture the dependencies in the data more comprehensively.

[0088] Furthermore, in step S3, a threshold number neural network (T3N) is constructed, and the training set is input into the T3N model. The specific steps are as follows:

[0089] S31. Build T3N model, model structure reference Figure 2 , the input layer receives preprocessed feature data, such as fuel characteristics, combustion process parameters, environmental conditions and operating conditions; the hidden layer uses multiple bidirectional attention modules, and each layer uses the ReLU activation function; the output layer is a plurality of neurons, including a plurality of output nodes, each output node corresponds to a predicted value of a target threshold number;

[0090] S32, the multi-level bidirectional multi-head attention module is composed of improved multi-head attention and bidirectional attention modules. First, the data output by the input layer is linearly transformed. The formula is as follows:

[0091] Q=XW Q ;

[0092] K=XW K ;

[0093] V=XW V ;

[0094] Wherein, X is the smoke pollutant data after the pretreatment, W Q , W K , W V is the weight matrix, Q, K, V are the values ​​after linear transformation;

[0095] The attention score calculation formula is as follows:

[0096]

[0097] In the formula, Q, K, and V are the values ​​after linear transformation. is the scaling factor;

[0098] Multi-head attention is achieved by computing h independent attention heads in parallel. The multi-head attention formula is as follows:

[0099] H i =Attention(Q i , K i , V i ), i=1, 2, 3..., h;

[0100] In the formula, h is the number of attention heads, H i is the score of the i-th attention head, Q i is the Q matrix of the i-th attention head, K i is the K matrix of the i-th attention head, Vi 为 V matrix of the i-th attention head;

[0101] The multi-head attention is concatenated and the final output is obtained through linear transformation:

[0102] MultiHead(Q,K,V)=Concat(H 1 H 2 ......, H h )·W O ;

[0103] Where MultiHead(Q, K, V) is the total output of the multi-head attention. For the bidirectional attention module, the positive multi-head attention is also called the forward attention, that is, H forward =MultiHead(Q, K, V), W O is the weight matrix; residual connection and layer normalization operations are added, and forward and backward attention are calculated for each layer respectively. The forward attention formula of the lth layer is as follows:

[0104] In the formula, is the forward attention weight parameter of the l-1th layer, is the l-th forward multi-head attention, and LayerNorm is the normalization operation on the layer weight parameters;

[0105] The formula for backward attention at layer l is as follows:

[0106]

[0107] in:

[0108] X backward =Reverse(X);

[0109] Q backward =X backward W Q ;

[0110] K backward =X backward W K ;

[0111] V backward =X backward W V ;

[0112] In the formula, Reverse is the reverse operation. is the backward attention weight parameter of the l-1th layer;

[0113] S33, the adaptive gating mechanism is used to dynamically adjust the weights of forward and backward attention, so as to flexibly capture important information at different time steps and features. In order to enhance this mechanism, the present invention adopts a complex long short-term memory network (LSTM) to calculate the gating value, which can more effectively process sequence data and introduce more context information in the gating process. The specific steps are as follows:

[0114] S331, the output H of the improved forward attention forward And the output H of the backward attention backward Concatenate as the subsequent input of LSTM:

[0115] H combined =Concat(H forward , H backward );

[0116] H combined(t-1) =H combined ·h t-1 ;

[0117] In the formula, H combined(t-1) It is the output of the t-1 layer network and the input of the t-th layer LSTM network;

[0118] LSTM introduces three gate mechanisms (input gate, forget gate and output gate), refer to Figure 3 As shown, a cell state is used to maintain and adjust the hidden state and cell state, thereby effectively capturing long-term dependencies;

[0119] S332, forget gate determines the previous cell state C t-1 Which information will be discarded? The calculation formula is as follows:

[0120] f t =σ(W f H combined(t-1) +b f );

[0121] In the formula, f tis the activation value of the forget gate, which is used to control which information in the cell state needs to be forgotten, σ is the Sigmoid activation function, W f is the weight matrix of the forget gate, b f The bias term of the forget gate, H combined(t-1) It is the output of the t-1 layer network and the input of the t-th layer LSTM network;

[0122] S333, input gate determines the current input H combined What information in will be written into the cell state? It includes two parts: the input gate activation i t and candidate memory cells The calculation formulas are as follows:

[0123] i t =σ(W i H combined(t-1) +b i );

[0124] In the formula, i t is the activation value of the input gate, which controls what new information will be written into the cell state, W i is the weight of the input gate, b i is the bias term of the input gate, H combined(t-1) It is the output of the t-1 layer network and the input of the t-th layer LSTM network;

[0125]

[0126] In the formula, is a candidate memory unit, representing the new potential cell state information, W C is the weight matrix of the candidate memory unit, b C is the bias term of the candidate memory unit, H combined(t-1) It is the output of the t-1 layer network and the input of the t-th layer LSTM network;

[0127] S334, update the cell state C according to the output of the forget gate and the input gate t , the calculation formula is as follows:

[0128]

[0129] In the formula, C t is the updated cell state, C t-1 is the cell state at the previous time step, f t is the activation value of the forget gate, i t is the activation value of the input gate, is a candidate memory unit;

[0130] S335, the final output gate calculation formula is as follows:

[0131] o t =σ(W o H commbine d( t-1 )+b o );

[0132] o t The activation value of the output gate controls which cell state information will affect the hidden state, W o is the weight matrix of the output gate, b o is the bias term of the input gate, H commbined(t-1) It is the output of the t-1 layer network and the input of the t-th layer LSTM network;

[0133] The new hidden state calculation formula is as follows:

[0134] h t =o t tanh(C t );

[0135] In the formula, o t The activation value of the output gate, C t is the updated cell state, h t is the new hidden state, used as the input of the t+1 layer LSTM network;

[0136] S336, the present invention will o t As the gate value to dynamically adjust the weights of forward and backward attention, the output formula of the final improved multi-level bidirectional multi-head attention module is as follows:

[0137] H final =o t ·H forward +(1-o t )·H backward ;

[0138] In the formula, o t The activation value of the output gate, H forward is the output of the forward attention, H backward It is the output of backward attention, so that the network can better capture the important information in the smoke pollution data.

[0139] S4. A smooth loss function is proposed. The loss function is related to the set threshold. The weight parameters of the T3N model are adjusted according to the size of the predicted value and the threshold. The Elastic-Net regularization term is used to reduce the problem of model overfitting. The Elastic-Net regularization term combines the regularization techniques of L1 and L2 regularization, aiming to have both the feature selection ability of L1 regularization and the stability of L2 regularization. Specifically, a smooth loss function is defined. The smooth loss function is as follows:

[0140]

[0141] In the formula, Loss is the output value of the total loss function, the target value to be minimized during the training process, M is the number of thresholds, N is the number of samples in the smoke pollutant dataset, is the predicted value of the nth sample predicted by the T3N model under the mth threshold number, R n is the actual value of the nth sample, ∈ is a constant, is the target threshold number T m The threshold loss function is , where u is the difference between the true sample value and the predicted sample value, and L is the regularization term. The formula is as follows:

[0142]

[0143] Where γ is the regularization coefficient, α is the mixed ratio parameter of L1 and L2 regularization (0≤α≤1), and β is the T3N model parameter. The weight of the T3N model is adjusted according to the set threshold value. The size of the threshold value not only determines the difference in the predicted value of the T3N model, but also affects the different ways of real-time regulation and optimization operation of artificial pollution control.

[0144] S5. The output of the improved multi-level bidirectional multi-head attention module is passed through the fully connected layer to obtain the final predicted value of the amount of smoke pollutant generated. Specifically, the fully connected layer performs a linear transformation on each input feature, adds a bias term, and then performs a nonlinear transformation through the activation function. The calculation formula of the fully connected layer is as follows:

[0145] y=σ(H final ·W final +b final );

[0146] In the formula, H final is the output value of the multi-level bidirectional multi-head attention module, W final is the weight parameter of the fully connected layer, b final is the bias vector of the fully connected layer, and σ is the ReLU activation function.

[0147] Furthermore, the network architecture of the T3N model is designed as follows: the input layer processes the preprocessed data, the input shape sequence length is 100, the feature dimension is 64, the hidden layer uses a multi-level bidirectional multi-head attention module, which contains 3 layers in total, each layer uses 4 attention heads, and the dimension of each head is 64. Residual connection is introduced after each layer of multi-head attention module to help alleviate the gradient vanishing problem and accelerate model convergence. At the same time, layer normalization is applied to stabilize the training process and improve model performance; the enhanced adaptive gating mechanism calculates the gating value by using the output gate of LSTM, and the dynamic The weights of forward and backward attention are adjusted dynamically; the output after the multi-head attention module is further processed by a fully connected layer. Two fully connected layers are used. The first layer contains 128 neurons and the second layer contains 64 neurons. Both use the ReLU activation function. The last fully connected layer outputs 20 neurons as the final prediction value. To prevent overfitting, Dropout (0.2) is used in the fully connected layer, and the Elastic-Net regularization term (L1 = 0.01, L2 = 0.01) is adopted, combining the advantages of L1 and L2 regularization.

[0148] Furthermore, the smooth loss function introduces a threshold value to adjust the weight of the T3N model. The value of the threshold value not only determines the difference in the predicted value of the T3N model, but also affects the different ways of real-time regulation and optimization of artificial pollutant control. The threshold value selected by the present invention is 0.6. The experimental results show that when the threshold value is 0.6, the flue gas pollutant generation and fuel cost are the most cost-effective after real-time regulation and optimization of artificial pollutant control. In the smooth loss function, It is a smooth Pinball loss function that can effectively handle large errors and outliers. Compared with the traditional mean square error (MSE), the Pinball loss function performs better in dealing with heteroskedasticity and asymmetric distribution because it is less sensitive to outliers, which makes the model more robust in the face of abnormal data.

[0149] Further, refer to Figure 4 , which is the experimental result diagram of a flue gas pollutant generation prediction method. Figure 4 Only part of the experimental results are shown, including the prediction of nitrogen dioxide concentration and carbon monoxide concentration. It can be seen from the figure that the T3N model has a good prediction effect, and the difference between its predicted value and the true value is small, indicating that the flue gas pollutant generation prediction method proposed in the present invention is effective.

Claims

1. A method for predicting the amount of flue gas pollutants generated, characterized in that: The following steps are involved: S1. Collect data related to the amount of smoke pollutants generated, including relevant attribute values ​​and the amount of harmful gas generated by the target variable; S2, fill the missing values ​​of the current data set by linear interpolation, and divide the filled data set into a training set and a test set; S3. A Threshold Number Neural Network (T3N) model is proposed. The hidden layer of the model adopts an improved multi-level bidirectional multi-head attention module. In this model, residual connections are introduced to alleviate the gradient vanishing problem and accelerate model convergence. Layer normalization is used to stabilize the training process and improve model performance. The adaptive gating mechanism is enhanced to dynamically adjust the weights of forward and backward attention. Through these improvements, the model can pay attention to the front and back information of the input sequence at the same time, thereby capturing the dependencies in the data more comprehensively. S31, use the improved multi-level bidirectional multi-head attention module for feature extraction; S32. Use a long short-term memory network (LSTM) to calculate the gating value for dynamically adjusting the weights of forward and backward attention, thereby flexibly capturing important information at different time steps and features, and introducing more contextual information in the gating process; S4. A smooth loss function is proposed. The loss function is related to the set threshold. The weight parameters of the T3N model are adjusted according to the difference between the predicted value and the threshold. At the same time, in order to reduce the overfitting problem of the model, the Elastic-Net regularization term is applied. This regularization term combines L1 and L2 regularization techniques, aiming to have both the feature selection ability of L1 regularization and the stability of L2 regularization. S5, the output of the improved multi-level bidirectional multi-head attention module in the T3N model is processed through a fully connected layer to finally generate a predicted value of the smoke pollutant; Among them, in step S3, the improved multi-level bidirectional multi-head attention module described in S31 is used to extract the time series features in the smoke pollutant data, and fully mine the important information in the data; the multi-level bidirectional multi-head attention module is composed of an improved multi-head attention and a bidirectional attention module. First, the data output by the input layer is linearly transformed, and the formula is as follows: ; ; ; In the formula, is the flue gas pollutant data after pretreatment. , , is the weight matrix, , , is the value after linear transformation; The attention score calculation formula is as follows: ; In the formula, , , is the value after linear transformation, is the scaling factor; Multi-head attention through parallel computing There are independent attention heads, and the multi-head attention formula is as follows: ; In the formula, is the number of attention heads, For the The score of head attention, For the Head attention matrix, For the Head attention matrix, For the Head attention matrix; The multi-head attention is concatenated and the final output is obtained through linear transformation: ; In the formula, is the total output of the multi-head attention. For the bidirectional attention module, the positive multi-head attention is also called the forward attention, that is, , is the weight matrix; Add residual connections and layer normalization operations, calculate forward and backward attention in each layer, The layer forward attention formula is as follows: ; In the formula, For the The forward attention weight parameter of the layer, For the The forward multi-head attention To normalize the layer weight parameters; No. The layer backward attention formula is as follows: ; in: ; ; ; ; In the formula, For reverse operation, For the The backward attention weight parameters of the layer.

2. A method for predicting the amount of flue gas pollutants generated according to claim 1, characterized in that: In step S2, the missing values ​​are filled using linear interpolation on the data set of step S1 in claim 1. , and , ,exist and Some point between The difference Calculated by the following formula: ; In the formula, and is a known flue gas pollutant data point, is the point to be interpolated and satisfies , For Data interpolation of flue gas pollutants at The current preprocessed data set is segmented and divided into a training set and a test set in a ratio of 7:3 of the total number of flue gas pollutant data.

3. A method for predicting the amount of flue gas pollutants generated according to claim 1, characterized in that: In step S3, S32 uses a long short-term memory network (LSTM) to calculate the gate value for dynamically adjusting the weights of forward and backward attention, including: S331, the output of the improved forward attention And the output of backward attention Concatenate as the subsequent input of LSTM: ; ; In the formula, for The output of the layer network is also the The input of the layer LSTM network; S332, forget gate determines the previous cell state Which information will be discarded? The calculation formula is as follows: ; In the formula, is the activation value of the forget gate, which is used to control which information in the cell state needs to be forgotten. is the Sigmoid activation function, is the weight matrix of the forget gate, The bias term of the forget gate; S333, input gate determines the current input What information will be written into the cell state, including two parts: input gate activation and candidate memory cells , the calculation formulas are as follows: ; In the formula, is the activation value of the input gate, which controls what new information will be written into the cell state, is the weight of the input gate, is the bias term of the input gate; ; In the formula, is a candidate memory unit, representing new potential cell state information, is the weight matrix of the candidate memory unit, is the bias term of the candidate memory unit; S334, update the cell state according to the output of the forget gate and the input gate , the calculation formula is as follows: ; In the formula, is the updated cell state, is the cell state at the previous time step, is the activation value of the forget gate, is the activation value of the input gate, is a candidate memory unit; S335, the final output gate calculation formula is as follows: ; The activation value of the output gate controls which cell state information will affect the hidden state. is the weight matrix of the output gate, is the bias term of the input gate; The new hidden state calculation formula is as follows: ; In the formula, The activation value of the output gate, is the updated cell state, is the new hidden state for Input of the LSTM network S336, the present invention will As the gate value to dynamically adjust the weights of forward and backward attention, the output formula of the final improved multi-level bidirectional multi-head attention module is as follows: ; In the formula, The activation value of the output gate, is the output of the forward attention, It is the output of backward attention, so that the network can better capture the important information in the smoke pollution data.

4. A method for predicting the amount of flue gas pollutants generated according to claim 1, characterized in that: In step S4, a smooth loss function is proposed, and the smooth loss function is as follows: ; ; ; In the formula, Output value of the total loss function, the target value to be minimized during training, is the number of threshold values, is the number of samples in the flue gas pollution data set, The predicted value of the nth sample predicted by the T3N model under the mth threshold number, For the The actual value of the samples, is a constant, The target threshold number The threshold loss function is is the difference between the true sample value and the predicted sample value, is the regularization term, and the formula is as follows: ; In the formula, is the regularization coefficient, is the mixing ratio parameter of L1 and L2 regularization (0 ≤ α ≤ 1), is the T3N model parameter; the weight of the T3N model is adjusted according to the set threshold value. The size of the threshold value not only determines the difference in the T3N model prediction value, but also affects the different ways of real-time regulation and optimization operation of artificial pollution control.

5. A method for predicting the amount of flue gas pollutants generated according to claim 1, characterized in that: In step S5, the output of the improved multi-level bidirectional multi-head attention module is passed through the fully connected layer to obtain the final predicted value of the amount of smoke pollutant generated. The fully connected layer performs a linear transformation on each input feature, adds a bias term, and then performs a nonlinear transformation through the activation function. The calculation formula of the fully connected layer is as follows: ; In the formula, is the output value of the multi-level bidirectional multi-head attention module, is the weight parameter of the fully connected layer, is the bias vector of the fully connected layer, is the ReLU activation function.

Citation Information

Patent Citations

  • Air pollutant concentration forecasting method and device and storage medium

    CN112766549A

  • Three-dimensional convolution attention neural network modeling method for nitric oxide in cement denitration process

    CN117217266A

  • Flue gas pollution source collecting and monitoring system and method and readable storage medium

    CN117744704A

  • Crop growth trend prediction method and system based on growth model

    CN117972329A