Multi-strategy optimization algorithm and mechanism feature fusion-based ammonia desulfurization outlet SO2 concentration intelligent prediction method

By integrating multi-strategy optimization algorithms and mechanistic characteristics into the ammonia-based desulfurization system of a coal-fired power plant, a Transformer neural network model was constructed. This solved the problems of measurement response lag and insufficient model robustness in existing technologies, achieving high-precision prediction of SO2 concentration, adapting to changes in coal type and load, and providing a reliable prediction basis.

CN122050573APending Publication Date: 2026-05-15WUHAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN UNIV OF TECH
Filing Date
2026-01-21
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies in ammonia desulfurization systems for coal-fired power plants suffer from problems such as measurement response lag, high equipment maintenance costs, low measurement accuracy, and insufficient predictive ability of models under complex operating conditions. In particular, the prediction error is large when the coal type changes or the load fluctuates, and the models lack mechanistic feature embedding and have insufficient robustness.

Method used

A multi-strategy optimization algorithm and mechanism feature fusion method is adopted. By collecting multi-source operating data, intermediate mechanism features such as effective ammonia concentration and SO2 mass transfer-reaction comprehensive factor are calculated to construct a Transformer neural network model. The prediction accuracy and robustness of the model are improved by using a multi-strategy improved particle swarm optimization algorithm and adversarial training.

Benefits of technology

It achieves high-precision, minute-level advance prediction of SO2 concentration at the outlet of ammonia desulfurization, improves the model's generalization ability, interpretability, and robustness, and provides a reliable basis for operation optimization and emission compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122050573A_ABST
    Figure CN122050573A_ABST
Patent Text Reader

Abstract

The invention discloses an ammonia desulfurization outlet SO2 concentration intelligent prediction method based on a multi-strategy optimization algorithm and mechanism feature fusion. The method comprises the steps that multi-source operation data of an ammonia desulfurization system is collected and preprocessed to construct a training sample set; calculating intermediate mechanism characteristics; constructing a Transform neural network model, and performing hyper-parameter optimization by adopting a multi-strategy improved particle swarm optimization algorithm; the optimized Transform neural network model is trained through the training sample set, and the trained Transform neural network model is used for actual SO2 concentration prediction. According to the method, multi-source operation data is combined, key mechanism characteristics are explicitly embedded into a depth prediction model, and model prediction precision and robustness are improved through multi-strategy particle swarm optimization and adversarial training, so that minute-level advanced prediction of the concentration of SO2 at an outlet is realized, and a reliable basis is provided for operation optimization and emission standard reaching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of SO2 concentration prediction. Specifically, this invention relates to an intelligent prediction method for SO2 concentration at the outlet of ammonia desulfurization based on the fusion of multi-strategy optimization algorithm and mechanism characteristics. Background Technology

[0002] Ammonia-based desulfurization technology for coal-fired power units is gradually being adopted in some coal-fired power plants due to its advantages such as high value of desulfurization byproducts and no secondary pollution. To meet increasingly stringent national and local ultra-low emission requirements, desulfurization systems need to balance desulfurization efficiency with desulfurizing agent consumption and operating costs. Therefore, accurate and proactive prediction of SO2 concentration at the desulfurization tower outlet is of significant engineering importance.

[0003] Currently, power plants generally use continuous emission monitoring systems (CEMS) to monitor SO2 concentration at the desulfurization tower outlet online. However, CEMS has the following problems: First, the measurement response has a certain lag, making it difficult to reflect the rapid changes in actual operating conditions in a timely manner; second, the equipment maintenance cost is high, and the measurement accuracy in low concentration ranges is easily reduced due to probe contamination and calibration cycles; third, it only provides the result value and lacks the ability to predict the trend of SO2 emission changes under complex operating conditions.

[0004] To address the aforementioned issues, existing technologies have proposed using traditional neural networks, support vector machines, ensemble learning, and other methods to perform data-driven modeling of desulfurization efficiency or outlet SO2 concentration, incorporating operating parameters such as liquid-to-gas ratio, pH, and load as input features. However, these methods generally suffer from the following shortcomings:

[0005] Feature selection focuses primarily on the operating parameters of the desulfurization tower itself, with less consideration given to the overall impact from coal quality and combustion to desulfurization. This results in limited extrapolation capabilities of the model, leading to a significant increase in prediction errors when coal type changes or load fluctuates dramatically.

[0006] Most of the models are "black box" structures, and key mechanisms such as alkali consumption of absorbent liquid, gas-liquid mass transfer and neutralization reaction have not been explicitly embedded into the model input, resulting in insufficient interpretability and engineering credibility of the models.

[0007] Model hyperparameters are often determined by experience or simple search methods, which can easily get trapped in local optima and make it difficult to balance prediction accuracy, training efficiency and stability.

[0008] The model does not adequately consider the robustness of unsteady conditions such as load fluctuations and coal quality fluctuations, and is prone to large deviations under sudden changes in conditions.

[0009] Publication No. CN110471291A, published on 2019-11-19, discloses a disturbance suppression prediction and control method for an ammonia-based desulfurization system. Although disturbance suppression is considered, its model is purely data-driven and does not embed the pollutant reaction mechanism into the model. Publication No. CN102693451A, published on 2012-09-26, discloses a multi-parameter ammonia-based flue gas desulfurization efficiency prediction method. It uses multiple artificial intelligence models to predict desulfurization efficiency, but it only targets SO2 as a single pollutant and does not consider the influence of coal quality characteristics.

[0010] Therefore, there is an urgent need for an intelligent prediction method for SO2 concentration at the outlet of ammonia desulfurization, which integrates mechanistic features and deep learning structures based on multi-source data throughout the entire process, and improves robustness through multi-strategy optimization and adversarial training, so as to achieve high-precision and advanced prediction of actual emission concentration.

[0011] Therefore, this invention proposes an intelligent prediction method for SO2 concentration at the outlet of ammonia-based desulfurization based on the fusion of multi-strategy optimization algorithm and mechanism characteristics. Summary of the Invention

[0012] This invention aims to overcome the shortcomings of existing technologies and proposes an intelligent prediction method for SO2 concentration at the outlet of ammonia desulfurization based on the fusion of multi-strategy optimization algorithms and mechanistic features, in order to achieve the following objectives: improve the prediction accuracy and robustness of the model, thereby realizing intelligent prediction of the outlet SO2 concentration.

[0013] To achieve the above objectives, the technical solution adopted by this invention is: an intelligent prediction method for SO2 concentration at the outlet of ammonia-based desulfurization based on the fusion of multi-strategy optimization algorithms and mechanistic features, the method comprising:

[0014] Step S1: Collect multi-source operating data of the ammonia desulfurization system and preprocess it to construct a training sample set. The multi-source operating data includes coal quality characteristics, combustion process characteristics, desulfurization process characteristics, and measured SO2 concentration at the desulfurization tower outlet.

[0015] Step S2: Calculate intermediate mechanism characteristics based on the desulfurization process characteristics, including: effective ammonia concentration and SO2 mass transfer-reaction comprehensive factor;

[0016] Step S3: Using the pre-treated coal quality characteristics, combustion process characteristics, desulfurization process characteristics, and intermediate mechanism characteristics as input samples, and the measured value of outlet SO2 concentration as output, construct a Transformer neural network model and use a multi-strategy improved particle swarm optimization algorithm to optimize its hyperparameters.

[0017] Step S4: Train the optimized Transformer neural network model using the training sample set, and use the trained Transformer neural network model for actual SO2 concentration prediction.

[0018] Preferably, the preprocessing in step S1 includes:

[0019] The multi-source operational data is time-aligned according to a unified sampling period, and linear interpolation is used to obtain the equivalent value of the time point for data whose sampling time deviates from the unified time point.

[0020] For any running data, if the number of consecutive missing points is less than a preset threshold, linear interpolation of the valid data before and after is used to fill the gap; if the number of consecutive missing points is greater than the preset threshold, the sample of the corresponding time slice is removed.

[0021] All multi-source operational data were normalized using a standardization method with zero mean and unit variance.

[0022] By constructing a sample sequence through a sliding time window, the multi-source operational data of the current moment and several moments before it are combined into an input sequence, and the measured value of the outlet SO2 concentration at the corresponding future preset time shift is used as a label, thus obtaining a labeled training sample set.

[0023] Preferably, in step S2, the effective ammonia concentration characteristic is used to quantify the impact of chloride ions in the slurry on the competitive consumption of the desulfurizing agent, and its calculation formula is as follows:

[0024] C NH3_effective =C NH3 ×(1-α×exp(-β×(N / Cl)));

[0025] Among them, C NH3 The N / Cl ratio represents the ammonia nitrogen concentration in the slurry; N / Cl represents the molar ratio of ammonia nitrogen to chloride ions in the slurry; α and β are positive coefficients obtained by regression from historical operating data, used to quantitatively characterize the competitive consumption of available ammonia by chloride ions.

[0026] Preferably, in step S2, the SO2 mass transfer-reaction comprehensive factor is a characteristic quantity calculated by integrating the following parameters:

[0027] ① pH value of slurry;

[0028] ② Slurry temperature;

[0029] ③ Empty gas velocity of the desulfurization tower;

[0030] ④ Liquid-to-gas ratio in the desulfurization system;

[0031] The SO2 mass transfer-reaction comprehensive factor is used to characterize the overall removal rate of SO2 from the gas phase to the liquid phase and undergoing neutralization reaction in the desulfurization tower.

[0032] Preferably, the SO2 mass transfer-reaction comprehensive factor is expressed as follows:

[0033] k mt = k La ·f(pH);

[0034] Where, k La Let be the overall gas-liquid volumetric mass transfer coefficient per unit volume of slurry under empty tower conditions, and its expression is as follows:

[0035] k La = k0 (L / G) m U g n exp [-E a / (T +273.15) / R1];

[0036] In the formula, k0 is an empirical coefficient; L / G is the liquid-to-gas ratio of the desulfurization system; Ug is the empty gas velocity of the desulfurization tower; T is the slurry temperature; Ea is the apparent activation energy; R1 is the gas constant; m and n are exponential parameters obtained by regression using historical operating data.

[0037] f(pH) is a correction function characterizing the pH dependence of the acid-base neutralization reaction rate, and its expression is as follows:

[0038] f(pH) = 1 + γ·(pH - pH0);

[0039] In the formula, γ is the fitting coefficient and pH0 is the reference pH value.

[0040] Preferably, in step S3, constructing the Transformer neural network model includes:

[0041] For each time step t, coal quality characteristics, combustion characteristics, desulfurization process characteristics, and mechanism characteristics are concatenated into a one-dimensional feature vector z_t ∈ R^F. Time windows of length L {z_{t-L+1},…,z_t} are stacked to form an input matrix Z ∈ R^{L×F}.

[0042] First, the features at each time step are mapped to the d_model dimensional embedding space using a linear mapping: H (0) = ZW in +b in ;

[0043] Among them W in ∈ R^{F×d_model},b inAs a trainable bias; then add a positional encoding P∈ R^{L×d_model} for each time step, resulting in:

[0044] ;

[0045] Then, N layers of encoders are stacked, each layer containing a multi-head self-attention sublayer and a feedforward fully connected sublayer, and residual connections and LayerNorm structures are used between the sublayers to achieve deep modeling of sequence features;

[0046] Finally, the hidden state h_L^{(N)} at the last time step is taken as the sequence representation and input into the regression output layer to obtain the predicted value of the outlet SO2 concentration.

[0047] Preferably, in step S3, the hyperparameter optimization using the multi-strategy improved particle swarm optimization algorithm includes: improving the particle swarm optimization algorithm using a dynamic inertia weight strategy, namely:

[0048] Let the particle swarm size be Np, the maximum number of iterations be Kmax, and the position vector xi of each particle i represent a set of hyperparameters to be optimized, with the corresponding velocity vector vi. In the k-th iteration, the velocity and position of particle i in the d-th dimension are updated as follows:

[0049] v i,d (k+1) = ω (k) v i,d (k) + c1 (k) r1(p i,d best - x i,d (k) ) + c2 (k) r2(g d best - x i,d (k) );

[0050] x i,d (k+1) = x i,d (k) + v i,d (k+1) ;

[0051] Where, p i,d best g represents the historical best position of particle i; d best The global optimal position is represented by r1 and r2, which are uniformly random numbers in the interval [0,1]. c1 is the cognitive factor, c2 is the social factor, and the inertia weight ω is the inertia weight. (k) Employ a dynamic decreasing strategy:

[0052] ω (k) = ω max - (ω max - ω min )·k / K max ;

[0053] Where ω max ω represents the preset maximum inertia weight. min This represents the preset minimum inertia weight;

[0054] The mean square error of SO2 prediction at the outlet is taken as the fitness function, and the smaller the error, the higher the fitness. After setting the number of iterations or meeting the convergence condition, the combination of hyperparameters of the globally optimal particle is taken as the optimal hyperparameters of the Transformer model.

[0055] Preferably, in step S3, the hyperparameter optimization using the multi-strategy improved particle swarm optimization algorithm further includes: introducing a random perturbation based on the Lévy flight mechanism during the particle position update process to enhance the ability to escape local optima, i.e.:

[0056] After each particle position update, a Levy flight perturbation is applied to the positions of some particles with a preset probability using PLF:

[0057] x i (k+1) | Levy= x i (k+1) + δ Levy(λ);

[0058] Where δ represents the step size coefficient; Levy(λ) represents element-wise multiplication; it is a random vector that follows a power-law distribution; x i (k+1) Indicates the updated particle position; x i (k+1) | Levy represents the updated particle position after introducing the Levy flight perturbation; λ is the stability index of the Levy distribution. Preferably, in step S3, the hyperparameter optimization using the multi-strategy improved particle swarm optimization algorithm further includes: adaptively adjusting the cognitive factor c1 and the social factor c2 according to the population convergence state, that is:

[0059] c1 (k) = c 1,max - (c 1,max - c 1,min )·k / K max ;

[0060] c2 (k) = c2,min + (c 2,max - c 2,min )·k / K max ;

[0061] Among them, c 1,max This represents the preset maximum cognitive factor c1 value, c 1,min This represents the preset minimum cognitive factor c1 value; c 2,max c represents the preset maximum cognitive factor c2 value. 2,min This represents the preset minimum cognitive factor c2 value.

[0062] Preferably, in step S4, an adversarial training strategy is used for model training, including:

[0063] Adversarial examples are constructed using the Fast Gradient Signed Method (FGSM): For each batch of input samples Z and their labels y, the gradient of the loss function L(Z, y) with respect to the input Z is calculated. Construct adversarial examples:

[0064] Z adv = Z + ε·sign( );

[0065] Where ε is the disturbance amplitude; then Z and Z adv The samples are merged as training input, using a loss L(Z,y) that includes the original sample loss and an adversarial sample loss L(Z). adv The parameters are updated using the weighted total loss of (y).

[0066] The technical effects of this invention are as follows: This invention realizes unified modeling of multi-source data from coal quality, combustion to desulfurization, explicitly embeds key mechanism features such as effective ammonia concentration and SO2 mass transfer-reaction comprehensive factor into the deep prediction model, and improves the model prediction accuracy and robustness through multi-strategy particle swarm optimization and adversarial training, thereby achieving minute-level advance prediction of outlet SO2 concentration, providing a reliable basis for operation optimization and emission compliance. Attached Figure Description

[0067] Figure 1 The flowchart of the intelligent prediction method for SO2 concentration at the outlet of ammonia desulfurization based on the fusion of multi-strategy optimization algorithm and mechanism features provided by the present invention is shown. Detailed Implementation

[0068] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. The purpose is to help those skilled in the art to have a more complete, accurate, and in-depth understanding of the inventive concept and technical solutions of the present invention, and to facilitate its implementation. It should be noted that the terms "first," "second," etc., used in this application are only for the convenience of describing the technical solutions and to distinguish components; the corresponding component configurations may be the same or different, and are not intended to limit the scope of this application.

[0069] This invention provides an intelligent prediction method for SO2 concentration at the outlet of ammonia-based desulfurization based on multi-strategy optimization algorithms and mechanistic feature fusion. It aims to overcome the shortcomings of existing ammonia-based desulfurization outlet SO2 concentration prediction technologies, such as insufficient feature coverage, inadequate utilization of mechanistic information, insufficient model parameter optimization, and poor robustness under unsteady conditions. This method achieves unified modeling of multi-source data from coal quality, combustion, to desulfurization, explicitly embedding key mechanistic features such as effective ammonia concentration and SO2 mass transfer-reaction comprehensive factors into a deep prediction model. Furthermore, it improves the model's prediction accuracy and robustness through multi-strategy particle swarm optimization and adversarial training, thereby enabling minute-level advance prediction of outlet SO2 concentration and providing a reliable basis for operational optimization and emission compliance.

[0070] like Figure 1 As shown, the method of the present invention includes the following steps:

[0071] Step S1: Collect multi-source operating data of the ammonia desulfurization system and preprocess it to construct a training sample set. The multi-source operating data includes coal quality characteristics, combustion process characteristics, desulfurization process characteristics, and measured SO2 concentration at the desulfurization tower outlet.

[0072] Step S2: Calculate intermediate mechanism characteristics based on the desulfurization process characteristics, including: effective ammonia concentration and SO2 mass transfer-reaction comprehensive factor;

[0073] Step S3: Using the pre-treated coal quality characteristics, combustion process characteristics, desulfurization process characteristics, and intermediate mechanism characteristics as input samples, and the measured value of outlet SO2 concentration as output, construct a Transformer neural network model and use a multi-strategy improved particle swarm optimization algorithm to optimize its hyperparameters.

[0074] Step S4: Train the optimized Transformer neural network model using the training sample set, and use the trained Transformer neural network model for actual SO2 concentration prediction.

[0075] Specifically, to make the technical solution of the present invention clearer, the present invention will be explained in detail through the following embodiments.

[0076] Referring to step S1, this embodiment selects one year of operating data from the ammonia desulfurization system of a 660 MW supercritical coal-fired unit, with a collection period of 1 minute. The collected multi-source operating data includes:

[0077] Coal quality characteristics: sulfur content (mass fraction) and chlorine content (mass fraction) of coal fed into the furnace.

[0078] Combustion characteristics: boiler load, total air volume, primary and secondary air ratio, oxygen content at furnace outlet, combustion temperature, etc.

[0079] Characteristics of the desulfurization process: flue gas temperature at the inlet of the desulfurization tower, pH of the circulating slurry, slurry temperature, liquid-to-gas ratio, ammonia flow rate, ammonia concentration, chloride ion concentration in the slurry, etc.

[0080] Measured SO2 concentration at the outlet of the desulfurization tower.

[0081] The multi-source operational data is then preprocessed, including:

[0082] The multi-source operational data is time-aligned according to a uniform sampling period (e.g., 1 min), and linear interpolation is used to obtain the equivalent value of the sampling time point for data whose sampling time deviates from the uniform time point.

[0083] For any running data, if the number of consecutive missing points is less than a preset threshold, linear interpolation of the valid data before and after is used to fill the gap; if the number of consecutive missing points is greater than the preset threshold, the sample of the corresponding time slice is removed.

[0084] All multi-source operational data were normalized using a standardization method with zero mean and unit variance.

[0085] A sample sequence is constructed using a sliding time window. The multi-source operational data from the current moment and L-1 moments prior forms an input sequence of length L. The measured SO2 concentration at the outlet at the corresponding preset future time shift is used as the label, resulting in a labeled training sample set covering the entire process of "coal quality—combustion—desulfurization—emission". The samples in the training sample set are divided into a training set, a validation set, and a test set in chronological order, with a ratio of approximately 7:2:1 (this ratio can be flexibly adjusted as needed during implementation).

[0086] Referring to step S2, this embodiment introduces mechanism features reflecting key physicochemical processes based on the above-mentioned desulfurization process characteristics, including at least the effective ammonia concentration and the SO2 mass transfer-reaction comprehensive factor.

[0087] The effective ammonia concentration characteristic is used to quantify the impact of chloride ions in the slurry on the competitive consumption of desulfurizing agents. The calculation formula is as follows:

[0088] C NH3_effective =C NH3×(1-α×exp(-β×(N / Cl)));

[0089] Among them, C NH3 The N / Cl ratio represents the ammonia nitrogen concentration in the slurry; N / Cl represents the molar ratio of ammonia nitrogen to chloride ions in the slurry; α and β are positive coefficients obtained through regression analysis of historical operating data, used to quantitatively characterize the competitive consumption of available ammonia by chloride ions. This characteristic reflects that, under the same dosing conditions, an increase in chloride ion concentration reduces the amount of available ammonia that can react with SO2.

[0090] The SO2 mass transfer-reaction comprehensive factor is a characteristic quantity calculated by combining the following parameters:

[0091] ① pH value of slurry;

[0092] ② Slurry temperature;

[0093] ③ Empty gas velocity of the desulfurization tower;

[0094] ④ Liquid-to-gas ratio in the desulfurization system;

[0095] The SO2 mass transfer-reaction comprehensive factor is used to characterize the overall removal rate of SO2 from the gas phase to the liquid phase and undergoing neutralization reaction within the desulfurization tower. This embodiment provides a method for representing the SO2 mass transfer-reaction comprehensive factor as follows:

[0096] k mt = k La ·f(pH);

[0097] Where, k La Let be the overall gas-liquid volumetric mass transfer coefficient per unit volume of slurry under empty tower conditions, and its expression is as follows:

[0098] k La = k0 (L / G) m U g n exp [-E a / (T +273.15) / R1];

[0099] In the formula, k0 is an empirical coefficient; L / G is the liquid-to-gas ratio of the desulfurization system; Ug is the empty gas velocity of the desulfurization tower; T is the slurry temperature; Ea is the apparent activation energy; R1 is the gas constant; m and n are exponential parameters obtained by regression using historical operating data.

[0100] f(pH) is a correction function characterizing the pH dependence of the acid-base neutralization reaction rate, and its approximately linear expression is as follows:

[0101] f(pH) = 1 + γ·(pH - pH0);

[0102] In the formula, γ is the fitting coefficient and pH0 is the reference pH value; the above parameters can be determined by using the least squares regression method based on the historical data of the unit.

[0103] This invention is not limited to the specific mathematical form described above. Those skilled in the art can select or establish other equivalent mass transfer-reaction correlations based on the characteristics of different units. In practical implementation, k can be... mt It can be added as an independent feature to the subsequent model input vector, or combined with the effective ammonia concentration C. NH3_effective Multiplication constructs a comprehensive mechanism characteristic k mt ·C NH3_effective It is used to simultaneously reflect the gas-liquid mass transfer capacity and the effective alkalinity supply capacity.

[0104] Referring to step S3, this embodiment uses the Transformer neural network model for SO2 prediction. The construction process of the Transformer neural network model is as follows:

[0105] For each time step t, coal quality characteristics, combustion characteristics, desulfurization process characteristics, and mechanism characteristics are concatenated into a one-dimensional feature vector z_t ∈ R^F. Time windows of length L {z_{t-L+1},…,z_t} are stacked to form an input matrix Z ∈ R^{L×F}.

[0106] First, the features at each time step are mapped to the d_model dimensional embedding space using a linear mapping: H (0) = ZW in +b in ;

[0107] Among them W in ∈ R^{F×d_model},b in As a trainable bias; then add a positional encoding P∈ R^{L×d_model} for each time step, resulting in:

[0108] ;

[0109] Then, N layers of encoders are stacked, each layer containing a multi-head self-attention sublayer and a feedforward fully connected sublayer, and residual connections and LayerNorm structures are used between the sublayers to achieve deep modeling of sequence features;

[0110] Finally, the hidden state h_L^{(N)} at the last time step is taken as the sequence representation and input into the regression output layer to obtain the predicted value of the outlet SO2 concentration.

[0111] For the constructed Transformer neural network model, this embodiment uses the particle swarm optimization algorithm to optimize hyperparameters such as the number of Transformer network layers, the number of attention heads, the hidden layer dimension, and the learning rate. Several strategies are introduced to improve the particle swarm optimization algorithm to avoid the problem of getting trapped in local optima by manual parameter tuning and simple search methods. In this embodiment, the multi-strategy improved particle swarm optimization algorithm includes the following improvement strategies:

[0112] A dynamic inertia weight strategy is adopted, in which the inertia weight decreases from the maximum value to the minimum value with the number of iterations, so as to balance global search capability and local search accuracy.

[0113] A random perturbation based on the Lévy flight mechanism is introduced during the particle position update process to enhance the ability to escape local extrema;

[0114] The cognitive and social factors are adaptively adjusted based on the population convergence state, enabling the algorithm to adaptively balance the influence of individual optimality and group optimality at different iteration stages.

[0115] Specifically, the particle swarm optimization algorithm improved based on the above strategy is as follows:

[0116] Let the particle swarm size be Np, the maximum number of iterations be Kmax, and the position vector xi of each particle i represent a set of hyperparameters to be optimized, with the corresponding velocity vector vi. In the k-th iteration, the velocity and position of particle i in the d-th dimension are updated as follows:

[0117] v i,d (k+1) = ω (k) v i,d (k) + c1 (k) r1(p i,d best - x i,d (k) ) + c2 (k) r2(g d best - x i,d (k) );

[0118] x i,d (k+1) = x i,d (k) + v i,d (k+1) ;

[0119] Where, p i,d best g represents the historical best position of particle i; d bestThe global optimal position is represented by r1 and r2, which are uniformly random numbers in the interval [0,1]. c1 is the cognitive factor, c2 is the social factor, and the inertia weight ω is the inertia weight. (k) Employ a dynamic decreasing strategy:

[0120] ω (k) = ω max - (ω max - ω min )·k / K max ;

[0121] Where ω max ω represents the preset maximum inertia weight. min This represents the preset minimum inertia weight, such as ω. max 0.9 can be taken, ω min 0.4 is acceptable.

[0122] To enhance the ability to escape local optima, a random perturbation based on the Lévy flight mechanism is introduced during the particle position update process. Specifically, after each particle position update, a Lévy flight perturbation is applied to the positions of some particles with a preset probability PLF.

[0123] x i (k+1) | Levy= x i (k+1) + δ Levy(λ);

[0124] Where δ represents the step size coefficient; Levy(λ) represents element-wise multiplication; it is a random vector that follows a power-law distribution; x i (k+1) Indicates the updated particle position; x i (k+1) | Levy represents the updated particle position after introducing the Levy flight perturbation; λ is the stability index of the Levy distribution.

[0125] Simultaneously, the cognitive factor c1 and the social factor c2 are adaptively adjusted based on the population convergence state, that is:

[0126] c1 (k) = c 1,max - (c 1,max - c 1,min )·k / K max ;

[0127] c2 (k) = c 2,min + (c 2,max - c 2,min )·k / Kmax ;

[0128] Among them, c 1,max This represents the preset maximum cognitive factor c1 value, c 1,min This represents the preset minimum cognitive factor c1 value; c 2,max c represents the preset maximum cognitive factor c2 value. 2,min This represents the preset minimum cognitive factor c2 value. For example, c 1,max 2.5 is acceptable, c 1,min Take 1.5, c 2,max Take 1.5, c 2,min Take 2.5.

[0129] Finally, the mean square error of SO2 prediction at the validation set outlet is used as the fitness function; the smaller the error, the higher the fitness. After setting the number of iterations or meeting the convergence condition, the combination of hyperparameters of the globally optimal particle is taken as the optimal hyperparameters of the Transformer model.

[0130] Referring to step S4, for the Transformer model with optimal hyperparameters, this embodiment trains it using the training sample set. Specifically, this embodiment introduces adversarial training based on the Fast Gradient Sign Method (FGSM) during the training process:

[0131] For each batch of input samples Z and their labels y, calculate the gradient of the loss function L(Z, y) with respect to the input Z. Construct adversarial examples:

[0132] Z adv = Z + ε·sign( );

[0133] Where ε is the disturbance amplitude; then Z and Z adv The samples are merged as training input, using a loss L(Z,y) that includes the original sample loss and an adversarial sample loss L(Z). adv The parameters are updated using the weighted total loss of y, thereby improving the model's robustness to small perturbations and anomalous fluctuations.

[0134] For example, the parameters in the model construction, optimization, and training process of this embodiment are as follows: Assume that the input feature dimension F=30, the time window length L=30, and the hyperparameters such as the embedding dimension dmodel, the number of encoder layers N, the number of attention heads h, the feedforward layer dimension dff, and the learning rate η are automatically optimized by multi-strategy PSO.

[0135] The particle swarm size Np=20, the maximum number of iterations Kmax=50, the inertia weights decrease linearly in the range [0.9, 0.4], c1 and c2 are adaptively adjusted, and the Levy flight perturbation probability PLF=0.2. Using the mean squared error on the validation set as the fitness function, an optimal combination of hyperparameters is obtained through particle swarm search, for example, N=3, h=4, dmodel=64, dff=256, η=5×10⁻⁶. -4 .

[0136] Under these optimal hyperparameters, the Adam optimizer is used to train the Transformer model, with a batch size of 128 and 80 training epochs. In each training epoch, adversarial training based on FGSM is introduced, with a perturbation amplitude ε=0.02 and adversarial sample weight λ=0.6.

[0137] Ultimately, this embodiment yields a fully trained and optimized Transformer model, which is then used for online prediction of actual SO2. Specifically, during the operation of the ammonia-based desulfurization system, the Transformer model, optimized through multi-strategy training and adversarial training, is deployed in the power plant's DCS or SIS system. It collects coal quality characteristics, combustion process characteristics, and desulfurization process characteristics in real time, calculates the intermediate mechanism characteristics in real time, and concatenates these real-time characteristics and intermediate mechanism characteristics into a model input sequence according to a preset time window. This sequence is then input into the trained and optimized Transformer neural network model to output a predicted SO2 concentration at the desulfurization tower outlet corresponding to a preset future time shift, thereby achieving advanced prediction of the outlet SO2 concentration.

[0138] To verify the effectiveness of the method in this embodiment, the following comparative models were trained on the same dataset:

[0139] ① An LSTM model using only the operating parameters of the desulfurization tower;

[0140] ② A Transformer model that uses full-process features but does not introduce mechanistic features;

[0141] ③ An unoptimized Transformer model using full-process and mechanistic features;

[0142] ④ The multi-strategy PSO-Transformer hybrid prediction model in this embodiment.

[0143] Test results show that:

[0144] ① The model in this embodiment can achieve a determination coefficient R² of over 0.97 on the test set, and the mean absolute error is approximately 2 mg / Nm³;

[0145] ② Compared to the LSTM model that only uses desulfurization operation parameters, the R² is improved by about 5 percentage points;

[0146] ③ Compared to the Transformer model without incorporating mechanistic features, the mean absolute error is reduced by approximately 15%;

[0147] ④ Under conditions of frequent load fluctuations and coal type switching, the maximum instantaneous error of the model in this embodiment is significantly lower than that of the model without adversarial training.

[0148] In summary, the intelligent prediction method for SO2 concentration at the outlet of ammonia desulfurization based on the fusion of multi-strategy optimization algorithm and mechanism features proposed in this invention can achieve high-precision and advanced prediction of SO2 concentration at the outlet while taking into account both physical mechanism and data-driven approach, and has good prospects for engineering application.

[0149] Compared with the prior art, this embodiment has the following beneficial effects:

[0150] 1. Achieve full-process modeling from coal quality to emissions, significantly improving the model's generalization ability.

[0151] This embodiment, based on the traditional desulfurization tower operating parameters, introduces coal quality characteristics such as sulfur content and chlorine content of the coal fed into the furnace, as well as combustion condition characteristics such as boiler load and primary and secondary air ratio. This enables the prediction model to simultaneously perceive changes at both the pollutant generation and treatment ends, providing better adaptability to varying coal types and load conditions. It also avoids the problem of insufficient extrapolation capability caused by modeling based solely on data from a single process segment.

[0152] 2. Improve the interpretability and reliability of prediction models by embedding mechanistic features.

[0153] The effective ammonia concentration characteristic proposed in this embodiment explicitly characterizes the competitive consumption of chloride ions by the ammonia-chlorine molar ratio. The proposed SO2 mass transfer-reaction integrated factor couples key operating parameters such as pH, temperature, empty tower gas velocity, and liquid-to-gas ratio to reflect the comprehensive removal rate of gas-liquid mass transfer and neutralization reaction. These mechanistic characteristics are embedded as independent inputs into the data-driven model, enabling the model not only to "calculate" but also to "explain the reasons," thus increasing its engineering credibility in operating condition extrapolation, parameter adjustment, and anomaly diagnosis.

[0154] 3. The multi-strategy PSO-Transformer structure significantly improves prediction accuracy and training stability.

[0155] By employing a particle swarm optimization algorithm with dynamic inertia weights, Lévy flight perturbations, and adaptive learning factors, the hyperparameters of the Transformer neural network are automatically optimized, avoiding the problems of manual parameter tuning and standard optimization algorithms easily getting trapped in local optima. Combined with an adversarial training strategy, the impact of perturbations such as load fluctuations and coal quality fluctuations on model performance is effectively suppressed. Validation results on one year of historical data from a 660MW unit show that the model of this invention achieves a determination coefficient R² of over 0.97 on the test set, with a mean absolute error of approximately 2mg / Nm³. Compared to comparative schemes that do not incorporate mechanistic features or use conventional models, the prediction accuracy is significantly improved.

[0156] 4. It possesses good engineering feasibility and application value.

[0157] The input parameters required in this embodiment all come from the power plant's existing online detection and operation monitoring system. No additional dedicated hardware equipment is required. After the model is trained, it can be directly deployed on the DCS or SIS system to achieve minute-level advance prediction of the outlet SO2 concentration. This provides quantitative basis for operators to adjust operating parameters such as liquid-gas ratio and ammonia injection volume in advance, thereby balancing desulfurization agent consumption and operating costs while ensuring compliance with environmental emission standards.

[0158] The present invention has been described above by way of example with reference to the accompanying drawings. Obviously, the specific implementation of the present invention is not limited to the above-described manner. Any non-substantial improvements made using the inventive concept and technical solution; or the direct application of the inventive concept and technical solution to other situations without modification, are all within the protection scope of the present invention.

Claims

1. A smart prediction method for SO2 concentration at the outlet of ammonia-based desulfurization based on the fusion of multi-strategy optimization algorithms and mechanistic features, characterized in that: The method includes: Step S1: Collect multi-source operating data of the ammonia desulfurization system and preprocess it to construct a training sample set. The multi-source operating data includes coal quality characteristics, combustion process characteristics, desulfurization process characteristics, and measured SO2 concentration at the desulfurization tower outlet. Step S2: Calculate intermediate mechanism characteristics based on the desulfurization process characteristics, including: effective ammonia concentration and SO2 mass transfer-reaction comprehensive factor; Step S3: Using the pre-treated coal quality characteristics, combustion process characteristics, desulfurization process characteristics, and intermediate mechanism characteristics as input samples, and the measured value of outlet SO2 concentration as output, construct a Transformer neural network model and use a multi-strategy improved particle swarm optimization algorithm to optimize its hyperparameters. Step S4: Train the optimized Transformer neural network model using the training sample set, and use the trained Transformer neural network model for actual SO2 concentration prediction.

2. The intelligent prediction method for SO2 concentration at the outlet of ammonia-based desulfurization based on the fusion of multi-strategy optimization algorithm and mechanism characteristics as described in claim 1, characterized in that: The preprocessing in step S1 includes: The multi-source operational data is time-aligned according to a unified sampling period, and linear interpolation is used to obtain the equivalent value of the time point for data whose sampling time deviates from the unified time point. For any running data, if the number of consecutive missing points is less than a preset threshold, linear interpolation of the valid data before and after is used to fill the gap; if the number of consecutive missing points is greater than the preset threshold, the sample of the corresponding time slice is removed. All multi-source operational data were normalized using a standardization method with zero mean and unit variance. By constructing a sample sequence through a sliding time window, the multi-source operational data of the current moment and several moments before it are combined into an input sequence, and the measured value of the outlet SO2 concentration at the corresponding future preset time shift is used as a label, thus obtaining a labeled training sample set.

3. The intelligent prediction method for SO2 concentration at the outlet of ammonia-based desulfurization based on the fusion of multi-strategy optimization algorithm and mechanism characteristics as described in claim 1, characterized in that: In step S2, the effective ammonia concentration characteristic is used to quantify the impact of chloride ions in the slurry on the competitive consumption of desulfurizing agents, and its calculation formula is as follows: C NH3_effective =C NH3 ×(1-α×exp(-β×(N / Cl))); Among them, C NH3 The N / Cl ratio represents the ammonia nitrogen concentration in the slurry; N / Cl represents the molar ratio of ammonia nitrogen to chloride ions in the slurry; α and β are positive coefficients obtained by regression from historical operating data, used to quantitatively characterize the competitive consumption of available ammonia by chloride ions.

4. The intelligent prediction method for SO2 concentration at the outlet of ammonia-based desulfurization based on the fusion of multi-strategy optimization algorithm and mechanism characteristics as described in claim 1, characterized in that: In step S2, the SO2 mass transfer-reaction comprehensive factor is a characteristic quantity calculated by integrating the following parameters: ① pH value of slurry; ② Slurry temperature; ③ Empty gas velocity of the desulfurization tower; ④ Liquid-to-gas ratio in the desulfurization system; The SO2 mass transfer-reaction comprehensive factor is used to characterize the overall removal rate of SO2 from the gas phase to the liquid phase and undergoing neutralization reaction in the desulfurization tower.

5. The intelligent prediction method for SO2 concentration at the outlet of ammonia-based desulfurization based on the fusion of multi-strategy optimization algorithm and mechanism characteristics as described in claim 4, characterized in that: The SO2 mass transfer-reaction comprehensive factor is expressed as follows: k mt = k La ·f(pH); Where, k La Let be the overall gas-liquid volumetric mass transfer coefficient per unit volume of slurry under empty tower conditions, and its expression is as follows: k La = k0 (L / G) m U g n exp [-E a / (T +273.15) / R1]; In the formula, k0 is an empirical coefficient; L / G is the liquid-to-gas ratio of the desulfurization system; Ug is the empty gas velocity of the desulfurization tower; T is the slurry temperature; Ea is the apparent activation energy; R1 is the gas constant; m and n are exponential parameters obtained by regression using historical operating data. f(pH) is a correction function characterizing the pH dependence of the acid-base neutralization reaction rate, and its expression is as follows: f(pH) = 1 + γ·(pH - pH0); In the formula, γ is the fitting coefficient and pH0 is the reference pH value.

6. The intelligent prediction method for SO2 concentration at the outlet of ammonia-based desulfurization based on the fusion of multi-strategy optimization algorithm and mechanism characteristics as described in claim 1, characterized in that: In step S3, constructing the Transformer neural network model includes: For each time step t, coal quality characteristics, combustion characteristics, desulfurization process characteristics, and mechanism characteristics are concatenated into a one-dimensional feature vector z_t ∈R^F, and time windows of length L {z_{t-L+1},…,z_t} are stacked to form an input matrix Z ∈ R^{L×F}. First, the features at each time step are mapped to the d_model dimensional embedding space using a linear mapping: H (0) = ZW in + b in ; Among them W in ∈ R^{F×d_model},b in As a trainable bias; then add a positional encoding P ∈ R^{L×d_model} for each time step, resulting in: ; Then, N layers of encoders are stacked, each layer containing a multi-head self-attention sublayer and a feedforward fully connected sublayer, and residual connections and LayerNorm structures are used between the sublayers to achieve deep modeling of sequence features; Finally, the hidden state h_L^{(N)} at the last time step is taken as the sequence representation and input into the regression output layer to obtain the predicted value of the outlet SO2 concentration.

7. The intelligent prediction method for SO2 concentration at the outlet of ammonia-based desulfurization based on the fusion of multi-strategy optimization algorithm and mechanism characteristics as described in claim 1, characterized in that: In step S3, the hyperparameter optimization using the multi-strategy improved particle swarm optimization algorithm includes: improving the particle swarm optimization algorithm using a dynamic inertia weight strategy, namely: Let the particle swarm size be Np, the maximum number of iterations be Kmax, and the position vector xi of each particle i represent a set of hyperparameters to be optimized, with the corresponding velocity vector vi. In the k-th iteration, the velocity and position of particle i in the d-th dimension are updated as follows: v i,d (k+1) = ω (k) v i,d (k) + c1 (k) r1(p i,d best - x i,d (k) ) + c2 (k) r2(g d best - x i,d (k) ); x i,d (k+1) = x i,d (k) + v i,d (k+1) ; Where, p i,d best g represents the historical best position of particle i; d best The global optimal position is represented by r1 and r2, which are uniformly random numbers in the interval [0,1]. c1 is the cognitive factor, c2 is the social factor, and the inertia weight ω is the inertia weight. (k) Employ a dynamic decreasing strategy: oh (k) = ω max - (oh max - oh min )·k / K max ; Where ω max ω represents the preset maximum inertia weight. min This represents the preset minimum inertia weight; The mean square error of SO2 prediction at the outlet is taken as the fitness function, and the smaller the error, the higher the fitness. After setting the number of iterations or meeting the convergence condition, the combination of hyperparameters of the globally optimal particle is taken as the optimal hyperparameters of the Transformer model.

8. The intelligent prediction method for SO2 concentration at the outlet of ammonia desulfurization based on the fusion of multi-strategy optimization algorithm and mechanism characteristics as described in claim 7, characterized in that: In step S3, the hyperparameter optimization using the multi-strategy improved particle swarm optimization algorithm further includes: introducing a random perturbation based on the Lévy flight mechanism during the particle position update process to enhance the ability to escape local optima, i.e.: After each particle position update, a Levy flight perturbation is applied to the positions of some particles with a preset probability using PLF: x i (k+1) |Lions= x i (k+1) +δ Lions (λ); Where δ represents the step size coefficient; Levy(λ) represents element-wise multiplication; it is a random vector that follows a power-law distribution; x i (k+1) Indicates the updated particle position; x i (k+1) |Levy represents the updated particle position after introducing the Levy flight perturbation; λ is the stability index of the Levy distribution.

9. A method for intelligent prediction of SO2 concentration at the outlet of ammonia-based desulfurization based on the fusion of multi-strategy optimization algorithm and mechanism characteristics, as described in claim 7 or 8, characterized in that: In step S3, the hyperparameter optimization using the multi-strategy improved particle swarm optimization algorithm further includes: adaptively adjusting the cognitive factor c1 and the social factor c2 according to the population convergence state, that is: c1 (k) = c 1,max - (c 1,max - c 1,min )·k / K max ; c2 (k) = c 2,min + (c 2,max - c 2,min )·k / K max ; Among them, c 1,max This represents the preset maximum cognitive factor c1 value, c 1,min This represents the preset minimum cognitive factor c1 value; c 2,max c represents the preset maximum cognitive factor c2 value. 2,min This represents the preset minimum cognitive factor c2 value.

10. The intelligent prediction method for SO2 concentration at the outlet of ammonia-based desulfurization based on the fusion of multi-strategy optimization algorithm and mechanism characteristics as described in claim 1, characterized in that: In step S4, an adversarial training strategy is used to train the model, including: Adversarial examples are constructed using the Fast Gradient Signed Method (FGSM): For each batch of input samples Z and their labels y, the gradient of the loss function L(Z, y) with respect to the input Z is calculated. Construct adversarial examples: WITH adv = Z + ε sign( ); Where ε is the disturbance amplitude; then Z and Z adv The samples are merged as training input, using a loss L(Z, y) that includes the original sample loss and the adversarial sample loss L(Z). adv The parameters are updated using the weighted total loss of (y).