SO2 concentration prediction method based on double-stage attention and double-flow memory adjustment GRU

By introducing a two-stage attention mechanism and a dual-current memory adjustment structure into the GRU network, the problems of information redundancy and timing information attenuation in complex industrial processes are solved, and the prediction accuracy and robustness are significantly improved, and the accurate prediction of SO2 concentration in the flue gas desulfurization process of thermal power plants is achieved.

CN120126601AActive Publication Date: 2025-06-10JIANGNAN UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510141462.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-10
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

Existing soft measurement technologies fail to fully consider the nonlinearity and dynamic nature of complex industrial processes, resulting in information redundancy and timing information attenuation, weakening the model's ability to capture key information and long-term dependencies, and thus reducing prediction accuracy.

Method used

A GRU network (FTA-DS-MRGRU) based on dual-stage attention and dual-stream memory regulation is adopted to introduce feature attention modules in the feature extraction stage to suppress uncorrelated features and noise; a timing attention module is introduced in the prediction output stage to capture long-term time series dependencies, and a complementary information flow is generated through the time-dependent and causal correlation dual-stream information extraction structure.

Benefits of technology

The model's ability to capture key information and long-term dependencies in complex industrial processes is significantly improved, the prediction accuracy and robustness are enhanced, and the net flue gas SO2 content after the flue gas desulfurization process of thermal power plants is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126601A_ABST
    Figure CN120126601A_ABST
Patent Text Reader

Abstract

The invention discloses a SO2 concentration prediction method based on dual-stage attention and double-flow memory adjustment GRU, which belongs to the field of industrial process soft measurement, and comprises the following steps: after constructing a DS-MRGRU network, introducing feature attention to capture key features and suppress irrelevant features and noise in a feature extraction stage; and introducing a time sequence attention to capture a long-time sequence dependency relationship in a prediction output stage, and constructing an FTA-DS-MRGRU soft measurement model with enhanced key information extraction capability for predicting the content of SO2 in purified flue gas. According to the method, historical information and currently input key features are extracted through DS-MRGRU to generate a complementary information stream to enhance the generalization ability and prediction performance of the model, feature attention and time sequence attention are fused to suppress redundant information, the interpretability of the model is improved, historical important moment information is mined, the long-time sequence prediction effect is improved, and the prediction efficiency is improved. And accurate prediction of the SO2 content of the purified flue gas is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a SO 2 concentration prediction method based on dual-stage attention and two-stream memory-regulated GRU, belonging to the field of soft sensors in modern industrial processes. Background Art

[0002] In complex industrial processes, it is necessary to measure certain key quality variables that are difficult to directly or real-time measure but are closely related to product quality and production safety in real time to meet the requirements of process control. Soft sensor technology can use easily measurable auxiliary variables in industrial processes to estimate difficult-to-measure dominant variables in real time. Due to its economic reliability, rapid response, and easy achievement of real-time monitoring and control of product quality in online detection, it has been widely used in the process manufacturing industry and the field of process control. For example, taking the desulfurization process of a thermal power plant as an example, in this process, it is necessary to measure the SO 2 flue gas emission concentration through a soft sensor model to ensure that its emissions meet the national air pollutant emission standards.

[0003] Soft sensor modeling methods can generally be divided into three categories: mechanism analysis-based modeling, data-driven modeling, and mechanism-data hybrid-driven modeling. Mechanism-included modeling methods are mainly based on a profound understanding of the process technology mechanism, and the modeling process is complex and difficult to solve. Data-driven modeling methods do not overly rely on the process technology mechanism and only need to extract key feature information from process data to establish a soft sensor model. Among them, the soft sensor method based on artificial neural networks has become a research hotspot in the field of soft sensors due to its excellent modeling performance. In recent years, deep neural networks have developed rapidly, and their feature representation ability has become stronger. The gated recurrent neural network (GRNN) can not only effectively describe the dynamic time series characteristics between data by introducing a time series feedback mechanism and a threshold unit structure, but also alleviate the problems of gradient disappearance and explosion existing in the standard RNN. Common GRNNs include long short-term memory (LSTM) neural networks and gated recurrent unit (GRU) neural networks. Among them, the GRU network is developed on the basis of the LSTM network and is an important variant of the long short-term memory network. It not only alleviates the long-term dependence problem existing in the RNN, but also has a simple internal structure and can improve the network calculation efficiency, which has attracted extensive attention from research scholars in the field.

[0004] The similarities between GRU and LSTM are that each neuron is a processing unit containing several threshold controllers, and each threshold controller is used to determine whether the input information is useful and control the inflow of relevant information; the difference is that each GRU processing unit only contains two threshold controllers and a sequential output, so it has fewer parameters and a simpler network structure, as Figure 2 shown. The GRU network threshold mechanism consists of two threshold controllers, the reset gate and the update gate, and the candidate hidden state. Among them, the reset gate is used to determine whether the candidate hidden state at the current moment depends on the hidden state at the previous moment and the degree of dependence; the role of the update gate is to control the update of the hidden state from the (t-1)th moment to the tth moment; the candidate hidden state is used to control the retention of relevant input information. The expressions of GRU are shown in the following formulas (1) to (3):

[0005] r t = σ(W r x t + U r h t-1 + b r ) (1)

[0006] z t = σ(W z x t + U z h t-1 + b z ) (2)

[0007]

[0008] Among them, x t and h t-1 represent the input at the current moment and the output of the hidden state at the previous moment, respectively; r t and z t are the outputs of the reset gate and the update gate at the current moment, respectively, and their activation function σ(·) is the sigmoid function; is the output of the candidate hidden state at the current moment, and its activation function tanh(·) is the hyperbolic tangent function; W r , W z and are the network input weight vectors, U r , U z and are the hidden weight vectors, b r , b z and are the bias vectors; ⊙ represents the Hadamard product operation. The expression of the hidden state h t at the current moment is shown in formula (4):

[0009]

[0010] As described above, the GRU neural network realizes the update and dynamic memory of information through its threshold controller and time series feedback mechanism, and realizes the persistence of information retention. By introducing a time series feedback mechanism and a threshold unit structure, the GRU neural network can not only effectively describe the dynamic time series characteristics between data, but also alleviate the problems of gradient disappearance and explosion existing in the standard RNN. Moreover, the GRU neural network is developed on the basis of the LSTM neural network, and its internal structure is simpler, which can effectively improve the network calculation efficiency.

[0011] However, although the GRU network can effectively mine the long-term dependence relationship between time series data, it is difficult to capture the different variable characteristics at different time steps, and the GRU network uses a single gating mechanism to update the hidden state, resulting in a linear constraint relationship during its calculation, restricting the full passage of information flow, thereby affecting the model prediction performance. At the same time, with the continuous development and application of industrial big data technology, it is found that in the actual industrial process modeling, the problems of information redundancy and time series information attenuation will weaken the model's ability to capture key information and long-term dependence relationships, resulting in poor prediction effects. Summary of the Invention

[0012] In order to solve the problem that the existing soft measurement technology at present does not fully consider the nonlinearity, dynamics of complex industrial processes, and information redundancy and time series information attenuation will weaken the soft measurement model's ability to capture key information and long-term dependence relationships, resulting in a decrease in the accuracy of soft measurement results, the present invention provides a SO 2 concentration prediction method based on dual-stage attention and two-stream memory-regulated GRU, and the technical solution is as follows:

[0013] As an aspect of the present invention, there is provided a SO 2 concentration prediction method based on dual-stage attention and two-stream memory-regulated GRU, comprising:

[0014] S100. After obtaining the historical data of the process variables and target variables in the flue gas desulfurization process of a thermal power plant, perform preprocessing to obtain several time series of historical data, and divide the several time series into a training set and a test set; wherein, the target variable is the net flue gas SO 2 concentration at the outlet of the secondary absorption tower, and the process variables are other parameters in the flue gas desulfurization process of the thermal power plant;

[0015] S200. Construct a DS-MRGRU network for extracting the time relationship and causal relationship between current information and historical information, including a time-related flow information extraction structure and a dynamic causal-related flow information extraction structure in parallel;

[0016] S300. Introduce a feature attention module in the feature extraction stage of the DS-MRGRU network to capture key features and suppress irrelevant features and noise, and introduce a temporal attention module in the prediction output stage of the DS-MRGRU network to capture long-term sequence dependencies, so as to construct an FTA-DS-MRGRU soft sensor model with enhanced key information extraction ability;

[0017] S400. Use the training set in S100 to train the FTA-DS-MRGRU soft sensor model through the backpropagation algorithm, and use the test set in S100 to test the FTA-DS-MRGRU soft sensor model to obtain the target FTA-DS-MRGRU soft sensor model;

[0018] S500. Obtain the real-time data of the process variables in the flue gas desulfurization process of the thermal power plant, input it into the target FTA-DS-MRGRU soft sensor model, and obtain the SO 2 content of the clean flue gas after the flue gas desulfurization process of the thermal power plant.

[0019] Furthermore, the time-related flow information extraction structure includes a first time-gate-based TGRU layer, a Dropout layer, and a second time-gate-based TGRU layer connected in sequence;

[0020] The expressions of the first time-gate-based TGRU layer and the second time-gate-based TGRU layer are:

[0021]

[0022] Among them, represents the input at time t of the TGRU layer, represents the time-related hidden state output at time t-1 in the input time series of the TGRU layer, σ(·) represents the activation function sigmoid function, represents the candidate time-related hidden state output at time t in the input time series of the TGRU layer, ⊙ represents the Hadamard product operation, W T represents the network input weight vector of the TGRU layer, U T represents the hidden weight vector of the TGRU layer, b T represents the bias vector of the TGRU layer, represents the update gate corresponding to the TGRU layer, represents the input weight vector of the update gate corresponding to the TGRU layer, represents the hidden weight vector of the update gate corresponding to the TGRU layer, represents the bias vector of the update gate corresponding to the TGRU layer.

[0023] Further, the dynamic causal correlation flow information extraction structure includes a first causal gate-based CGRU layer, a Dropout layer, and a second causal gate-based CGRU layer connected in sequence;

[0024] The expressions of the first causal gate-based CGRU layer and the second causal gate-based CGRU layer are:

[0025]

[0026] Among them, represents the input at time t of the CGRU layer, represents the causal correlation hidden state output at time t - 1 in the input time series of the TGRU layer, and σ(·) represents the activation function sigmoid function, represents the candidate causal correlation hidden state output at time t in the input time series of the CGRU layer, and ⊙ represents the Hadamard product operation, W c represents the network input weight vector of the CGRU layer, U c represents the hidden weight vector of the CGRU layer, b C represents the bias vector of the CGRU layer, represents the update gate corresponding to the CGRU layer, represents the input weight vector of the update gate corresponding to the CGRU layer, represents the hidden weight vector of the update gate corresponding to the CGRU layer, represents the bias vector of the update gate corresponding to the CGRU layer.

[0027] Further, S300 includes:

[0028] S310. After the feature attention module assigns feature attention weights to the input time series, the time series is weighted according to the feature attention weights, and a weighted time series is output;

[0029] S320. The time gate in the time correlation flow information extraction structure extracts the temporal relationship in the weighted time series to obtain a time correlation hidden state output, and the causal gate in the dynamic causal correlation flow information extraction structure extracts the causal relationship of the weighted time series to obtain a causal correlation hidden state output;

[0030] S330. After the first temporal attention module assigns first temporal attention weights to the time correlation hidden state output and performs weighting, a comprehensive time feature focusing on historical key points is obtained, and after the second temporal attention module assigns second temporal attention weights to the causal correlation hidden state output and performs weighting, a comprehensive causal feature focusing on historical key points is obtained;

[0031] S340. After fusing the comprehensive time feature and the comprehensive causal feature through the merging layer, output through the fully connected layer to obtain the final output result.

[0032] Furthermore, the expression of the feature attention module is:

[0033] q t1 = W q 1 x t + b q 1

[0034] e t i = V 1 i σ(W 1 i q t1 + U 1 i x t i + b 1 i ), 1 ≤ i ≤ n

[0035]

[0036] Among them, x t represents the input at time t of the feature attention module, x t i represents the i-th variable at time t in the input time series, q t1 represents the query vector of the feature attention module, e t i represents the feature attention score corresponding to the i-th variable at time t in the input time series, α t i represents the feature attention weight corresponding to the i-th variable at time t in the input time series, represents the input vector after being weighted by the feature attention module, n represents the number of input features, W q 1 、b q 1 、W 1 i 、U 1 i 、b 1 i 、V 1 i all represent the feature attention parameters that are continuously updated and optimized during the model training process.

[0037] Furthermore, the expressions of the first time series attention module and the second attention module are:

[0038] q t2 = W q 2 h t + b q 2

[0039] g j = V 2 j σ(W 2 j q t2 + U 2 j h t-T+j + b 2 j ), 1 ≤ j ≤ T

[0040]

[0041]

[0042] Among them, h t represents the input at time t of the temporal attention module, q t2 represents the query vector of the temporal attention module, g j represents the temporal attention score corresponding to the hidden state at the j-th time step in the input time series, β j represents the temporal attention weight corresponding to the hidden state at the j-th time step in the input time series, c t represents the vector after weighted summation of the outputs of each hidden state, T represents the size of the time step taken, W q 2 , b q 2 , W 2 j , U 2 j , b 2 j , V 2 j all represent the temporal attention parameters updated and optimized during the model training process.

[0043] Furthermore, S100 includes:

[0044] S110. After obtaining the historical data of the process variables and target variables in the flue gas desulfurization process of the thermal power plant, perform missing value processing, outlier processing, and data standardization in sequence;

[0045] S120. Use the sliding window technique to convert the historical data after data standardization into several time series with a length of T, and divide the several time series into a training set and a test set.

[0046] Further, the training of the FTA-DS-MRGRU soft sensor model using the training set in S100 through the backpropagation algorithm in S400 includes:

[0047] Performing forward calculation on the FTA-DS-MRGRU soft sensor model using the training set in S100 to calculate the network weights of the FTA-DS-MRGRU soft sensor model;

[0048] Calculating the loss function value of the FTA-DS-MRGRU soft sensor model through the mean square root error loss function to obtain the corresponding error term, and the calculation formula is:

[0049]

[0050] where N represents the total number of test set samples, y t represents the true value of the SO 2 concentration in the clean flue gas, represents the predicted value of the SO 2 concentration in the clean flue gas;

[0051] Based on the corresponding error term, the Adam optimization algorithm is used to update the gradient of the network weights.

[0052] As another aspect of the present invention, a target detection method is provided. The method uses the above method to detect the gas to be detected, and the gas to be detected is the gas generated in the industrial process.

[0053] As another aspect of the present invention, a computer device is provided, including a processor and a memory. The memory is used to store computer instructions, and the computer instructions are loaded and executed by the processor to implement the above-mentioned SO 2 concentration prediction method based on dual-stage attention and dual-stream memory-regulated GRU. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0055] Figure 1 It is a brief flowchart of the SO 2 concentration prediction method based on dual-stage attention and dual-stream memory-regulated GRU provided by the present invention;

[0056] Figure 2 It is a structural diagram of the gated recurrent unit GRU in the soft sensor modeling method related to the present invention;

[0057] Figure 3 The structural diagram of MRGRU provided by the present invention;

[0058] Figure 4 The structural diagram of the feature attention module provided by the present invention;

[0059] Figure 5 The structural diagram of the temporal attention module provided by the present invention;

[0060] Figure 6 The structural diagram of FTA-DS-MRGRU provided by the present invention;

[0061] Figure 7 A brief schematic diagram of the flue gas desulfurization process of a certain thermal power plant provided by the present invention;

[0062] Figure 8 For the net flue gas SO 2 The feature attention weight diagram in the concentration emission prediction at the outlet of the secondary absorption tower;

[0063] Figure 9 For the net flue gas SO 2 The temporal attention weight diagram in the concentration emission prediction at the outlet of the secondary absorption tower.

[0064] Figure 10A For the comparison diagram of the true value and the predicted value of the net flue gas SO 2 concentration emission at the outlet of the secondary absorption tower based on GRU;

[0065] Figure 10B For the comparison diagram of the true value and the predicted value of the net flue gas SO 2 concentration emission at the outlet of the secondary absorption tower based on MRGRUGRU;

[0066] Figure 10C For the comparison diagram of the true value and the predicted value of the net flue gas SO 2 concentration emission at the outlet of the secondary absorption tower based on DS-MRGRU;

[0067] Figure 10D For the comparison diagram of the true value and the predicted value of the net flue gas SO 2 concentration emission at the outlet of the secondary absorption tower based on TA-DS-MRGRU;

[0068] Figure 10E For the comparison diagram of the true value and the predicted value of the net flue gas SO 2 concentration emission at the outlet of the secondary absorption tower based on FA-DS-MRGRU;

[0069] Figure 10F For the comparison diagram of the true value and the predicted value of the net flue gas SO 2Comparison chart of the true value and predicted value of concentration emissions.

[0070] Figure 11A For the predicted error distribution map of the net flue gas SO 2 concentration emissions at the outlet of the secondary absorption tower based on GRU;

[0071] Figure 11B For the predicted error distribution map of the net flue gas SO 2 concentration emissions at the outlet of the secondary absorption tower based on MRGRU;

[0072] Figure 11C For the predicted error distribution map of the net flue gas SO 2 concentration emissions at the outlet of the secondary absorption tower based on DS-MRGRU;

[0073] Figure 11D For the predicted error distribution map of the net flue gas SO 2 concentration emissions at the outlet of the secondary absorption tower based on TA-DS-MRGRU;

[0074] Figure 11E For the predicted error distribution map of the net flue gas SO 2 concentration emissions at the outlet of the secondary absorption tower based on FA-DS-MRGRU;

[0075] Figure 11F For the predicted error distribution map of the net flue gas SO 2 concentration emissions at the outlet of the secondary absorption tower based on FTA-DS-MRGRU. Detailed implementation method

[0076] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the accompanying drawings.

[0077] Example 1:

[0078] The embodiment of the present invention provides a method for predicting SO 2 concentration based on dual-stage attention and two-stream memory-regulated GRU, as Figure 1 shown, including:

[0079] S100. After obtaining the historical data of the process variables and target variables in the flue gas desulfurization process of a thermal power plant, perform preprocessing to obtain several time series of the historical data, and divide the several time series into a training set and a test set; wherein, the target variable is the net flue gas SO 2 concentration at the outlet of the secondary absorption tower, and the process variables are other parameters in the flue gas desulfurization process of the thermal power plant;

[0080] Specifically, the method for predicting SO 2The concentration prediction method selects, through mechanism analysis and expert experience, important process variables that may affect the prediction of the SO 2 concentration at the outlet of the secondary absorption tower from the industrial distributed control system of flue gas desulfurization in thermal power plants as input variables, that is, process variables, and the SO 2 concentration at the outlet of the secondary absorption tower as the output variable, that is, the target variable. Continuous and uniform sampling is carried out at preset time intervals, and then missing value processing, outlier processing based on the Pauta criterion (3σ criterion), and data standardization processing based on the normalization method (z-score method) are carried out in sequence to obtain several time series of historical data.

[0081] S200. Construct a DS-MRGRU (Dual-Stream MRGRU) network for extracting the temporal relationship and causal relationship between current information and historical information, including a parallel temporal correlation stream information extraction structure and a dynamic causal correlation stream information extraction structure;

[0082] In the embodiment of the present invention, a time gate and a causal gate are introduced into the hidden state of the traditional GRU network to optimize the network structure, and a Memory Regulation Gated Recurrent Unit (MRGRU) is constructed to cooperate with the original gating strategy to capture key feature information in historical information and current input. Then, based on the MRGRU, a Time-Aware Gated Recurrent Unit (TGRU) with only a time gate and a Causal-Aware Gated Recurrent Unit (CGRU) are differentiated, and a dual-stream information extraction network including a temporal correlation stream information extraction structure and a dynamic causal correlation stream information extraction structure is constructed, that is, a DS-MRGRU network.

[0083] S300. Introduce a feature attention module in the feature extraction stage of the DS-MRGRU network to capture key features and suppress irrelevant features and noise, and introduce a temporal attention module in the prediction output stage of the DS-MRGRU network to capture long-term time series dependence relationships, and construct an FTA-DS-MRGRU soft sensor model with enhanced key information extraction ability;

[0084] Specifically, the embodiment of the present invention adopts a two-stage attention mechanism to extract key information. The feature attention module introduced in the model feature extraction stage is mainly used to capture key features and suppress the influence of irrelevant features and noise; the temporal attention module introduced in the prediction output stage is mainly used to adaptively focus on the key time point information in historical data and capture long-time series dependence relationships. The attention weights provide an intuitive quantitative explanation of the contribution degree of each input data to the model prediction result, which helps to understand the model decision-making process and enhance the transparency and interpretability of the model.

[0085] Embed the two-stage attention mechanism into the DS-MRGRU network to capture complex time series features and establish a soft measurement model with stronger prediction performance. First, take X = {x t-T+1 , x t-T+2 , …, x t} as input data and input it into the feature attention module. Here, the input data is a time series with a time step of T, and x t represents the historical data of n process variables at time t in this time series. The feature attention module assigns different weights α = {α 1 , α 2 , …, α n} to the n process variables through a weighting mechanism to obtain the data processed by the feature attention module. Then, for and respectively, two information flows independently learn time-related features and dynamic causal-related features, and generate T hidden state information with different representations and respectively. Then, the temporal attention module is used to assign weights to the and hidden states respectively. The weights assigned to the hidden states corresponding to the time flow and the causal flow are specifically: and And two comprehensive vectors c t1 and c t2 are generated in a summing manner, and then merged and integrated into the final vector c t . Finally, the network continuously optimizes the parameters through backpropagation, and obtains the output variable t according to c The expression of the final vector c t is shown in the following formula (5):

[0086]

[0087] S400. Use the training set in S100 to train the FTA-DS-MRGRU soft sensor model through the backpropagation algorithm, and use the test set in S100 to test the FTA-DS-MRGRU soft sensor model to obtain the target FTA-DS-MRGRU soft sensor model;

[0088] S500. Obtain the real-time data of the process variables in the flue gas desulfurization process of the thermal power plant, and input it into the target FTA-DS-MRGRU soft sensor model to obtain the SO 2 content of the clean flue gas after the flue gas desulfurization process of the thermal power plant.

[0089] In summary, based on the existing gating strategy of GRU, the present invention designs two auxiliary memory adjustment gates: a time gate and a causal gate, which are used to finely regulate past information and current information respectively. And further construct a two-stream information extraction structure related to time and causality to generate complementary information flows, significantly enhancing its generalization ability and prediction performance in different industrial scenarios. At the same time, considering that in the modeling of nonlinear dynamic industrial processes, information redundancy and temporal information attenuation problems will weaken the model's ability to capture key information and long-term dependence relationships, a two-stage attention mechanism is embedded in the DS-MRGRU network, and a method for predicting the SO 2 concentration of clean flue gas based on two-stage attention and two-stream memory adjustment GRU is proposed. This method adaptively extracts key feature variables using a feature attention mechanism in the feature extraction stage, suppresses the influence of redundant information, and improves the interpretability of the model; and uses a temporal attention mechanism to mine historical important moment information in the prediction output stage to improve the prediction effect of the model in long time series, realizing the accurate prediction of the SO 2 content of the clean flue gas after the flue gas desulfurization process of the thermal power plant, which is a key quality variable.

[0090] Embodiment 2:

[0091] S100 further includes:

[0092] S110. After obtaining the historical data of the process variables and target variables in the flue gas desulfurization process of the thermal power plant, perform missing value processing, outlier processing, and data standardization in sequence;

[0093] Specifically, the method for predicting the SO 2 concentration based on two-stage attention and two-stream memory adjustment GRU provided by the embodiment of the present invention selects important process variables that may affect the prediction of the SO 2 concentration of the clean flue gas at the outlet of the secondary absorption tower as input variables, that is, process variables, through mechanism analysis and expert experience. The SO 2The concentration is used as the output variable, that is, the target variable, and continuous and uniform sampling is carried out on it at time intervals of 5 minutes to obtain the input-output variable dataset denoted as (X all , Y all ). After that, the collected (X all , Y all ) are successively subjected to missing value processing, outlier processing, and data standardization. It includes: for variables that only contain partial time points, if there are too many missing data and they cannot be supplemented, such variables will be deleted, and variables with all constant values in the sample will be deleted; for variables with some data being null values, the null values will be replaced with the average of the two adjacent data; secondly, according to the process requirements and operation experience, the operation range of the original data variables is summarized, and then a part of the samples outside this range is removed by using the maximum and minimum clipping methods, and outliers are removed according to the 3σ criterion. The 3σ criterion means that first assume that a set of measured data only contains random errors, calculate and process it to obtain the standard deviation, determine an interval according to a certain probability, and consider that any error exceeding this interval does not belong to random error but gross error, and the data containing this error should be removed. First, the measured variable is measured with equal precision to independently obtain x 1 , x 2 , …, x N , calculate its arithmetic mean and the residual error (i = 1, 2, …, N), and calculate the standard deviation σ. If the residual error v i of a certain measured value x t (1 ≤ t ≤ N) satisfies the following formula (6):

[0094]

[0095] then it is considered that this error belongs to gross error, and the data x t containing this error should be removed;

[0096] Finally, the input variables and output variables are standardized by the z-score method.

[0097] S120. Use the sliding window technique to convert the historical data after data standardization into several time series with a length of T, and use the first 80% of the time series as the training set and the remaining 20% of the time series as the test set.

[0098] S200. Construct a DS-MRGRU network for extracting the temporal relationship and causal relationship between current information and historical information, including a parallel temporal correlation flow information extraction structure and a dynamic causal correlation flow information extraction structure;

[0099] Specifically, introduce the temporal gate T t and the causal gate Ct Optimize the network structure and retain z t and 1 - z t as mutually exclusive relationship gates. As shown in Figure 3 , construct MRGRU. The time gate screens the key information in the hidden state at the previous moment to achieve memory regulation of historical information and effectively capture information features related to time; the causal gate is used to select important information in the current candidate state, regulate the influence ability of the current information on the hidden state, and then capture the potential dynamic causal relationship between information, significantly enhancing the flexibility of the gating strategy. Based on this, MRGRU effectively changes the linear constraint relationship while retaining the mutually exclusive relationship, can cooperate with the original gating strategy to capture key feature information in historical information and current input, enriches the transmitted information, finely adjusts the information transmission, improves the information processing ability of the model, and can more effectively adapt to the dynamic changes in the industrial environment. The state update equation of MRGRU is shown in the following formulas (7) to (9):

[0100] T t = σ(W T x t + U T h t-1 + b T ) (7)

[0101] c t = σ(W C x t + U C h t-1 + b C ) (8)

[0102]

[0103] Among them, T t represents the time gate, C t represents the causal gate, W T represents the time gate network input weight vector, W c represents the causal gate network input weight vector, U T represents the time gate implicit weight vector, U c represents the causal gate implicit weight vector, b T represents the time gate bias vector, b C represents the causal gate bias vector.

[0104] Based on MRGRU, a TGRU layer based on time gates and a CGRU layer based on causal gates are differentiated. A time-related flow information extraction structure is constructed according to the TGRU layer based on time gates, and a dynamic causal-related flow information extraction structure is constructed according to the CGRU layer based on causal gates, obtaining a dual-stream information extraction network DS-MRGRU. The time-related flow information extraction structure is regulated by the time gate T t to control the memory and forgetting of historical information, so as to focus on extracting the temporal relationship in the time series; the dynamic causal-related flow information extraction structure is controlled by the causal gate C t to control the influence of the current input information on the prediction result, aiming to capture the dynamic causal relationship between the current input and the prediction result, and ensure that the model can accurately respond to the changes in the current environment.

[0105] Specifically, as Figure 6 shown, the time-related flow information extraction structure includes a first TGRU layer based on time gates, a Dropout layer, and a second TGRU layer based on time gates connected in sequence;

[0106] The expressions of the first TGRU layer based on time gates and the second TGRU layer based on time gates are shown in formulas (10) and (11):

[0107]

[0108] Among them, represents the input at time t of the TGRU layer, represents the output of the time-related hidden state at time t - 1 in the input time series of the TGRU layer, σ(·) represents the activation function sigmoid function, represents the output of the candidate time-related hidden state at time t in the input time series of the TGRU layer, ⊙ represents the Hadamard product operation, W T represents the network input weight vector of the TGRU layer, U T represents the hidden weight vector of the TGRU layer, b T represents the bias vector of the TGRU layer, represents the update gate corresponding to the TGRU layer,, represents the input weight vector of the update gate corresponding to the TGRU layer, represents the hidden weight vector of the update gate corresponding to the TGRU layer, represents the bias vector of the update gate corresponding to the TGRU layer.

[0109] As Figure 6 shown, the dynamic causal-related flow information extraction structure includes a first CGRU layer based on causal gates, a Dropout layer, and a second CGRU layer based on causal gates connected in sequence;

[0110] The expressions of the first causal-gate based CGRU layer and the second causal-gate based CGRU layer are shown in formulas (12) and (13):

[0111]

[0112] where denotes the input of the CGRU layer at time t, denotes the causally related hidden state output at time t-1 in the input time series of the TGRU layer, σ(·) denotes the activation function sigmoid function, denotes the candidate causally related hidden state output at time t in the input time series of the CGRU layer, ⊙ denotes the Hadamard product operation, W c denotes the network input weight vector of the CGRU layer, U c denotes the hidden weight vector of the CGRU layer, b C denotes the bias vector of the CGRU layer, denotes the update gate corresponding to the CGRU layer, denotes the input weight vector of the update gate corresponding to the CGRU layer, denotes the hidden weight vector of the update gate corresponding to the CGRU layer, denotes the bias vector of the update gate corresponding to the CGRU layer.

[0113] S300. Introduce a feature attention module in the feature extraction stage of the DS-MRGRU network to capture key features and suppress irrelevant features and noise, and introduce a temporal attention module in the prediction output stage of the DS-MRGRU network to capture long-time series dependency relationships, and construct an FTA-DS-MRGRU soft measurement model with enhanced key information extraction ability;

[0114] As Figure 6 shown, the feature attention module aims to dynamically assign attention weights to input features before the input data enters the dual information flow, mine out the features that have an important impact on the model prediction results, avoid the problem of loss of feature correlation information existing in the traditional correlation analysis method, and accelerate the training convergence process, effectively improving the learning efficiency of the model;

[0115] As the length of the input sequence increases, some key information may be lost. Therefore, a temporal attention mechanism is introduced to directly focus on long-distance dependencies without the need to transmit information through multiple time steps, effectively alleviating the problem of loss of key information as the length of the input sequence increases. The FTA-DS-MRGRU soft measurement model uses the temporal attention mechanism after the dual information flow structure to capture temporal dependencies, adaptively assign weights to the received hidden state information, evaluate the importance of different historical time points for the moment to be predicted, and thus select the information at key time points.

[0116] S300 further includes:

[0117] S310, after assigning feature attention weights to the input time series through the feature attention module, weighting the time series according to the feature attention weights, and outputting a weighted time series;

[0118] Specifically, as Figure 4 shown, the expression of the feature attention module is shown in formulas (14) to (17):

[0119] q t1 = W q 1 x t + b q 1 (14)

[0120] e t i = V 1 i σ(W 1 i q t1 + U 1 i x t i + b 1 i ), 1 ≤ i ≤ n (15)

[0121]

[0122] Among them, x t represents the input at time t of the feature attention module, x t i represents the i-th variable at time t in the input time series, q t1 represents the query vector of the feature attention module, e t i represents the feature attention score corresponding to the i-th variable at time t in the input time series, α t i represents the feature attention weight corresponding to the i-th variable at time t in the input time series, represents the input vector after being weighted by the feature attention module, n represents the number of input features, W q 1 、b q 1 、W 1 i 、U 1 i 、b 1 i 、V 1i Both represent the feature attention parameters that are continuously updated and optimized during the model training process.

[0123] S320. Extract the temporal relationship in the weighted time series through the time gate in the time-related flow information extraction structure to obtain a time-related hidden state output, and extract the causal relationship of the weighted time series through the causal gate in the dynamic causal-related flow information extraction structure to obtain a causal-related hidden state output;

[0124] S330. After assigning the first temporal attention weight to the time-related hidden state output through the first temporal attention module and performing weighting, obtain a comprehensive time feature that focuses on historical key points, and after assigning the second temporal attention weight to the causal-related hidden state output through the second temporal attention module and performing weighting, obtain a comprehensive causal feature that focuses on historical key points;

[0125] Specifically, as Figure 5 shown, the expressions of the first temporal attention module and the second temporal attention module are as shown in the following formulas (18) to (21):

[0126] q t2 = W q 2 h t + b q 2 (18)

[0127] g j = V 2 j σ(W 2 j q t2 + U 2 j h t-T+j + b 2 j ), 1 ≤ j ≤ T (19)

[0128]

[0129] Among them, h t represents the input at the t-th moment of the temporal attention module, q t2 represents the query vector of the temporal attention module, g j represents the temporal attention score corresponding to the hidden state at the j-th time step in the input time series, β j represents the temporal attention weight corresponding to the hidden state at the j-th time step in the input time series, c t represents the vector obtained by weighted summation of each hidden state output, T represents the size of the taken time step, W q 2 、b q2 , W 2 j , U 2 j , b 2 j , V 2 j All represent the temporal attention parameters updated and optimized during the model training process.

[0130] S340. After fusing the comprehensive time feature and the comprehensive causal feature through the merging layer, it is output through the fully connected layer to obtain the final output result.

[0131] The embodiment of the present invention adopts a two-stage attention mechanism to extract key information: a feature attention module is introduced in the model feature extraction stage to capture key features and suppress the influence of irrelevant features and noise; a temporal attention module is introduced in the prediction output stage to adaptively focus on the key time point information in historical data and capture long-term sequence dependencies. The attention weights provide an intuitive quantitative explanation of the contribution degree of each input data to the model prediction result, which helps to understand the model decision-making process and enhance the transparency and interpretability of the model.

[0132] S400. Use the training set in S100 to train the FTA-DS-MRGRU soft sensor model through the backpropagation algorithm, and use the test set in S100 to test the FTA-DS-MRGRU soft sensor model to obtain the target FTA-DS-MRGRU soft sensor model;

[0133] S400 specifically includes:

[0134] S410. Forward calculation: Calculate the weights of each gating unit and the two-stage attention mechanism, and its calculation formula is shown in formulas (10) to (21).

[0135] S420. Backward calculation: Calculate its loss function value. The model loss function is the mean squared error (MSE), and the calculation method is shown in formula (22):

[0136]

[0137] Among them, N represents the total number of test set samples, y t represents the true value of the SO 2 concentration in the clean flue gas, represents the predicted value of the SO 2 concentration in the clean flue gas;

[0138] S430. Gradient update: Based on the corresponding error terms, the Adam optimization algorithm is used to update the network weights. The Adam optimization algorithm is a first-order optimization algorithm that can replace the traditional stochastic gradient descent algorithm. It has higher computational efficiency and better convergence performance within the same training cycle, and requires less computational space at the same time.

[0139] S500. Obtain the real-time data of the process variables in the flue gas desulfurization process of the thermal power plant, input it into the target FTA-DS-MRGRU soft sensor model, and obtain the SO 2 content of the clean flue gas after the flue gas desulfurization process of the thermal power plant.

[0140] In summary, based on the existing gating strategy of GRU, the present invention designs two auxiliary memory adjustment gates: a time gate and a causal gate, which are used to finely regulate past information and current information respectively. And further construct a two-stream information extraction structure related to time and causality to generate complementary information flows, significantly enhancing its generalization ability and prediction performance in different industrial scenarios. At the same time, considering that in the modeling of nonlinear dynamic industrial processes, information redundancy and temporal information attenuation problems will weaken the model's ability to capture key information and long-term dependence relationships, a two-stage attention mechanism is embedded in the DS-MRGRU network, and a method for predicting the SO 2 concentration of clean flue gas based on two-stage attention and two-stream memory adjustment GRU is proposed. This method adaptively extracts key feature variables using a feature attention mechanism in the feature extraction stage, suppresses the influence of redundant information, and improves the interpretability of the model; and uses a temporal attention mechanism to mine historical important moment information in the prediction output stage to improve the prediction effect of the model in long time series, realizing the accurate prediction of the key quality variable of the SO 2 content of the clean flue gas after the flue gas desulfurization process of the thermal power plant.

[0141] To verify the effectiveness and superiority of the SO 2 concentration prediction method based on two-stage attention and two-stream memory adjustment GRU provided by the embodiments of the present invention, it is specifically described through experiments. The experimental data comes from the data acquisition system of the desulfurization process of a certain thermal power plant, and the purpose is to perform soft sensor modeling on the SO 2 concentration at the outlet of the secondary absorption tower of this process.

[0142] The flow chart of the flue gas desulfurization process of this thermal power plant is as Figure 7 shown. After researching the desulfurization process of this thermal power plant and preprocessing the data analysis, a candidate input variable set consisting of 18 auxiliary variables, that is, process variables, is finally determined, as shown in Table 1:

[0143] Table 1 Candidate input variables for the flue gas desulfurization process of coal-fired power plants

[0144]

[0145] The FTA-DS-MRGRU soft-sensing model uses a two-stage attention mechanism to dynamically adjust the attention weights, achieving the importance assessment of the spatial and temporal dimensions. By visualizing the attention weight distribution inside the model, the specific contributions of each feature and moment to the model's prediction performance are quantitatively analyzed.

[0146] The proportion of the feature attention weights of the process variables is as Figure 8 shown, which quantifies the importance of each input variable when predicting the net flue gas SO 2 concentration at the outlet of the secondary absorption tower. Among them, the 2nd, 3rd, 11th, and 13th variables are particularly prominent, and the proportion of the attention weights all exceeds 10%, corresponding to the SO 2 concentration of the flue gas at the outlet of the primary absorption tower, the pH value of the gypsum slurry in the primary absorption tower, the net flue gas pressure at the inlet of the furnace chimney, and the intermediate value of the unit's desulfurization efficiency respectively. These variables have significant effects during the desulfurization process: the SO 2 concentration of the flue gas at the outlet of the absorption tower directly affects the desulfurization effect; the pH value of the gypsum slurry in the absorption tower is related to the chemical reaction environment of the gypsum slurry and the desulfurization efficiency; the net flue gas pressure at the inlet of the furnace chimney can reflect the flue gas flow state and the working load of the desulfurization system; the intermediate value of the unit's desulfurization efficiency directly reflects the real-time effect of the desulfurization process. Therefore, these 4 variables have a key impact on the prediction result of the SO 2 concentration and are important considerations for optimizing the desulfurization process and improving the prediction accuracy.

[0147] As Figure 9 shown, the temporal attention weight mechanism effectively quantifies the importance of the hidden states at each historical time point when predicting the SO 2 concentration. The 0 on the abscissa represents the current time point, and 6 corresponds to the time point farthest from the current. It can be clearly observed from the figure that the hidden states closer to the current time point obtain higher attention weights, indicating that they carry more key information and have a greater impact on the prediction result. This result not only verifies the effectiveness of the temporal attention weight mechanism in capturing key information in time series data, but also shows that the sampling data in the recent few moments during the power plant desulfurization process can more accurately reflect the current system's dynamic changes and operating status.

[0148] To further verify the superiority of the proposed model, corresponding ablation experiments are carried out. Figures 10A to 10F 、 Figures 11A to 11F and Table 2 respectively show the prediction fitting curves of each model for the SO 2 concentration, the frequency distribution of the errors, and the performance index data. Specifically, the mean square error (MSE), the mean absolute error (MAE), and the correlation index (R 2) As an evaluation index for the performance of the prediction model, the mean squared error is specifically shown in Formula (22), and the calculation formulas for the mean absolute error (MAE) and the correlation index (R 2 ) are respectively shown in Formulas (23) to (24) as follows:

[0149]

[0150] Among them, N represents the total number of samples in the test set, and y t represents the true value of the net flue gas SO 2 concentration, represents the predicted value of the net flue gas SO 2 concentration, represents the average value of y t .

[0151] Table 2 SO 2 concentration prediction results of different prediction methods

[0152]

[0153] Among them, Figures 11A to 11B σ represents the standard deviation, which reflects the dispersion degree of the prediction error. The larger the value, the higher the stability and reliability of the model. Compared with the traditional GRU network, MRGRU realizes the precise control of information transmission by introducing the time gate and the causal gate, effectively reducing the interference of redundant information, thereby improving the prediction performance of the model. On this basis, DS-MRGRU significantly enhances the feature extraction ability of the model by constructing a dual information flow structure, namely, a parallel time-related flow and a dynamic causal-related flow. FTA-DS-MRGRU further introduces a two-stage attention mechanism, in which the feature attention mechanism assigns higher weights to important auxiliary variables, while the time-series attention mechanism focuses on extracting the hidden state information of important historical time points, significantly improving the prediction accuracy and robustness of the model. The experimental results fully demonstrate the performance advantages of this method.

[0154] Through the soft measurement of the net flue gas SO 2 concentration in the flue gas desulfurization process of a certain power plant, the effectiveness and superiority of the net flue gas SO 2 concentration prediction method provided by the embodiment of the present invention based on two-stage attention and dual-stream memory-regulated GRU compared with other methods are verified, and the prediction accuracy of the model can be effectively improved.

[0155] Some steps in the embodiments of the present invention can be implemented by software, and the corresponding software program can be stored in a readable storage medium, such as an optical disc or a hard disk, etc.

[0156] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. SO2 concentration prediction method based on dual-stage attention and dual-stream memory regulated GRU, characterized in that: include: S100, obtaining historical data of process variables and target variables in the flue gas desulfurization process of a thermal power plant, performing preprocessing, obtaining several time series of the historical data, and dividing the several time series into a training set and a test set; wherein the target variable is the net flue gas SO2 concentration at the outlet of the secondary absorption tower, and the process variables are other parameters in the flue gas desulfurization process of the thermal power plant; S200, constructing a DS-MRGRU network for extracting the temporal relationship and causal relationship between current information and historical information, including a parallel temporal correlation flow information extraction structure and a dynamic causal correlation flow information extraction structure; S300, introducing a feature attention module in the feature extraction stage of the DS-MRGRU network to capture key features and suppress irrelevant features and noise, and introducing a temporal attention module in the prediction output stage of the DS-MRGRU network to capture long time series dependencies, and constructing an FTA-DS-MRGRU soft measurement model with enhanced key information extraction capability; S400, training the FTA-DS-MRGRU soft measurement model using the training set in S100 through a back propagation algorithm, and testing the FTA-DS-MRGRU soft measurement model using the test set in S100 to obtain a target FTA-DS-MRGRU soft measurement model; S500, obtaining real-time data of process variables in the flue gas desulfurization process of the thermal power plant, inputting the target FTA-DS-MRGRU soft sensor model, and obtaining the SO2 content of the clean flue gas after the flue gas desulfurization process of the thermal power plant.

2. The method according to claim 1, characterized in that The time-dependent stream information extraction structure includes a first time-gate-based TGRU layer, a Dropout layer, and a second time-gate-based TGRU layer connected in sequence; The expressions of the first time-gate-based TGRU layer and the second time-gate-based TGRU layer are: in, represents the input of the TGRU layer at time t, represents the time-dependent hidden state output at time t-1 in the TGRU layer input time series, σ(·) represents the activation function sigmoid function, represents the candidate time-dependent hidden state output at time t in the TGRU layer input time series, ⊙ represents the Hadamard product operation, and W T represents the TGRU layer network input weight vector, U T represents the hidden weight vector of the TGRU layer, b T represents the TGRU layer bias vector, represents the update gate corresponding to the TGRU layer, represents the input weight vector of the update gate corresponding to the TGRU layer, represents the hidden weight vector of the update gate corresponding to the TGRU layer, Represents the bias vector of the update gate corresponding to the TGRU layer.

3. The method according to claim 1, characterized in that The dynamic causal correlation flow information extraction structure includes a first CGRU layer based on a causal gate, a Dropout layer, and a second CGRU layer based on a causal gate connected in sequence; The expressions of the first causal gate-based CGRU layer and the second causal gate-based CGRU layer are: in, represents the input of the CGRU layer at time t, represents the causal related hidden state output at time t-1 in the TGRU layer input time series, σ(·) represents the activation function sigmoid function, represents the candidate causal related hidden state output at time t in the CGRU layer input time series, ⊙ represents the Hadamard product operation, W c Represents the CGRU layer network input weight vector, U c represents the hidden weight vector of the CGRU layer, b C represents the CGRU layer bias vector, Represents the update gate corresponding to the CGRU layer, Represents the input weight vector of the update gate corresponding to the CGRU layer, represents the hidden weight vector of the update gate corresponding to the CGRU layer, Represents the bias vector of the update gate corresponding to the CGRU layer.

4. The method according to claim 1, characterized in that: The S300 includes: S310, after assigning feature attention weights to the input time series through the feature attention module, weighting the time series according to the feature attention weights, and outputting a weighted time series; S320, extracting the temporal relationship in the weighted time series through the time gate in the time-related flow information extraction structure to obtain a time-related hidden state output, and extracting the causal relationship of the weighted time series through the causal gate in the dynamic causal-related flow information extraction structure to obtain a causal-related hidden state output; S330, assigning a first temporal attention weight to the time-related hidden state output through a first temporal attention module and weighting the weights to obtain a comprehensive time feature of the historical key point, and assigning a second temporal attention weight to the causal-related hidden state output through a second temporal attention module and weighting the weights to obtain a comprehensive causal feature of the historical key point; S340, the comprehensive time feature and the comprehensive causal feature are fused through a merging layer and then output through a fully connected layer to obtain a final output result.

5. The method according to claim 4, characterized in that The expression of the feature attention module is: q t1 =W q 1 x t +b q 1 e i t =V1 i σ(W1 i qt1+U1 i x t i +b1 i ),1≤i≤n Among them, x t represents the input of the feature attention module at time t, x t i represents the i-th variable at time t in the input time series, q t1 represents the feature attention module query vector, e t i represents the feature attention score corresponding to the i-th variable at time t in the input time series, α t i represents the feature attention weight corresponding to the i-th variable at time t in the input time series, represents the input vector after weighting by the feature attention module, n represents the number of input features, and W q 1 、b q 1 、W1 i 、U1 i 、b1 i 、V1 i Both indicate that the model training process continuously updates the optimized feature attention parameters.

6. The method according to claim 5, characterized in that The expressions of the first temporal attention module and the second attention module are: q t2 =W q 2 h t +b q 2 g j =V2 j σ(W2 j q t2 +U2 j h t-T+j +b2 j ),1≤j≤T Among them, h t represents the input of the temporal attention module at time t, q t2 represents the query vector of the temporal attention module, g j represents the temporal attention score corresponding to the hidden state of the jth time step in the input time series, β j represents the temporal attention weight corresponding to the hidden state of the jth time step in the input time series, c t represents the weighted sum of the outputs of each hidden state, T represents the time step size, and W q 2 、b q 2 、W2 j 、U2 j 、b2 j 、V2 j Both represent the updated and optimized temporal attention parameters during model training.

7. The method according to claim 1, characterized in that The S100 includes: S110, after obtaining historical data of process variables and target variables in the flue gas desulfurization process of the thermal power plant, perform missing value processing, outlier processing and data standardization in sequence; S120, using a sliding window technology to convert the standardized historical data into a number of time series with a length of T, and dividing the time series into a training set and a test set.

8. The method according to claim 1, characterized in that The step of training the FTA-DS-MRGRU soft sensor model by using the training set in S100 through a back propagation algorithm in S400 includes: Performing forward calculation on the FTA-DS-MRGRU soft measurement model using the training set in S100 to calculate the network weight of the FTA-DS-MRGRU soft measurement model; The corresponding error term is obtained by calculating the loss function value of the FTA-DS-MRGRU soft sensor model through the semi-mean square error loss function, and the calculation formula is: Among them, N represents the total number of test set samples, y t Indicates the true value of SO2 concentration in clean flue gas, Indicates the predicted value of SO2 concentration in the clean flue gas; Based on the corresponding error term, the Adam optimization algorithm is used to perform gradient update on the network weights.

9. A target detection method, characterized in that: The method adopts the method described in any one of claims 1 to 8 to detect the gas to be detected, and the gas to be detected is a gas generated in an industrial process.

10. A computer device, characterized in that: The computer device includes a processor and a memory, the memory is used to store computer instructions, and the computer instructions are loaded and executed by the processor to implement the SO2 concentration prediction method based on dual-stage attention and dual-stream memory regulated GRU as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Air quality prediction method based on variational recursive network and self-attention mechanism

    CN113095550A

  • Short-term load prediction model and method based on double attention mechanism and LSTM

    CN113902202A

  • Method for predicting concentration of purified flue gas SO2 in flue gas desulfurization process of coal-fired power plant

    CN117313936A

  • Public building air conditioner load decomposition analysis method and device based on two-stage attention mechanism fused convolutional neural network and long short-term memory network, and electronic equipment

    CN119026291A

  • Hard disk fault prediction method and system based on improved GRU and storage medium

    CN119201565A