Short-term Air Quality Prediction Method Based on Sparrow Search Algorithm and Decomposition Error Correction

By using the sparrow search algorithm and decomposition error correction method in air quality prediction, combined with the CEEMDAN decomposition algorithm and deep learning model, the problem of insufficient prediction error and parameter simplification in the existing technology is solved, and more efficient and accurate air quality prediction is achieved.

CN115660167BActive Publication Date: 2025-06-24NANCHANG INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211308007.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-25
Publication Date
2025-06-24
Estimated Expiration
2042-10-25

AI Technical Summary

Technical Problem

The existing air quality prediction methods have shortcomings in prediction errors and parameter simplification, resulting in low prediction accuracy and long operation time.

Method used

The short-term air quality prediction method based on sparrow search algorithm and decomposition error correction is adopted. The data is decomposed through the CEEMDAN decomposition algorithm, combined with the deep learning model to predict, and the prediction errors of multiple deep learning models are processed using the simple average method, MAE reciprocal method and Lagrangian multiplier method to obtain the weight of each model, and the prediction results of each model are combined.

Benefits of technology

It improves the accuracy and efficiency of air quality prediction, reduces prediction error and calculation time, and achieves more efficient air quality prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115660167B_ABST
    Figure CN115660167B_ABST
Patent Text Reader

Abstract

The present invention proposes a short-term air quality prediction method based on the sparrow search algorithm and decomposition error correction, comprising the following steps: S1, decomposing the data using the CEEMDAN decomposition algorithm to obtain a number of IMFs; S2, using a deep learning model to predict the number of IMFs, processing the prediction errors of multiple deep learning models to obtain the weights of each deep learning model; the number of deep learning models is at least two; S3, combining the prediction results of each deep learning model according to the weights to obtain a comprehensive prediction result, that is, the result of multi-model combined prediction. By combining air quality prediction with time series decomposition and adopting a combined prediction method, the present invention combines the prediction results of each model using weights, and can obtain a higher-precision air quality prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of air pollution prediction, and particularly to a short-term air quality prediction method based on the sparrow search algorithm and decomposition error correction. Background Art

[0002] Long-term exposure of humans to polluted air can lead to health problems. According to the research of the World Health Organization (WHO), the global disease burden associated with exposure to air pollution has caused huge losses to human health worldwide. It is estimated that the disease burden caused by air pollution is comparable to other major global health risks such as unhealthy diet and smoking, and air pollution is now considered the greatest environmental threat to human health. Therefore, predicting future air quality based on past air quality data has great research significance. It is not only reflected in the early warning of air pollution, but also in better planning and decision-making for the development direction of cities and better protection of human physical health.

[0003] Currently, the methods for air quality prediction mainly focus on the prediction accuracy of future air quality data, while ignoring prediction errors and parameter simplification, which can lead to low prediction accuracy and long operation time. Summary of the Invention

[0004] The present invention aims to at least solve the technical problems existing in the prior art, and particularly innovatively proposes a short-term air quality prediction method based on the sparrow search algorithm and decomposition error correction.

[0005] To achieve the above object of the present invention, the present invention provides a short-term air quality prediction method based on the sparrow search algorithm and decomposition error correction, including the following steps:

[0006] S1, decomposing the data using the CEEMDAN decomposition algorithm to obtain a number of IMFs;

[0007] S2, using a deep learning model to predict the number of IMFs, processing the prediction errors of multiple deep learning models to obtain the weights of each deep learning model; the number of deep learning models is at least two;

[0008] S3, combining the prediction results of each deep learning model according to the weights to obtain a comprehensive prediction result, that is, the result of multi-model combined prediction.

[0009] Further, the deep learning model includes: LSTM, Bi-LSTM, GRU, and Bi-GRU.

[0010] Further, S2 includes the following steps:

[0011] S2-1. Use different deep learning models to predict the several IMFs, and obtain multiple prediction results; meanwhile, obtain the predicted data and the values of statistical evaluation indicators.

[0012] S2-2. Use any one of the simple average method, the reciprocal of MAE method, and the Lagrange multiplier method to process the prediction errors of each model, and obtain the weights of each deep learning model.

[0013] When using the simple average method, the weights of each model are equal.

[0014] When using the reciprocal of MAE method, the weights of each model are expressed as follows:

[0015]

[0016] Where S represents the sum of the reciprocals of the MAEs of the prediction results of H models, and H is the total number of models.

[0017] a i represents the weight of the i-th model.

[0018] MAE i represents the mean absolute error of the i-th model.

[0019] When using the Lagrange multiplier method, the weights of each model are obtained by solving the following formula:

[0020]

[0021] Where a i represents the weight of the i-th model, i = (1, 2,..., N);

[0022] N is the total number of models;

[0023] represents taking the partial derivative of a in the Lagrangian function L(a, λ).

[0024] S2-3. Finally, combine the prediction results of each deep learning model with the weights to generate the results of multi-model combined prediction.

[0025] Furthermore, before step S1, preprocess the data, including the following steps:

[0026] S0-1. Use cubic spline interpolation to fill in the missing values;

[0027] S0-2. Use the moving average method to process the data after cubic spline interpolation.

[0028] Preprocessing the data can not only handle the anomalies and missing data in the original data, but also eliminate the influence of periodic and seasonal factors in the original data, thereby improving the accuracy of prediction.

[0029] Furthermore, the deep learning model is obtained by training with the sparrow search algorithm, including the following steps:

[0030] SA, Normalize the data decomposed by CEEMDAN, and then classify it into a uniform distribution of statistical probabilities between 0 and 1; The normalization operation not only speeds up the convergence rate of the neural network, but also eliminates the adverse effects caused by odd sample data.

[0031] SB, Divide the classified data into a training data set and a test data set;

[0032] SC, Use the sparrow search algorithm SSA to optimize the hyperparameters of the deep learning model, and the hyperparameters include the number of neurons in each layer of the neural network, the number of iterations, and the learning rate;

[0033] SD, Use the deep learning model to predict the test data set and evaluate it according to the evaluation metrics. If the hyperparameters of the model do not overfit and the error of the prediction result is the smallest, the optimal hyperparameters are obtained, and thus the optimal deep learning model. If the hyperparameters of the model overfit or do not meet the minimum error of the prediction result, jump to step SC.

[0034] The sparrow search algorithm enables the model to adaptively determine the parameters of the deep learning model, so that manual setting is not required, thus realizing the simplification of parameters.

[0035] Furthermore, the evaluation metrics include any combination of mean absolute error, root mean square error, correlation coefficient, and mean absolute percentage error.

[0036] Furthermore, before step S1, there is S0 to obtain data, and this data is meteorological data.

[0037] Furthermore, the data is short-term meteorological data, and the collection time is 24 to 72 consecutive hours.

[0038] In summary, due to the adoption of the above technical solutions, the present invention combines air quality prediction with time series decomposition, and adopts a combined prediction method, using weights to combine the prediction results of each model, thereby obtaining a higher-precision air quality prediction result.

[0039] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. Brief Description of the Drawings

[0040] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where:

[0041] Figure 1 is a schematic diagram of the overall structure of the present invention.

[0042] Figure 2 is a schematic diagram of the network structures of LSTM, Bi-LSTM, GRU, and Bi-GRU. Figure 2 (a) is a structural diagram of the LSTM network. Figure 2 (b) is a structural diagram of the BiLSTM network. Figure 2 (c) is a structural diagram of the GRU network. Figure 2 (d) is the structure of the Bi-GRU network.

[0043] Figure 3 is a schematic diagram of the air quality index data of Beijing and the locations of monitoring stations in Beijing in a specific embodiment. Figure 3 (a) is the air quality index data of Beijing. Figure 3 (b) is the locations of the monitoring stations in Beijing.

[0044] Figure 4 is a schematic diagram of the results of data preprocessing of the present invention and a multi-step deep learning model for processing time series. Figure 4 (a) is a partial result diagram of cubic spline interpolation. Figure 4 (b) is a result diagram of the moving average method. Figure 4 (c) is a diagram of a multi-step deep learning model for processing time series. Figure 4 (d) is a local diagram from Day 0 to Day 120 after using the moving average method.

[0045] Figure 5 is the decomposition result and normalization result of CEEMDAN in a specific embodiment. Figure 5 (a) is the decomposition result of CEEMDAN. Figure 5 (b) is the result after normalizing the decomposition result of CEEMDAN.

[0046] Figure 6 is a schematic diagram of the results of whether to use the decomposition algorithm on different deep neural models of the present invention. Figure 6 (a), Figure 6 (b) is a schematic diagram of the prediction results of the model before and after using the decomposition algorithm. Figure 6 (c) is a schematic diagram of the error of the prediction results.

[0047] Figure 7 is a schematic diagram of the changes in the statistical indicators and prediction results of each model.

[0048] Figure 8 Schematic diagram of prediction results of different method combinations Figure 8 (a), Figure 8 (b) are graphs of prediction results of different method combinations Figure 8 (c) is a statistical graph of prediction errors of different method combinations

[0049] Figure 9 Schematic diagram of changes in model statistical indicators using different combination methods Figure 9 (a) is a schematic diagram of changes in statistical indicators of the model that combines "bg - g - bl - l" Figure 9 (b) is a schematic diagram of changes in statistical indicators of the model that combines into "imf" Figure 9 (c) is a schematic diagram of changes in statistical indicators of the model that combines into "ceemdan" Specific implementation manners

[0050] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention

[0051] The present invention proposes a short - term air quality prediction method based on the sparrow search algorithm and decomposition error correction, as Figure 1 shown, which includes the following steps

[0052] S100: Data processing

[0053] S100 - 1. Pre - processing the data can not only handle the abnormal and missing data in the original data, but also eliminate the influence of periodic and seasonal factors in the original data, thereby improving the prediction accuracy. The pre - processing includes using cubic spline interpolation to fill in the missing values; adopting the moving average method to process the data after cubic spline interpolation

[0054] S100 - 2. Using the decomposition algorithm to decompose the pre - processed data into a finite number of intrinsic mode functions IMF. The decomposed components are respectively used as the data for the subsequent parts. The decomposition algorithm is CEEMDAN with adaptive noise

[0055] S200: Training the model and predicting the data

[0056] S200 - 1. Normalize the IMF. After normalization, the IMF is distributed between 0 and 1. The normalization operation not only speeds up the convergence rate of the neural network, but also eliminates the adverse effects caused by odd - numbered sample data

[0057] S200-2, the normalized IMFs are divided into a training dataset and a test dataset. The training data accounts for 80% of the data. Five-fold cross-validation is used on the training set to obtain a reliable and stable model.

[0058] S200-3, Use the Sparrow Search Algorithm (SSA) and the training dataset to train deep learning models (i.e., neural networks LSTM, Bi-LSTM, GRU, and Bi-GRU). The Sparrow Search Algorithm (SSA) is used to optimize the hyperparameters of the neural network, including the number of neurons in each layer, the number of iterations, and the learning rate. These hyperparameters are used to obtain the current optimal neural network. Subsequently, this neural network is used to predict the test dataset and is evaluated according to the evaluation metrics. The hyperparameters of the model with no overfitting and the smallest prediction error are the optimal hyperparameters.

[0059] S200-4, Construct each neural network with the optimal hyperparameters and use the neural network to predict the dataset. At the same time, obtain the predicted data and the values of the statistical evaluation metrics.

[0060] S200-5, Use the simple average method, the reciprocal of MAE method, and the Lagrange multiplier method to correct the errors of multiple models (i.e., error correction) and obtain the weights of each deep learning model.

[0061] S300, Calculate the combined prediction results of multiple models according to the weights. Compare and analyze the combined prediction results with the prediction results without combination.

[0062] The four deep learning models are namely LSTM, Bi-LSTM, GRU, and Bi-GRU.

[0063] The Long Short-Term Memory Network (LSTM) is an improved Recurrent Neural Network (RNN). It is a progressive RNN, mainly designed to solve the problems of gradient disappearance and gradient explosion during the training of long sequences, and can be used for time series modeling and prediction. The structure of the LSTM unit is as follows Figure 2 (as shown in (a)). The LSTM consists of a state storage unit C t and three logic gates (input gate i t , output gate o t , forget gate f t ), and in addition, x t is the input of the LSTM, and h t is the output of the LSTM. Among them, represents the element-wise multiplication of matrices, represents the addition of matrices.

[0064] f t = σ(W f[h t-1 ,x t +b f ) (1)

[0065] i t =σ(W i [h t-1 ,x t +b i ) (2)

[0066] C t ′=tanh(W c [h t-1 ,x t +b c ) (3)

[0067] C t =f t *C t-1 +i t *C t ′ (4)

[0068] o t =σ(W o [h t-1 ,x t +b o ) (5)

[0069] h t =o t *tanh(C t )(6) In equations (1)-(6), W is the weight matrix corresponding to each module, W f is the weight matrix of the forget gate, W i is the weight matrix of the input gate, W o is the weight matrix of the output gate, W c is the weight matrix of the tanh module. b represents the bias error corresponding to each module, b f is the bias error of the forget gate, b i is the bias error of the input gate, b o is the bias error of the output gate, b c is the bias error of the tanh module. [h t-1 ,x t represents the concatenation of h t-1 ,x t , h t-1 represents the output of the LSTM at time t-1, x t represents the input of the LSTM at time t. σ is the sigmoid activation function, and tanh is the hyperbolic tangent function.

[0070]

[0071] The state storage unit C is used in the LSTM t which enables more effective prediction of time series. At time t, the LSTM obtains the output h of the previous unit t-1 and the new input x t When this happens, first the forget gate f t controls whether to retain the state information C of the previous unit t-1 ; the input gate i t determines which contents in h t-1 and x t can be used for the calculation of this unit state, reducing irrelevant contents; the tanh module generates a new state C t-1 ' through h t and x t ; then the state information is updated, and these results are combined to obtain the current state information C t ; finally, the output gate o t and the current state information C t generate the output h t .

[0072] The bidirectional long short-term memory model (Bi-LSTM) is a special LSTM that consists of a two-layer network with a forward LSTM and a backward LSTM. Bi-LSTM can consider both past and future information of the data simultaneously. Figure 2 (b) shows the structure of the Bi-LSTM network. Bi-LSTM can obtain two hidden states in reverse and connect these two states to obtain the same output. The forward LSTM and backward LSTM layers can regain the past information and future information of the input sequence. Bi-LSTM can achieve better prediction results than LSTM.

[0073] The gated recurrent unit (GRU) is a recurrent neural network (RNN) similar to LSTM. It uses a gating mechanism to control input, memory, and other information for prediction at the current time step, which solves problems such as long-term memory and gradients in backpropagation. At the same time, the biggest difference between GRU and LSTM is that it has a simpler structure with only two gates (an update gate for updating the state and a reset gate for resetting the state), and there is no unit state C in GRU like in LSTM t . Generally speaking, the update gate in GRU controls whether past information can be passed backward, while the reset gate controls the degree to which past information is forgotten. Figure 2 (c) shows the structure of the GRU network.

[0074] Equations (9) to (12) represent the state of the GRU cell at time t.

[0075] r t = σ(W r [h t-1 , x t + b r ) (9)

[0076] z t = σ(W z [h t-1 , x t + b z ) (10)

[0077]

[0078] where r t represents the reset gate, z t represents the update gate, x t represents the input at time t, h t represents the hidden state at time t, represents the partial state information at time t, h t represents the state information at time t, W represents the weight, W r represents the reset gate weight, W z represents the update gate weight, W h represents the weight used to calculate the current state information in the tanh module; b represents the bias term of the input gate, b r represents the bias term of the reset gate, b z represents the bias term of the update gate, b h represents the bias term in the tanh module; σ(·) represents the sigmoid activation function widely used in neural network operations; tanh(·) represents the tanh activation function. Both the sigmoid and tanh functions can transform data values into a certain range, and σ(x) and tanh(x) can be expressed as Equations (7) and (8).

[0079] The Bidirectional gate recurrent unit (Bi-GRU) is a special type of GRU, whose structure is similar to that of Bi-LSTM and consists of GRU networks for forward and backward propagation. Figure 2 (d) shows the structure of the Bi-GRU network.

[0080] The decomposition algorithm is preferably the Complete Ensemble Empirical Mode Decomposition with Adaptive Noise (CEEMDAN). The EMD algorithm or the EEMD algorithm can also be used. The EMD algorithm, i.e., the empirical mode decomposition algorithm, gradually decomposes a non-stationary signal into several intrinsic mode functions (IMFs) and a residue (Res). EMD has obvious advantages in dealing with non-stationary and non-linear data and is suitable for analyzing non-linear and non-stationary signal sequences. However, EMD is affected by the end effect and mode mixing. Therefore, EEMD was proposed. EEMD adds white Gaussian noise to the original sequence to make the distribution of signal poles more uniform and effectively suppresses mode mixing. However, the computational complexity of the EEMD algorithm is very high, and the amplitude and number of iterations of adding white noise in EEMD need to be set according to experience. Therefore, Torres et al. proposed the Complete Ensemble Empirical Mode Decomposition with Adaptive Noise CEEMDAN. CEEMDAN not only solves the mode mixing problem in EMD, but also reduces the computational cost and improves the execution efficiency by adaptively adding white noise and setting the number of iterations.

[0081] The CEEMDAN algorithm includes: assuming that y(t) is the original signal; v j (t) is a Gaussian white noise signal that satisfies the standard normal distribution; j is the number of times white noise is added, j = 1, 2,..., N; E i (·) is the function that decomposes the i-th IMF using EMD; I i (t) is the i-th component obtained by EMD decomposition; is the i-th component obtained by CEEMDAN decomposition; ε is the amplitude of the white noise. The steps of CEEMDAN decomposition are as follows:

[0082] The first step: Add Gaussian white noise to the original signal y(t) to obtain a new signal After that, use EMD to decompose the new signal to obtain the first-order component I1(t). Where res represents the residue; j is the number of decompositions.

[0083]

[0084] Where represents the new signal obtained by adding white noise j times;

[0085] represents the signal obtained by decomposing the signal after adding white noise j times using the EMD algorithm It represents the first-order component obtained by using EMD decomposition after adding white noise j times.

[0086] res j It represents the residual obtained by using EMD decomposition after adding white noise j times.

[0087] Step 2: The overall average of N components produces the first-order component of CEEMDAN decomposition After removing the first-order component of the decomposition, the residual signal r1(t) is obtained.

[0088]

[0089] where j is the number of times white noise is added, j = 1, 2,..., N;

[0090] Step 3: Gaussian white noise is added to the residual signal r1(t) to obtain a new residual signal After that, the new residual signal is decomposed using EMD to obtain the second-order component I2(t).

[0091]

[0092] where ε is the amplitude of the white noise;

[0093] E1[v j (t)] represents the Gaussian white noise signal v added j times decomposed using EMD j (t);

[0094] Step 4: Repeat Step 2 to obtain the second-order component of CEEMDAN and the residual r2(t) after decomposition.

[0095] Step 5: Repeat the above steps until the obtained residual signal r k (t) is a monotonic function that cannot be further decomposed, then the algorithm ends. At this time, the decomposition result of CEEMDAN The k-th order component in can be expressed by the following formula.

[0096]

[0097] where j is the number of times white noise is added, j = 1, 2,..., N;

[0098] At this time, the residual signal of CEEMDAN can be expressed as:

[0099]

[0100] The result of CEEMDAN decomposition can be expressed as:

[0101]

[0102] The Sparrow Search Algorithm (SSA) is a swarm intelligence optimization model inspired by the swarm intelligence, foraging, and anti-predation behaviors of sparrows. It was proposed by Xue J and Shen B in 2020. SSA describes two types of sparrow behaviors in a sparrow population. One type of sparrow is called the producer, and its main task is to find food, so it tends to fly to places with abundant food. Another type of sparrow is called the scrounger, which continuously selects and follows the best producers to find food.

[0103] When the producer (scrounger) discovers a predator while foraging, it will emit an alarm signal. When the alarm signal exceeds a certain threshold, the position of the producer will be updated as follows:

[0104]

[0105] Where, represents the position value of the i-th (i = 1, 2,..., n, where n represents the number of sparrows) sparrow at the t-th (t = 1, 2,..., P, where P represents the maximum number of iterations) iteration of the j-th (j = 1, 2,..., M, where M represents the number of parameters to be optimized) parameter to be optimized. α and Q are random numbers. The difference is that Q follows a normal distribution, while the value of α is between 0 and 1. W and S represent the alarm value and the safety threshold respectively. C represents a column matrix with M rows, and the value of each row is 1.

[0106] The position of the scrounger is related to the producer, and its position calculation method is as follows:

[0107]

[0108] Where, represents the best position among the sparrows of the producer at the (t + 1)-th iteration; represents the worst position among the sparrows of the producer at the t-th iteration, and |·| is the absolute value symbol. A′ = A T (AA T ) -1 , and A represents a column matrix with M rows, and the value of each row is 1 or -1.

[0109] The sparrows on the edge, that is, the sparrows with uncertain identities on the periphery of all sparrows, move towards the middle when encountering danger, and the position update at this time is as follows:

[0110]

[0111] Among them represents the global optimal position of the producer among the sparrows at the t-th iteration; β and K are random numbers. The difference is that β follows a normal distribution with a mean of 0 and a standard deviation of 1, while K controls the direction and step size of the specific movement of the sparrow, and its value range is between -1 and 1; λ is a very small number to ensure that the denominator is not zero; F i , F g and F w represent the values of the i-th sparrow, the best sparrow, and the worst sparrow, respectively. The best position of the sparrow at the end of the algorithm iteration represents the best value of the parameter to be searched and queried.

[0112] The S200-5 includes: using different neural networks to separately predict several decomposed IMF rows and input them into the network as a single IMF; to obtain multiple prediction results. Then, the prediction errors of multiple neural networks are processed to obtain the weights of each prediction model. Finally, the prediction results of each neural network are combined with the weights to generate the result of multi-model combined prediction.

[0113] In formulas (26) and (27), Y t represents the combined prediction result of the multi-model at time t; represents the prediction result of the H-th model at time t; a i represents the weight of the i-th model, and ξ represents the error of the multi-model combination.

[0114]

[0115] Among them, F(·) represents combining the prediction results of the multi-model in a certain way.

[0116] If the time range of the prediction result is t′~t′+T, then the combined prediction result Y t of the multi-model and the true value y t The mean-square error (MSE) between them can be expressed as:

[0117]

[0118] y t represents the true value at time t; represents the combined prediction result of the multi-model; y represents the true value;

[0119] In this article, the following three combined prediction methods are considered:

[0120] (1) Simple average combination method

[0121] The final multi-model combined prediction is the average of the predictions of multiple models, where ξ represents the error of the multi-model combination.

[0122]

[0123] represents the predicted value of the i-th model at time t;

[0124] ξ represents the error of the multi-model combination;

[0125] (2) Reciprocal combination method of MAE

[0126] In the multi-model combined prediction, the weight of each model can be determined by the reciprocal of the mean absolute error (MAE). The formula for calculating the weight is as follows:

[0127]

[0128] where S represents the sum of the reciprocals of the mean absolute errors (MAE) of the prediction results of H models;

[0129] a i represents the weight of the i-th model,

[0130] The mean absolute error of the i-th model is represented by MAE i The formula (27) shows the calculation method of the final multi-model comprehensive prediction result.

[0131] (3) Combination method based on Lagrange multiplier method

[0132] The Lagrange multiplier method is a method for finding the extreme value of a multi-variable function under a set of constraints. In this part, the weight a i of the i-th model is regarded as the variable to be solved again, and the constraint condition is as shown in formula (32).

[0133]

[0134] With the introduction of error, this can be expressed as:

[0135]

[0136] where, represents the minimum value of the mean square error (MSE) between the combined prediction result of the multi-model and the true value, see formula (28);

[0137] is the predicted value of the i-th model; y represents the true value; err i represents the error of the i-th model.

[0138] After introducing the Lagrange multipliers, the constrained optimization problem is transformed into an unconstrained optimization problem.

[0139] L(a,λ) = E[(a1err1 + a2err2 +... + a N err N ) 2 + λ(a1 + a2 +... + a N - 1) (35)

[0140] where L(a,λ) represents the Lagrangian function; a is the quantity to be determined, and λ is the Lagrange multiplier;

[0141] E[] represents the expectation;

[0142] The Karush - Kuhn - Tucker (KKT) conditions are introduced:

[0143]

[0144] The combined weights can be solved by formula (36).

[0145] The specific embodiments are as follows:

[0146] First, data is collected: A total of 8,695 hours of Air Quality Index (AQI) data from January 1, 2021 to December 31, 2021 in Beijing area is used ( Figure 3 (a) is the AQI data of Beijing area; Table 1 lists the statistical characteristics of the AQI data), and these data are measured by 24 air monitoring stations in Beijing ( Figure 3 (b) shows the locations of monitoring stations such as Dingling, Dongsi, Temple of Heaven, Wanliu in Haidian, Huairou Town, and Changping Town), and are published by the China National Environmental Monitoring Center (GEMS) on the National Urban Air Quality Real - time Release Platform (https: / / air.cnemc.cn:18007 / ).

[0147] As shown in Table 1, there are 65 missing values in the AQI data of Beijing area, and the maximum value in the data is 500, reaching the limit in the AQI definition. To reduce the impact of missing values and outliers in the data on subsequent predictions, it is necessary to pre - process the data to improve the prediction accuracy.

[0148] The Air Quality Index (AQI) is a quantitative description of air quality, which can illustrate the cleanliness or pollution degree of the air and the impact of air pollution on health. Its calculation formula is as follows:

[0149] AQI = max{IAQI1, IAQI2,…, IAQI n} (37)

[0150]

[0151] Among them, n is the pollutant item; IAQI p is the IAQI of pollutant item P. The air quality index (IAQI) of a single pollutant represents the air quality index of the pollutant. In formula (38), C p represents the mass concentration value of pollutant item P; BP H and BP L represent the upper limit and the lower limit of the mass concentration of pollutant item P respectively; IAQI H and IAQI L represent the air quality sub - indices corresponding to BP H and BP L respectively.

[0152] Then pre - process the data: Use cubic spline interpolation to fill in the missing values. Figure 4 (a) shows partial results of cubic spline interpolation (from 00:00 on January 1, 2021 to 23:00 on January 2, 2021).

[0153] After that, use the moving average method to process the data after cubic spline interpolation. It can not only eliminate the irregular variations and other variations in the time series, but also reveal the long - term trend of the time series and improve the prediction accuracy. The calculation formula for the moving average of the time series y=(y1, y2, …, y T ) is:

[0154]

[0155] where G represents the number of moving average terms, and G < T. Let y1 be the AQI value on Day 0, then the value on the last day is Day 8736.

[0156] Figure 4 (b) and Figure 4 (d) show the results of using the moving average method when G = 24. Table 1 shows the statistical characteristics of the results after the above operations on the original data.

[0157] Table 1 Statistical description of the data

[0158]

[0159] In addition, the data needs to be divided before training the model. The training set accounts for 80%, and the test set accounts for 20%. In Figure 4 (c), x t represents the AQI value at time t and is the input value; y t+1Denotes the AQI value at time t+1 (the next time), which is the output value; T represents the period, and T = 24 means predicting the future AQI value using the AQI values of 24 consecutive hours. Let the first day of the training dataset be Day 0, and the last day of the test dataset be Day 8712.

[0160] Then, evaluate the prediction results of the model: To determine which model has a better prediction effect, the prediction effect of the proposed model is evaluated using the following evaluation metrics: mean absolute error (MAE), root mean squared error (RMSE), coefficient of determination (R 2 ), and mean absolute percentage error (MAPE). Their definitions are as follows:

[0161]

[0162]

[0163] Among them, is the predicted value; y = (y1, y2,..., y N ) is the true value.

[0164] The mean absolute error (MAE) represents the average of the absolute errors between the predicted value and the observed value, and its range is [0, +∞). The smaller the MAE, the better the prediction effect of the model.

[0165] The root mean squared error (RMSE) is a typical indicator of the regression model, used to show the difference between the actual data and the predicted data of the model, and its range is [0, +∞). The smaller the RMSE, the better the prediction effect of the model.

[0166] R 2 refers to the coefficient of determination, which can be used to measure the fitting degree of the regression problem. The closer R 2 is to 1, the better the prediction effect of the model.

[0167] The mean absolute percentage error (MAPE) is a relative indicator. If MAPE = 0%, the proposed model is perfect; if MAPE > 100%, it means that the prediction effect of the model is very poor.

[0168] Meanwhile, if the prediction effect of the model is better, the parameters of the model are also more excellent.

[0169] Finally, the experimental results were analyzed and compared: The experiment consisted of two parts. In the first part of the experiment, the LSTM, Bi-LSTM, GRU, and Bi-GRU models were first used to predict the AQI value. Then, the data decomposed by the CEEMDAN decomposition algorithm was used as the training data, and the above models were used for prediction in turn. Finally, their prediction results were compared. In the second part of the experiment, the combined weights were obtained by the simple average method, the reciprocal of MAE method, and the Lagrange multiplier method respectively. Secondly, the prediction results of the four deep learning models were combined according to the weights. Finally, the prediction results of the three combinations were compared respectively, and these prediction results were compared with the results of the first part.

[0170] (1) Decomposition results of CEEMDAN

[0171] Figure 5 (a) shows the decomposition results of CEEMDAN, Figure 5 (b) shows the normalization results of the CEEMDAN decomposition components. In the following text, the normalized CEEMDAN decomposition components (IMFs) are used as data.

[0172] The subsequent experimental process was similar to that without using the CEEMDAN decomposition algorithm. The above four deep learning models were trained respectively using the training dataset combined with the SSA optimization algorithm to obtain the optimal deep learning network models. The prediction results combining the CEEMDAN decomposition algorithm, the SSA optimization algorithm, and the deep learning models are shown in Table 2 and Figure 6 .

[0173] Table 2 Prediction results of GRU, Bi-GRU, LSTM, and Bi-LSTM on different evaluation metrics

[0174]

[0175] Figure 6 (a) shows the prediction results of the four best deep learning models with and without using the CEEMDAN decomposition algorithm, Figure 6 (b) shows the scatter plot of the re - sults. Figure 6 (c) The error symbols indicate the magnitude of the predicted data relative to the original data. When the error is positive, it means the prediction result is greater than the original data, and vice versa, it means the prediction result is less than the original data.

[0176] From Figure 6It can be seen that the prediction results of SSA-GRU and SSA-Bi-GRU are distributed on both sides of the original data. At the same time, the prediction error of SSA-Bi-GRU is smaller than that of SSA-GRU. The prediction results of SSA-LSTM and SSA-Bi-LSTM are mostly smaller than the original data. The CEEMDAN decomposition algorithm also greatly reduces the predicted values. It makes most of the predicted values smaller than the original data. However, the results of SSA-Bi-GRU-CEEMDAN are concentrated around the original data. It can be inferred that SSA-Bi-GRU-CEEMDAN has better prediction results. Figure 6 (c) proves this inference. Most of the errors of SSA-Bi-GRU-CEEMDAN are distributed around 0.

[0177] Figure 7 Shows the evolution of the statistical evaluation indicators of the prediction results of the above models. Combining Table 2 and Figure 7 It can be concluded that after decomposing the data using the CEEMDAN decomposition algorithm, the prediction results of the models are improved. The RMSE is reduced by 11.245% on average, the R2 is increased by 0.866% on average, and the MAE is reduced by 8.165% on average. In particular, after using the CEEMDAN decomposition model, the SSA-Bi-GRU model achieves the best prediction results. Its RMSE and MAE are the smallest, the R2 is the largest among all models, and its RMSE and MAE are 4.810319 and 3.657710 respectively, a reduction of 14.356206% and 12.123751%. Its R2 increases from 0.978011 to 0.983518, and its MAEP remains around 0.06%.

[0178] (2) Error correction and weight acquisition

[0179] In the experiment, the combined weights of multiple models are re-obtained through the integration of the simple average method, the reciprocal of MAE method, and the Lagrange multiplier method.

[0180] Figure 8 Shows the prediction results of the multi-model combination. Each figure shows the results of the "bg-g-bl-l" method, the "imf" method, and the "ceemdan" method from left to right. "bg-g-bl-l" in the legend represents the result of combining four deep learning models without CEEMDAN decomposition, as shown in formula (47). "imf" in the legend means that the IMF components predicted by four deep learning networks are first combined and then merged into the final prediction, as shown in formulas (48) and (49). "ceemdan" in the legend means that the IMF components predicted by each deep learning network are first combined and then merged into the final prediction, as shown in formula (50).

[0181] Wbg_g_bl_l = [w bg , w g , w bl , w l (44)

[0182]

[0183] W ceemdan = [w bg , w g , w bl , w l (46)

[0184]

[0185]

[0186] Where Y represents the combined prediction result of multiple models; Y i represents the combined prediction result of the i-th (i = 1, 2,..., n) IMF component; y i,j represents the result predicted by the i-th IMF component through model j (j = bi_gru, gru, bi_lstm, lstm); y j represents the prediction result of the j-th model; W bg_g_bl_l , W imf , and W ceemdan represent weights. The combined prediction model obtained by these methods is as follows:

[0187] a. Simple average method

[0188] The weights of the combined models represented by "bg-g-bl-l", "imf", and "ceemdan" are shown in Equation (51), Equation (52), and Equation (53), respectively.

[0189]

[0190] b. MAE reciprocal method

[0191] The weights of the combined models represented by "bg-g-bl-l", "imf", and "ceemdan" are shown in Equation (54), Equation (55), and Equation (56), respectively.

[0192] W bg_g_bl_l = [0.3266 0.2752 0.2153 0.1829] (54)

[0193]

[0194] W ceemdan= [0.3077 0.2662 0.2464 0.1797] (56)

[0195] c. Lagrange multiplier method

[0196] The weights of the combined models represented by "bg - g - bl - l", "imf", and "ceemdan" are shown in Equation (57), Equation (58), and Equation (59) respectively.

[0197] W bg_g_bl_l = [4.2063 -2.7239 -0.9561 0.4737] (57)

[0198]

[0199] W ceemdan = [0.9193 0.0673 -0.0035 0.0169] (59)

[0200] Figure 8 It shows that the combined prediction results using the Lagrange multiplier method are better than those of the combined prediction results of the reciprocal of MAE method and the simple average method. Figure 8 (b) Clearly shows that the multi - model combined prediction results using the Lagrange multiplier method on the far right are closely distributed on both sides of the original data. Figure 8 (c) shows that the prediction errors after using the multi - model combination are all concentrated around 0, and only a very low proportion of the absolute values of the errors are greater than 10. This proportion is also much smaller than Figure 6 (c). This indicates that the prediction results of the multi - model combination are much better than those of the single model. From the values in Table 3, when the models are combined in the same way, after obtaining the weights, using the Lagrange multiplier method to predict the multi - model combination can obtain the best prediction results. Such a conclusion can also be drawn from Figure 9 this.

[0201] Table 3 Evaluation indicators for combined model prediction

[0202]

[0203]

[0204] In Table 3, take the statistical indicators of the prediction results of the model combination method "imf" (see the above explanation of the model combination method) as an example. The weights are obtained using the simple average method to get the final prediction results of the combined model. The values of its internal indicators RMSE, R2, MAE, and MAPE are 6.256345, 0.972717, 5.499121, and 0.102130 respectively. At the same time, the values of the indicators RMSE, R2, MAE, and MAPE of the combined model obtained by using the MAE inversion method are 5.880983, 0.975893, 5.294846, and 0.104462 respectively. Conversely, the values of the indicators RMSE, R2, MAE, and MAPE of the combined model using the Lagrange multiplication to obtain weights are 4.784601, 0.984043, 3.674884, and 0.079240 for RMSE, R2, MAE, and MAPE respectively. For RMSE, it changed by -5.999701% and -18.642836%; R2 changed by 1.183331% and 0.835132%; MAE changed by -3.714685% and -30.595073%; MAPE changed by 2.283364% and -24.144665%. At the same time, after using the Lagrange multiplier method to obtain weights, the prediction results obtained by using other model combinations have also been significantly improved.

[0205] In Table 2, the CEEMDAN decomposition algorithm is used to decompose the data, and then the best results of predicting using the deep learning model are SSA-Bi-GRU-CEEMDAN, with RMSE and R2 being 4.810319 and 0.983518 respectively. The best results of directly predicting using the deep learning model are SSA-Bi-GRU, with RMSE and R2 being 5.616658 and 0.978011 respectively. However, combining the results of Table 3, through the use of the CEEMDAN combination algorithm and the integration using the Lagrange multiplication, multiple deep learning models are combined. The RMSE of the best prediction obtained is 4.764355, a decrease of 0.955529%, and R2 is 0.984178, an increase of 0.067106%. The RMSE of the prediction obtained by using this method to combine multiple directly used deep learning models is 3.858319, a decrease of 31.305787%, and R2 is 0.989624, an increase of 1.187410%.

[0206] To sum up: (1) GRU, Bi-GRU, LSTM, and Bi-LSTM are used to predict the undecomposed data and the data decomposed by CEEMADN. The results show that the CEEMDAN decomposition algorithm can improve the prediction effect. Specifically, the average RMSE is reduced by 11.245%, the average R2 is increased by 0.866%, and the average MAE is reduced by 8.165%.

[0207] (2) A multi - model combination method based on the Lagrange multiplier method is designed. This method can obtain the weights of each deep - learning model, and these weights can combine multiple models. The result of the multi - model combination is better than that of a single model.

[0208] (3) The Lagrange multiplication is compared with the simple average combination model and the MAE reverse combination model. The experimental results show that the result obtained by using the Lagrange multiplier method is better than the other two methods.

[0209] A short - term air quality prediction model of multi - model simplified combination using the sparrow search algorithm and decomposition error correction is proposed in this paper. This analytical combination model is used to predict air quality. On the one hand, using the CEEMDAN decomposition algorithm to decompose the data and then using the deep - learning model for prediction can obtain better results. On the other hand, this combination model has obvious improvement compared with the prediction result of using a single model. At the same time, its prediction result is also better than other traditional combination methods.

[0210] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. A short-term air quality prediction method based on the sparrow search algorithm and decomposition error correction, characterized in that, It includes the following steps: S1. Decompose the data using the CEEMDAN decomposition algorithm to obtain several IMFs; S2. Use a deep learning model to predict the several IMFs, process the prediction errors of multiple deep learning models to obtain the weights of each deep learning model; the number of deep learning models is at least two; S2-1. Use different deep learning models to predict the several IMFs to obtain multiple prediction results; S2-2. Use any one of the simple average method, the reciprocal of MAE method, and the Lagrange multiplier method to process the prediction errors of each model to obtain the weights of each deep learning model; S2-3. Finally, combine the prediction results of each deep learning model with the weights to generate the result of multi-model combined prediction; S3. Combine the prediction results of each deep learning model according to the weights to obtain the comprehensive prediction result, that is, the result of multi-model combined prediction; The deep learning models include: LSTM, Bi-LSTM, GRU, and Bi-GRU; The deep learning model is obtained by training with the sparrow search algorithm, including the following steps: SA. Normalize the data decomposed by CEEMDAN, and then classify it into a uniform distribution of statistical probabilities between 0 and 1; SB. Divide the classified data into a training data set and a test data set; SC. Use the sparrow search algorithm SSA to optimize the hyperparameters of the deep learning model, and the hyperparameters include the number of neurons in each layer of the neural network, the number of iterations, and the learning rate; SD. Use the deep learning model to predict the test data set and evaluate it according to the evaluation index. If the hyperparameters of the model do not overfit and the error of the prediction result is the smallest, the optimal hyperparameters and thus the optimal deep learning model are obtained.

2. The short-term air quality prediction method based on the sparrow search algorithm and decomposition error correction according to claim 1, wherein Before step S1, preprocess the data, including the following steps: S0-1. Use cubic spline interpolation to fill in the missing values; S0-2. Use the moving average method to process the data after cubic spline interpolation.

3. A short-term air quality prediction method based on the sparrow search algorithm and decomposition error correction according to claim 1, characterized in that The evaluation indexes include any combination of mean absolute error, root mean square error, correlation coefficient, and mean absolute percentage error.

4. A short-term air quality prediction method based on the sparrow search algorithm and decomposition error correction according to claim 1, characterized in that, Before step S1, it includes S0 to obtain data, and the data is meteorological data.

Citation Information

Patent Citations

  • Lake TN prediction method based on VMD-CSSA-LSTM-MLR combined model

    CN113762078A

  • Short-term offshore wind power prediction method based on CEEMDAN-SSA-BiLSTM

    CN114611808A