Air Quality Prediction Method Based on Secondary EEMD Decomposition Combined with GAFSA-LSTM

By using the combined method of secondary EEMD decomposition and GAFSA-LSTM model in air quality prediction, the problems of low prediction accuracy and large calculation amount in the prior art are solved, and more efficient and accurate air quality prediction is achieved.

CN115372550BActive Publication Date: 2025-06-13HUAIAN FANYUN SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210856812.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-20
Publication Date
2025-06-13
Estimated Expiration
2042-07-20

AI Technical Summary

Technical Problem

The existing air quality prediction methods have problems such as low accuracy of prediction results, large time loss rate, poor general use performance and low reliability.

Method used

The air quality prediction method based on secondary EEMD decomposition combined with GAFSA-LSTM was adopted, and the characteristics of the AQI sequence were extracted through EEMD decomposition, and the GAFSA-optimized LSTM model was constructed to perform model training and prediction.

Benefits of technology

It improves the accuracy of air quality prediction, reduces the calculation amount of LSTM network model, is suitable for small and medium-sized batch data sets, and is highly versatile.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115372550B_ABST
    Figure CN115372550B_ABST
Patent Text Reader

Abstract

The present invention discloses an air quality prediction method based on secondary EEMD decomposition combined with GAFSA-LSTM. EEMD is used for the data itself, and combined with GAFSA-optimized LSTM to achieve the prediction of air quality values. Using the air quality data with a time length of T segments, the AQI value of the air quality at the next time point t is predicted. The AQI column in the original dataset is used as the signal column to be decomposed by secondary EEMD. Through secondary EEMD decomposition, the original AQI column is decomposed into N columns of IMF components and 1 column of residual components. Then, the residual components are subjected to secondary EEMD decomposition to obtain K columns of secondary decomposition IMF components and 1 column of residual components. These N+K+1 column component data are used as the result values and are respectively brought into the LSTM model optimized by GAFSA for training and prediction to obtain the predicted values of each component. Finally, the N+K+1 column predicted values are linearly accumulated to achieve reconstruction, and the final air quality prediction value is obtained. The obtained result is compared with the original AQI value. The present invention can obtain more accurate predicted values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the prediction of the Air Quality Index (AQI), and particularly to an air quality prediction method based on quadratic Ensemble Empirical Mode Decomposition (EEMD) combined with Gravitational Search Algorithm - Long Short - Term Memory (GAFSA - LSTM). Background Technique

[0002] To objectively evaluate the environmental quality status of a region, it is necessary to comprehensively consider the intricate relationships among various factors and between influencing factors and environmental quality. Currently, there are methods such as statistical prediction, numerical prediction, and artificial intelligence algorithms represented by deep learning for prediction. For example, Li Yonghui used spatial statistical methods to predict the air quality in Beijing and applied the spatial autoregressive model and spatial autocorrelation to air quality analysis, achieving certain results; Zhang Taixin used data provided by the geostationary satellite Himawari - 8 / AHI to construct an AOD inversion model for predicting PM2.5, and on this basis, established a factor screening mechanism to improve the generalization performance of the model; Pu Guolin et al. proposed an air quality prediction based on an improved neural network; Bai Shengnan et al. proposed a PM2.5 prediction based on the LSTM recurrent neural network; Cheng Rong et al. proposed a local air quality prediction model based on neural random forests. The above - mentioned methods all have defects such as low prediction accuracy, high time loss rate, weak generalization performance, and low reliability. Summary of the Invention

[0003] Object of the Invention: Aiming at the problems pointed out in the background technique, the present invention provides a prediction method based on quadratic EEMD decomposition combined with GAFSA - LSTM. The EEMD decomposition target is set as the AQI sequence in the dataset, a GAFSA - optimized LSTM model is constructed, and the AQI sequence after two - time EEMD decomposition is used as the predicted label value and input into the GAFSA - LSTM network model to train the GAFSA - LSTM model, thereby realizing the prediction of air quality.

[0004] Technical Solution: The present invention discloses an air quality prediction method based on quadratic EEMD decomposition combined with GAFSA - LSTM, including the following steps:

[0005] Step 1: For the air quality situation in a certain region, based on the air quality detection platform system in this region, monitor the PM 2.5 concentration, PM 10 concentration, SO 2 concentration, CO concentration, NO 2 concentration, O 3 concentration, and AQI value, obtain the above - mentioned initial air quality sequence x(t), after cleaning and pre - processing this sequence, obtain the dataset D, and split it into a training set and a test set;

[0006] Step 2: Based on the initial air quality sequence x(t) obtained in Step 1, perform decomposition on the AQI column in the initial air quality sequence using Ensemble Empirical Mode Decomposition (EEMD) to obtain several column components and one column of residual component r1(t).

[0007] Step 3: Based on the one column of residual component r1(t) obtained after decomposition in Step 2, perform decomposition on this residual component again using Ensemble Empirical Mode Decomposition (EEMD) to obtain several column components and one column of residual component r2(t).

[0008] Step 4: Establish a GAFSA-LSTM network model. Among them, the role of GAFSA is to iteratively optimize the number of hidden layer neurons V and the learning rate C in LSTM. Use the GAFSA-LSTM network model to train and predict the decomposed components. Respectively use the component values obtained by performing EEMD decomposition on the AQI column in Step 2 and Step 3 as the input labels of the training set of the GAFSA-LSTM model and input them into the model for training. Among them, the residual component r2(t) is retained and not input into the model.

[0009] Step 5: Output the results of the training set, perform a reconstruction operation on the result values, and verify the prediction effect of the model on the training set.

[0010] Step 6: According to Steps 2, 3, 4, and 5, verify the established prediction model on the test set to achieve the prediction of the air quality AQI value.

[0011] Furthermore, the specific operation of the one-time Ensemble Empirical Mode Decomposition (EEMD) in Step 2 is as follows:

[0012] (1) Split the dataset D obtained in Step 1 into a training set and a test set. Among them, the training set is D 1 , and the test set is D 2 . The training set D 1 : D 1 = {(x 1 , y 1 ), (x 2 , y 2 ),..., (x n , y n )}, where x 1 , x 2 ,..., x n are air quality factor values, including PM 2.5 concentration, PM 10 concentration, SO 2 concentration, CO concentration, NO 2 concentration, O 3 concentration, and y1 , y 2 , ..., y n is the Air Quality Index (AQI) value;

[0013] (2) The amplitude of EEMD decomposition is set to a, the decomposition times of Gaussian white noise n(t) is set to M, and this noise follows the normal distribution law;

[0014] (3) Add this Gaussian white noise n(t) to the initial air quality sequence x(t), and perform the m-th Empirical Mode Decomposition (EMD);

[0015] 1) Find all local minimum points and maximum points in the sequence after adding noise, and use the 5th order spline interpolation function to fit to obtain the upper envelope sequence u(t) and the lower envelope sequence v(t), and take their average to get the average sequence m(t) = (u(t) + v(t)) / 2;

[0016] 2) Subtract the average sequence from the sequence after adding noise to get the intermediate sequence h(t) = x(t) - m(t);

[0017] 3) Judge whether h(t) meets the Intrinsic Mode Function (IMF). If it meets, define it as IMF1. Otherwise, regard it as the new x(t) and repeat steps 1) - 2) until the m-th EMD decomposition is completed;

[0018] 4) Calculate the residual component r1(t) = x(t) - h(t); Repeat steps 1) - 3) for the residual component r1(t) until the decomposition termination condition is met and then stop;

[0019] (4) Perform ensemble averaging on the IMF components obtained from M times of EMD decomposition to cancel the amplitude influence brought by the added white noise;

[0020] (5) Perform one - time EEMD decomposition on the AQI column in the dataset D 1 obtained in step 1, to get N columns of IMF components of one - time EEMD decomposition and 1 column of the residual component r1(t) of one - time EEMD decomposition.

[0021] Furthermore, the second - order ensemble empirical mode decomposition in step 3 is specifically described as follows:

[0022] (1) Use the residual component r1(t) after EEMD decomposition in step 2 as the input value for the second - order EEMD decomposition and input it;

[0023] (2) Set the parameters of the second - order EEMD decomposition model the same as in step 2, which are respectively set as: the amplitude is set to a, the decomposition times of Gaussian white noise n(t) is set to M, and this noise follows the normal distribution law;

[0024] (3) The method of changing the empirical mode decomposition (EMD) model to the ensemble empirical mode decomposition (EEMD) model by adding Gaussian white noise n(t) is the same as in step 2, which are respectively expressed as:

[0025] 1) Find all the local minima and maxima in the sequence after adding noise, and use the 5th - order spline interpolation function to fit and obtain the upper envelope sequence u(t) and the lower envelope sequence v(t), and take their average to get the average sequence m(t) = (u(t) + v(t)) / 2;

[0026] 2) Subtract the average sequence from the sequence after adding noise to get the intermediate sequence h(t) = x(t) - m(t);

[0027] 3) Determine whether h(t) satisfies the intrinsic mode function (IMF). If it does, define it as IMF2; otherwise, regard it as the new x(t) and repeat steps 1) - 2) until the m - th EMD decomposition is completed;

[0028] 4) Calculate the residual component r2(t) = x(t) - h(t); repeat steps 1) - 3) for the residual component r2(t) until the decomposition termination condition is met and then stop;

[0029] (4) Perform an ensemble average on the IMF components obtained from the M - th EMD decomposition to offset the amplitude influence brought by the added white noise;

[0030] (5) Substitute the first - order decomposition residual component obtained in step 2 into the EEMD model again for the second - order EEMD decomposition to obtain K columns of IMF components from the second - order EEMD decomposition and 1 column of the residual component r2(t) from the second - order EEMD decomposition.

[0031] Furthermore, step 4 constructs a GAFSA - LSTM network model, which is specifically described as:

[0032] (1) Use the GAFSA - LSTM network model to input the components obtained after the decomposition in step 3 into the GAFSA - LSTM model respectively and make predictions respectively to obtain the prediction results;

[0033] (2) Use GAFSA to iteratively optimize the number of hidden - layer neurons V and the learning rate C in the LSTM network model. The construction method of the global artificial fish swarm algorithm (GAFSA) is as follows:

[0034] 1) Initialize the GAFSA parameters, including: the initial position of the artificial fish, the visual field Visual of the artificial fish, the maximum step size Step for the artificial fish to move, the maximum number of iterations Iter, the total number N of individual artificial fish GAFSA , the number of attempts TryNum, the crowding factor δ, the number of hidden - layer neurons V and the learning rate C;

[0035] 2) GAFSA sets the parameters to be optimized, sets each artificial fish as a parameter combination of a group (V, C), and initializes them.

[0036] 3) GAFSA sets the target value, brings the parameters to be optimized into the LSTM network model, combines the training set and test set of air quality data, obtains the mean absolute percentage error MAPE, and takes the minimum value of this mean absolute percentage error as the target value for parameter optimization, where

[0037] 4) GAFSA fitness calculation, each artificial fish performs swarming behavior, following behavior, foraging behavior, and random behavior respectively, calculates the fitness of each artificial fish, compares all the artificial fish, obtains the best one and records it, where the descriptions of each behavior of GAFSA are as follows:

[0038] ① Foraging behavior: Let the current state of artificial fish i GAFSA be X GAFSAi , randomly select a state X GAFSAj within its perception range, and the mathematical expression is as follows:

[0039] X GAFSAj = X GAFSAi + Visual·Rand();

[0040] In the problem of finding the maximum value, if Y GAFSAi < Y GAFSAj , then take a step forward in the vector sum direction of this position and the global optimal position X best ; if this condition is not met, reselect the state of X j , and the mathematical expression is as follows: if Y GAFSAi < Y GAFSAj is true, then execute:

[0041]

[0042] Otherwise, execute: X i / next = X GAFSAi · Step·Rand();

[0043] If the maximum number of attempts is exceeded, that is, when i GAFSA > TryNum and this condition is still not met, then move randomly one step;

[0044] ② Swarming behavior: Let the current state of artificial fish i GAFSA be X GAFSAi , find the number of partners nf and the center position X i,i within its neighborhood d GAFSAc < Visual, if It indicates that the food concentration at the center of the fish school is high and the crowding degree is low, then move one step forward in the direction of the vector sum of this position and the global optimal position X best ; otherwise, perform the foraging behavior. The mathematical expression is as follows: If When, then execute:

[0045]

[0046] Otherwise, perform the foraging behavior;

[0047] ③ Following behavior: Let the current artificial fish i GAFSA be in the state of X GAFSAi , and there is an artificial fish X i,j with the maximum food concentration in its neighborhood d max . If it indicates that this fish X max currently has a relatively high food concentration and a low surrounding crowding degree, then move one step forward in the direction of the vector sum of the direction of this fish and the global optimal position X best ;

[0048] Otherwise, perform the foraging behavior. The mathematical expression is as follows: If When, then execute:

[0049]

[0050] Otherwise, perform the foraging behavior.

[0051] ④ Random behavior: Randomly select a state within the field of vision to execute, and then move in the direction of this state. It is a default of the foraging behavior. The mathematical expression is as follows:

[0052] X i / next =(X max -X min )·Rand + X min ;

[0053] 5) GAFSA selects and executes behaviors. Simulate and evaluate the aggregation behavior and following behavior for each artificial fish, and actually execute the behavior with a high evaluation degree. The foraging behavior and random behavior represent defaults;

[0054] 6) GAFSA updates the position information. After each artificial fish executes the selected behavior, update the position information of the artificial fish based on the global information and local information;

[0055] 7) GAFSA updates the state and updates the state of the global optimal artificial fish;

[0056] 8) GAFSA iteration and termination: If the set number of iterations is reached, jump out of the loop, terminate, and output the current optimal artificial fish state, i.e., the parameter combination of the optimal (V, C), and substitute it into LSTM to obtain its prediction result; otherwise, jump to 4) and continue to execute;

[0057] (3) Use the decomposed AQI sequence as the label value in the training set of the GAFSA-LSTM model. There are a total of N + K + 1 columns and input them into the GAFSA-LSTM model.

[0058] Furthermore, the signal reconstruction method is used in step 5, and the specific description is as follows:

[0059] (1) Output the result value obtained by the GAFSA-LSTM model;

[0060] (2) Reconstruct the above result value. By using the linear accumulation method, the predicted value Fes(t) of the training set can be obtained, and its expression is:

[0061]

[0062] Among them, imf i (t) represents the IMF component result obtained by one GAFSA-EEMD decomposition, and imf2 i (t) represents the IMF component obtained by the second GAFSA-EEMD decomposition, and r2 n (t) represents the residual component obtained after the second GAFSA-EEMD decomposition.

[0063] Furthermore, in step 6, the prediction model trained in steps 2 - 5 is applied to the test set to realize the prediction of air quality. The predicted value is compared with the actual value to comprehensively evaluate the model effect. The evaluation indicators include mean square error MSE, R2 determination coefficient, mean absolute error MAE, and mean absolute percentage error MAPE. The specific calculation formulas are as follows:

[0064]

[0065]

[0066]

[0067]

[0068] Beneficial effects:

[0069] The present invention relates to an air quality prediction method based on secondary EEMD decomposition combined with GAFSA-LSTM, which makes full use of the respective advantages of EEMD, GAFSA, and LSTM to improve the accuracy of prediction results. For the air quality situation in a certain area, based on the air quality detection platform system in this area, the PM 2.5 concentration, PM 10 concentration, SO 2 concentration, CO concentration, NO 2 concentration, O 3 concentration, and AQI value are monitored. The air quality prediction method based on secondary EEMD decomposition combined with GAFSA-LSTM is used to achieve the prediction of the AQI value of air quality. The prediction accuracy obtained by bringing the secondary EEMD decomposition into the GAFSA-LSTM model has a certain improvement compared with no decomposition and single EEMD decomposition; the decomposed AQI sequence reduces the computational complexity of the LSTM network model and is applicable to medium and small batch data sets; this method has strong versatility and is applicable to air quality prediction in different regions. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 FIG. is the overall flowchart of the air quality prediction method based on secondary EEMD decomposition combined with LSTM;

[0071] Figure 2 FIG. is the original data graph of the air quality in Shanghai;

[0072] Figure 3 FIG. is the data graph of the air quality test set in Shanghai;

[0073] Figure 4 FIG. is the single EEMD decomposition graph;

[0074] Figure 5 FIG. is the secondary EEMD decomposition graph;

[0075] Figure 6 FIG. is the result comparison graph of the air quality prediction method based on secondary EEMD decomposition combined with GAFSA-LSTM. DETAILED DESCRIPTION OF THE INVENTION

[0076] The following further clarifies the present invention with specific examples. It should be understood that these examples are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications made by those skilled in the art to the present invention fall within the scope defined by the appended claims of this application.

[0077] The present invention discloses an air quality prediction method based on secondary EEMD decomposition combined with GAFSA-LSTM. Refer to the attached Figure 1 figures, which specifically include the following steps:

[0078] Step 1: Sample and preprocess the data required for the experiment. The specific operations are as follows:

[0079] Step 1.1: As shown in the appendix Figure 1 For the air quality situation in a certain area, based on the air quality detection platform system in this area, monitor the PM 2.5 concentration, PM 10 concentration, SO 2 concentration, CO concentration, NO 2 concentration, O 3 concentration, and AQI value within a certain time interval in this area. Sample the air quality data D through an external air monitor.

[0080] Step 1.2: Perform two EEMD decompositions on the AQI column in the sampled data D. Put the decomposed sequences together with the air quality prediction factors into the GAFSA-LSTM network model for training and prediction. Reconstruct the obtained prediction values to obtain the final prediction values, that is, the air quality prediction values for this area.

[0081] Step 1.3: The dataset obtained by sampling is the air quality data of Shanghai from December 2013 to December 2020. The detection interval is 1 day, with a total of 2,500 data. Some data are shown in Table 1; among them, the air quality factors are: PM 2.5 concentration, PM 10 concentration, SO 2 concentration, CO concentration, NO 2 concentration, O 3 concentration. According to the "Technical Regulations on Ambient Air Quality Index (AQI) (Trial)" issued by the state, cancel the original Air Pollution Index (API), and instead adopt the Air Quality Index (AQI). Among them, the air quality index AQI is divided into six levels, namely: excellent, good, light pollution, moderate pollution, moderate pollution, and severe pollution;

[0082] Step 1.4: Split the dataset D into a training set D1 and a test set D2, and divide them according to a ratio of 8:2. The training set D 1 consists of the first 2,000 days of data, and the test set D 2 consists of the last 500 days of data; among them, PM 2.5 concentration, PM 10 concentration, SO 2 concentration, CO concentration, NO 2 concentration, O 3 concentration are air quality impact factors, which are used as the input values of the objective function, and the AQI value is used as the label value of the objective function. Some data of the air quality in Shanghai are shown in Table 1. The original dataset D of the air quality in Shanghai and the data of the test set D2 , such as Figure 2 、 Figure 3 shown.

[0083] Table 1: Part of the air quality data in Shanghai

[0084]

[0085]

[0086]

[0087] Step 2: The specific operation of the first - stage EEMD decomposition is as follows:

[0088] Step 2.1: Set the amplitude of the EEMD decomposition as a, the decomposition times of the Gaussian white noise n(t) as M, and this noise follows the normal distribution law.

[0089] Step 2.2: Add the Gaussian white noise n(t) to the initial air quality sequence x(t) and perform the m - th empirical mode decomposition EMD.

[0090] Step 2.3: Find all the local minimum points and maximum points in the sequence after adding noise, and use the 5 - th order spline interpolation function to fit to obtain the upper envelope sequence u(t) and the lower envelope sequence v(t), and take their average to get the average sequence m(t)=(u(t)+v(t)) / 2.

[0091] Step 2.4: Subtract the average sequence from the sequence after adding noise to get the intermediate sequence h(t)=x(t)-m(t).

[0092] Step 2.5: Judge whether h(t) satisfies the intrinsic mode function IMF. If it satisfies, define it as IMF1; otherwise, regard it as the new x(t) and repeat steps 1) - 2) until the m - th EMD decomposition is completed.

[0093] Step 2.6: Calculate the residual component r1(t)=x(t)-h(t); repeat steps 1) - 3) for the residual component r1(t) until the decomposition termination condition is met and then stop.

[0094] Step 2.7: Perform the first - stage EEMD decomposition on the AQI column in the dataset D 1 obtained in step 1, to get 10 columns of IMF components of the first - stage EEMD decomposition and 1 column of the residual component r1(t) of the first - stage EEMD decomposition. The decomposition results are as Figure 4 shown.

[0095] Step 3: The specific operation of the second - stage EEMD decomposition is as follows:

[0096] Step 3.1: Take the residual component r2(t) after EEMD decomposition in Step 2 as the input value for the second EEMD decomposition and input it therein.

[0097] Step 3.2: The settings of each parameter of the second EEMD decomposition model are the same as those in Step 2, and are respectively set as follows: the amplitude is set as a, the decomposition times of the Gaussian white noise n(t) are set as M, and the noise follows the normal distribution law.

[0098] Step 3.3: The method of changing the empirical mode decomposition EMD model to the ensemble empirical mode decomposition EEMD model by adding the Gaussian white noise n(t) is the same as that in Step 2.

[0099] Step 3.4: Find all the local minimum points and maximum points in the sequence after adding the noise, and use the 5th-order spline interpolation function to fit to obtain the upper envelope sequence u(t) and the lower envelope sequence v(t), and take their average value to obtain the average value sequence m(t) = (u(t) + v(t)) / 2.

[0100] Step 3.5: Subtract the average value sequence from the sequence after adding the noise to obtain the intermediate sequence h(t) = x(t) - m(t).

[0101] Step 3.6: Judge whether h(t) satisfies the intrinsic mode function IMF. If it satisfies, define it as IMF1; otherwise, regard it as the new x(t) and repeat Steps 1)-2) until the mth EMD decomposition is completed.

[0102] Step 3.7: Calculate the residual component r2(t) = x(t) - h(t); repeat Steps 1)-3) for the residual component r2(t) until the decomposition termination condition is met and then stop.

[0103] Step 3.8: Perform ensemble averaging on the IMF components obtained from the M times of EMD decomposition to cancel the amplitude influence brought by the added white noise.

[0104] Step 3.9: Substitute the primary decomposition residual component obtained in Step 2 into the EEMD model again for the second EEMD decomposition to obtain 9 columns of IMF components of the second EEMD decomposition and 1 column of the residual component r2(t) of the second EEMD decomposition. The decomposition result is as Figure 5 shown.

[0105] Step 4: Input the decomposed air quality impact factors, 19 columns of IMF components and 1 column of residual component of the AQI sequence into the GAFSA-LSTM network for training. The specific operations are as follows:

[0106] Step 4.1: The number of hidden layer neurons V and the learning rate C in the LSTM network model are iteratively optimized using GAFSA. The optimization results are that the number of hidden layer neurons V is set to 1 and the learning rate C is set to 0.0012. The other parameter settings are as follows: the feature dimension of the input value is 10 + 9; the single-step prediction node of the output layer is 1, and the window length is 36; the number of three-step prediction nodes is 3, and the window length is 72.

[0107] The construction method of the global artificial fish swarm GAFSA is as follows:

[0108] 1) Initialize the GAFSA parameters, and set the parameters including: the initial position of the artificial fish, the visual field Visual of the artificial fish, the maximum step length Step for the artificial fish to move, the maximum number of iterations Iter, the total number N of individual artificial fish GAFSA , the number of attempts TryNum, the crowding factor δ, the number of hidden layer neurons V, and the learning rate C;

[0109] 2) Set the parameters to be optimized in GAFSA. Set each artificial fish as a parameter combination of a group (V, C) and initialize it;

[0110] 3) Set the target value in GAFSA. Substitute the parameters to be optimized into the LSTM network model, combine the training set and the test set of the air quality data, and obtain the mean absolute percentage error MAPE. Take the minimum value of this mean absolute percentage error as the target value for parameter optimization, where

[0111] 4) Calculate the fitness in GAFSA. Each artificial fish performs the aggregation behavior, the following behavior, the foraging behavior, and the random behavior respectively, and calculate the fitness of each artificial fish. Compare all the artificial fish to obtain the optimal one and record it. The descriptions of each behavior in GAFSA are as follows:

[0112] ① Foraging behavior: Let the current artificial fish i GAFSA be in the state X GAFSAi . Randomly select a state X GAFSAj within its perception range. The mathematical expression is as follows:

[0113] X GAFSAj = X GAFSAi + Visual·Rand();

[0114] In the problem of finding the maximum value, if Y GAFSAi < Y GAFSAj , then take a step forward in the vector sum direction of this position and the global optimal position X best ; if this condition is not met, reselect the state of X j . The mathematical expression is as follows: If Y GAFSAi< Y GAFSAj then execute:

[0115]

[0116] otherwise execute: X i / next = X GAFSAi ·Step·Rand();

[0117] If the maximum number of attempts is exceeded, that is, i GAFSA > TryNum and the condition is still not satisfied, then move randomly one step;

[0118] ② Swarming behavior: Let the current artificial fish i GAFSA be in the state X GAFSAi , and look for the number of partners n i,j and the center position X f within its neighborhood d GAFSAc < Visual. If indicates that the food concentration at the center of the fish school is high and the crowding degree is low, then take one step forward in the direction of the vector sum of this position and the global optimal position X best ; otherwise execute the foraging behavior. The mathematical expression is as follows: If then execute:

[0119]

[0120] otherwise execute the foraging behavior;

[0121] ③ Following behavior: Let the current artificial fish i GAFSA be in the state X GAFSAi , and there is an artificial fish X i,j with the maximum food concentration within its neighborhood d max < Visual. If indicates that this fish X max currently has a relatively high food concentration and a low surrounding crowding degree, then take one step forward in the direction of the vector sum of the direction of this fish and the global optimal position X best ;

[0122] otherwise execute the foraging behavior. The mathematical expression is as follows: If then execute:

[0123]

[0124] otherwise execute the foraging behavior.

[0125] ④ Random behavior: Randomly select a state within the field of vision to execute, and then move in the direction of this state. It is a kind of default of the foraging behavior. The mathematical expression is as follows:

[0126] Xi / next =(X max -X min )·Rand + X min ;

[0127] 5) GAFSA selects and executes behaviors. Simulate and evaluate the execution of clustering behavior and following behavior for each artificial fish, and actually execute the behavior with a high evaluation degree. Foraging behavior and random behavior are represented as default;

[0128] 6) GAFSA updates the position information. After each artificial fish executes the selected behavior, update the position information of the artificial fish based on the global information and local information;

[0129] 7) GAFSA updates the state and updates the state of the global optimal artificial fish;

[0130] 8) GAFSA iterates and terminates. If the set number of iterations is reached, jump out of the loop, terminate and output the current optimal artificial fish state, that is, the parameter combination of the optimal (V, C), and bring it into LSTM to obtain its prediction result; otherwise, jump to 4) and continue to execute.

[0131] Step 4.2: Use the decomposed AQI sequence as the label value in the training set of the LSTM model. There are a total of 20 columns and input them into the GAFSA-LSTM model.

[0132] Step 4.3: Use the set GAFSA-LSTM network model for training.

[0133] Step 5: Process the AQI sequence after two EEMD decompositions, specifically as follows:

[0134] Step 5.1: Output the result value obtained by the GAFSA-LSTM model;

[0135] Step 5.2: Reconstruct the above result value. By using the linear accumulation method, the predicted value res(t) of the training set can be obtained, and its expression is:

[0136]

[0137] where, imf i (t) represents the IMF component result obtained by one GAFSA-EEMD decomposition, imf2 i (t) represents the IMF component obtained by the second GAFSA-EEMD decomposition, and r2 n (t) represents the residual component obtained after the second GAFSA-EEMD decomposition.

[0138] Step 6: Verify the model effect on the test set, specifically as follows:

[0139] Step 6.1: Substitute the test set D 2 into the model for testing.

[0140] Step 6.2: Reconstruct the results of the test set.

[0141] Step 6.3: Comprehensively evaluate the performance of the model. The evaluation metrics include the mean squared error MSE, the R2 coefficient of determination, the mean absolute error MAE, and the mean absolute percentage error MAPE. Compare with the recurrent neural network RNN, LSTM, and the GAFSA-LSTM network model combined with the first EEMD decomposition. The results of the evaluation metrics are shown in Table 2, and the fitting results are as Figure 6 shown.

[0142] The evaluation metrics mean squared error MSE, R2 coefficient of determination, mean absolute error MAE, and mean absolute percentage error MAPE have the following specific calculation formulas:

[0143]

[0144]

[0145]

[0146]

[0147] Table 2: Evaluation index values of each model

[0148]

[0149] The letter meanings of the formulas involved in the above steps are as shown in Table 3 below:

[0150] Table 3: Variable table

[0151]

[0152]

[0153]

[0154] Through Figure 6As can be seen from Table 2, compared with RNN, LSTM has a better prediction effect and is also more excellent in the fitting effect on the result fitting graph. Therefore, compared with RNN, LSTM is more suitable for air quality prediction work. Compared with LSTM, the prediction effect after EEMD decomposition is more excellent, and both the result fitting and the average error are more excellent than the single LSTM algorithm. Thus, it can be seen that the combined algorithm of EEMD-LSTM is advisable and effective for air quality prediction. Further, compared with EEMD-LSTM, the algorithm of the present invention based on double EEMD decomposition combined with GAFSA-LSTM model (DoubleEEMD-GAFSA-LSTM) has overall improved the overall performance of the former, specifically manifested in that the average error value is lower and the degree of fitting of the result curve is higher, which means that the prediction effect of the Double EEMD-GAFSA-LSTM proposed by the present invention is better than the LSTM model used in previous studies, and alleviates the overfitting phenomenon that may occur in the model training stage of LSTM. It can also be seen from the result graph that at some "isolated points", the Double EEMD-GAFSA-LSTM of the present invention can also predict well, while the results of the comparative models at these points are somewhat unsatisfactory.

[0155] The above only expresses the preferred embodiments of the present invention, and its description is relatively specific and detailed, but it cannot be understood as a limitation to the scope of the patent of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations, improvements and substitutions can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.

[0156] In addition, the present invention can be combined with a computer system to complete the prediction of air quality. An air quality prediction method based on double EEMD decomposition combined with GAFSA-LSTM disclosed by the present invention can not only be used for air quality prediction, but also for the prediction of other types of data.

[0157] The above embodiments are only for explaining the technical concept and characteristics of the present invention, and their purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and cannot be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.

Claims

1. An air quality prediction method based on secondary EEMD decomposition combined with GAFSA-LSTM, characterized in that, it includes the following steps: Step 1: Monitor the PM concentration, PM concentration, SO concentration, CO concentration, NO concentration, O concentration, and AQI value in a certain area within a certain time interval, and obtain the initial air quality sequence x(t). After cleaning and preprocessing the sequence, obtain the dataset D and split it into a training set and a test set; 2.5 concentration, PM 10 concentration, SO 2 concentration, CO concentration, NO 2 concentration, O 3 concentration, AQI value, and obtain the initial air quality sequence x(t). After cleaning and preprocessing the sequence, obtain the dataset D and split it into a training set and a test set; Step 2: Based on the initial air quality sequence x(t) described in Step 1, use ensemble empirical mode decomposition EEMD to decompose the AQI column in the initial air quality sequence, and decompose it into several columns of components and 1 column of residual component r1(t); The specific operation of the one-time ensemble empirical mode decomposition EEMD in Step 2 is as follows: (1) Split the dataset D obtained in step 1 into a training set and a test set, where the training set is D 1 , and the test set is D 2 . The training set D 1 : D 1 = {(x 1 , y 1 ), (x 2 , y 2 ), …, (x n , y n )}, where x 1 , x 2 , …, x n are air quality factor values, including PM 2.5 concentration, PM 10 concentration, SO 2 concentration, CO concentration, NO 2 concentration, O 3 concentration, and y 1 , y 2 , …, y n are air quality coefficient AQI values; (2) The amplitude setting of EEMD decomposition is a, the decomposition times of Gaussian white noise n(t) is set to M, and this noise follows the normal distribution law; (3) Add this Gaussian white noise n(t) to the initial air quality sequence x(t) and perform the m-th empirical mode decomposition EMD; (4) Ensemble average the IMF components obtained by M times of EMD decomposition to offset the amplitude influence brought by the added white noise; Perform one - time EEMD decomposition on the AQI column in the dataset D obtained in step 1 1 to obtain N columns of IMF components from the one - time EEMD decomposition and 1 column of residual component r1(t) from the one - time EEMD decomposition; Step 3: Based on the 1 column of residual component r1(t) obtained after decomposition in Step 2, use ensemble empirical mode decomposition EEMD to decompose this residual component again, and decompose it into several columns of components and 1 column of residual component r2(t); The specific operation of the secondary ensemble empirical mode decomposition in Step 3 is as follows: (1) Use the residual component r1(t) after EEMD decomposition in Step 2 as the input value of the secondary EEMD decomposition and input it therein; (2) The parameters of the secondary EEMD decomposition model are respectively set as: the amplitude is set to a, the decomposition times of Gaussian white noise n(t) is set to M, and this noise follows the normal distribution law; (3) Add Gaussian white noise n(t) to change the empirical mode decomposition EMD model into an ensemble empirical mode decomposition EEMD model; (4) Ensemble average the IMF components obtained by M times of EMD decomposition to offset the amplitude influence brought by the added white noise; (5) Bring the one-time decomposition residual component obtained in Step 2 into the EEMD model again, perform secondary EEMD decomposition, and obtain K columns of IMF components of the secondary EEMD decomposition and 1 column of residual component r2(t) of the secondary EEMD decomposition; Step 4: Establish a GAFSA-LSTM network model, use the GAFSA-LSTM network model to train and predict the decomposed components, and use the component values obtained by performing EEMD decomposition on the AQI column in Steps 2 and 3 as the input labels of the training set of the GAFSA-LSTM model and input them into the model for training. Among them, the residual component r2(t) is retained and not input into the model; The specific operation of constructing the GAFSA-LSTM network model in Step 4 is as follows: (1) Use the global artificial fish swarm GAFSA to iteratively optimize the number of hidden layer neurons V and the learning rate C in the LSTM network model; (2) Use the GAFSA-LSTM network model to input the components obtained after decomposition in Step 3 into the GAFSA-LSTM model respectively and perform predictions respectively to obtain prediction results; (3) Use the decomposed AQI sequence as the label values in the training set of the GAFSA-LSTM model, with a total of N+K+1 columns, and input it into the GAFSA-LSTM model; Step 5: Output the results of the training set, perform a reconstruction operation on the result values, and verify the prediction effect of the model on the training set; Step 6: According to Steps 2 to 5, verify the established prediction model on the test set to achieve the prediction of the air quality AQI value.

2. The air quality prediction method based on quadratic EEMD decomposition combined with GAFSA-LSTM according to claim 1, characterized in that, The specific operation of performing the m-th empirical mode decomposition EMD is as follows: 1) Find all the local minimum points and maximum points in the sequence after adding noise, and use the 5th-order spline interpolation function to fit to obtain the upper envelope sequence u(t) and the lower envelope sequence v(t), and take their average value to obtain the average sequence m(t)=(u(t)+v(t)) / 2; 2) Subtract the average sequence from the sequence after adding noise to obtain the intermediate sequence h(t)=x(t)-m(t); 3) Judge whether h(t) satisfies the intrinsic mode function IMF. If it satisfies, define it as IMF1. Otherwise, regard it as the new x(t) and repeat Steps 1)-2) until the m-th EMD decomposition is completed; 4) Calculate the residual component r1(t)=x(t)-h(t); repeat Steps 1)-3) for the residual component r1(t) until the decomposition termination condition is met and stop.

3. The air quality prediction method based on quadratic EEMD decomposition combined with GAFSA-LSTM according to claim 1, characterized in that, The specific operation of changing the empirical mode decomposition EMD model to the ensemble empirical mode decomposition EEMD model is as follows: 1) Find all the local minimum points and maximum points in the sequence after adding noise, and use the 5th-order spline interpolation function to fit to obtain the upper envelope sequence u(t) and the lower envelope sequence v(t), and take their average value to obtain the average sequence m(t)=(u(t)+v(t)) / 2; 2) Subtract the average sequence from the sequence after adding noise to obtain the intermediate sequence h(t)=x(t)-m(t); 3) Judge whether h(t) satisfies the intrinsic mode function IMF. If it satisfies, define it as IMF2. Otherwise, regard it as the new x(t) and repeat Steps 1)-2) until the m-th EMD decomposition is completed; 4) Calculate the residual component r2(t)=x(t)-h(t); repeat Steps 1)-3) for the residual component r2(t) until the decomposition termination condition is met and stop.

4. The air quality prediction method based on quadratic EEMD decomposition combined with GAFSA-LSTM according to claim 1, characterized in that, The construction method of the global artificial fish swarm GAFSA is as follows: 1) Initialize the GAFSA parameters, and set the parameters including: the initial position of the artificial fish, the visual field Visual of the artificial fish, the maximum step length Step for the artificial fish to move, the maximum number of iterations Iter, the total number N of individual artificial fish GAFSA , the number of attempts TryNum, the crowding factor δ, the number V of hidden layer neurons, and the learning rate C; 2) GAFSA sets the parameters to be optimized, sets each artificial fish as a parameter combination of (V,C), and performs initialization; 3) GAFSA sets the target value, brings the parameters to be optimized into the LSTM network model, combines the training set and test set of air quality data, obtains the mean absolute percentage error MAPE, and takes the minimum value of the mean absolute percentage error as the target value for parameter optimization, where 4) GAFSA fitness calculation: Each artificial fish performs the aggregation behavior, following behavior, foraging behavior, and random behavior respectively, calculates the fitness of each artificial fish, compares all the artificial fish, and obtains and records the optimal one among them; 5) GAFSA selects the execution behavior: Simulates and evaluates the aggregation behavior and following behavior for each artificial fish, and actually executes the behavior with a high evaluation degree. The foraging behavior and random behavior are default; 6) GAFSA updates the position information: After each artificial fish executes the selected behavior, updates the position information of the artificial fish based on the global information and local information; 7) GAFSA updates the state: Updates the state of the globally optimal artificial fish; 8) GAFSA iteration and termination: If the set number of iterations is reached, jump out of the loop, terminate and output the current optimal artificial fish state, that is, the parameter combination of the optimal (V, C), and bring it into LSTM to obtain its prediction result; otherwise, jump to 4) and continue to execute.

5. The air quality prediction method based on quadratic EEMD decomposition combined with GAFSA-LSTM according to claim 1, characterized in that, the step 5 uses a signal reconstruction method, and the specific description is as follows: (1) Output the result value obtained by the GAFSA-LSTM model; (2) Reconstruct the above result value by using the linear accumulation method to obtain the predicted value x(t) of the training set, and its expression is: Among them, imf i (t) represents the IMF component result obtained by one - time GAFSA - EEMD decomposition, and imf2 i (t) represents the IMF component obtained by two - time GAFSA - EEMD decomposition, and r2 n (t) represents the residual component obtained after two - time GAFSA - EEMD decomposition.

6. The air quality prediction method based on quadratic EEMD decomposition combined with GAFSA-LSTM according to claim 1, characterized in that, the step 6 applies the prediction model trained in steps 2-5 to the test set to realize the prediction of air quality, compares the predicted value with the actual value, and comprehensively evaluates the model effect. The evaluation indexes include mean square error MSE, R2 determination coefficient, mean absolute error MAE, and mean absolute percentage error MAPE. The specific calculation formulas are as follows: