An ozone concentration prediction method based on peak perception and impact factor analysis
By applying the Transformer model based on peak perception and impact factor analysis in Yunnan, the limitations of traditional methods in ozone concentration prediction are solved, and higher forecast accuracy and rapid and accurate forecasting of ozone heavy pollution processes are achieved.
Patent Information
- Application Number
- CN202411797797.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-09
AI Technical Summary
Traditional ozone concentration prediction methods have limitations in dealing with complex nonlinear relationships and long-term data, making it difficult to accurately predict ozone concentration and its influencing factors.
The Transformer model based on peak perception and impact factor analysis is adopted to train historical observation data in the large database of pollutants and meteorological data in Yunnan, establish an ozone forecast model for historical meteorological, pollutants and future ozone in Yunnan, and output n-main control factors through deep learning optimization model.
The accuracy of ozone forecasting is improved, especially the forecasting of the time and concentration of ozone peak occurrence, the forecasting effect of the ozone heavy pollution process is optimized, and a fast and accurate forecasting model suitable for Yunnan is established.
Smart Images

Figure CN119623754B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of environmental prediction, and particularly relates to an ozone concentration forecasting method based on peak perception and impact factor analysis. Background Art
[0002] Ozone is a strong oxidant. Its presence in the atmosphere can not only cause harm to human health, but also have a negative impact on crops and ecosystems. Traditional ozone concentration forecasting methods often rely on statistical models and empirical formulas, and these methods have limitations in dealing with the complex non-linear relationships and long-time series data of predicted ozone and its impact factors. Therefore, in order to solve the above problems, this paper proposes an ozone concentration forecasting method based on peak perception and impact factor analysis. Summary of the Invention
[0003] In order to solve the above technical problems, the present invention designs an ozone concentration forecasting method based on peak perception and impact factor analysis.
[0004] In order to achieve the above technical effects, the present invention is realized through the following technical solutions: An ozone concentration forecasting method based on peak perception and impact factor analysis, characterized by comprising the following steps:
[0005] S1. Obtain the monitoring data of pollutants and their meteorology at the hourly scale in Yunnan region for many years through the method of network data collection, and combine the ECMWF meteorological forecast products to provide the hourly scale data of boundary layer height, radiation, sea level pressure, and vertical wind speed. Then, perform outlier processing, null value filling, data normalization, and dataset division on the pollutants and their meteorological monitoring data to construct a large database of pollutants and meteorological data in Yunnan region to support deep learning;
[0006] S2. Use a Transformer model optimized based on impact factor analysis and peak perception to train the historical observation data in the large database of pollutants and meteorological data in Yunnan region and the dataset of additional meteorological parameters provided by ECMWF, establish an ozone forecasting model in Yunnan region for historical meteorology, pollutants and future ozone, and optimize the ozone forecasting model in Yunnan region through the method of peak perception and impact factor analysis based on deep learning, and output n - main control factors;
[0007] S3. Use the evaluation result of the optimized ozone forecasting model in Yunnan region in the validation set as an index to evaluate the quality of the deep learning model for ozone concentration evaluation; The evaluation indexes include mean square error MSE, mean absolute error MAE, determination coefficient R 2 , select the evaluation index
[0008] Metric = (MSE + MAE) / 2 - R 2As an evaluation index.
[0009] Select the model with the lowest Metric on the validation set as the optimal model, and use it for fast and accurate prediction; where MSE is the mean square error, MAE is the mean absolute error, and R 2 is the coefficient of determination.
[0010] Furthermore, in S1, the establishment of the large database of pollutants and meteorological data in Yunnan region is as follows:
[0011] Through network data download, collect and organize the continuous online observation data of six conventional monitoring pollutants (O 3 , CO, SO 2 , NO 2 , PM 2.5 , PM 10 ) in Yunnan region from 2015 to the present and meteorological observation data, and combine the ECMWF meteorological forecast products to provide hourly-scale meteorological numerical simulation data of boundary layer height, radiation, sea level pressure, and vertical wind speed, and construct a long-time-span historical dataset of ozone and meteorology, and other pollutants; and perform data outlier processing, null value filling, data normalization and dataset partitioning on the dataset to make it meet the data format required by deep learning, and establish a large database of pollutants and meteorological data in Yunnan region to support deep learning;
[0012] During the project implementation period, continuously organize the incoming data of meteorological and pollutant observation data to realize the dynamic update of the large database.
[0013] Furthermore, the above-mentioned data outlier processing, null value filling, data normalization and dataset partitioning are as follows:
[0014] Data outlier processing: If the monitoring data is too large due to abnormal reasons such as equipment damage or the equipment is insensitive, resulting in the monitoring data being continuously 0 for a period of time, this part of the data content in the conventional monitoring pollutant data and meteorological data will be excluded;
[0015] Null value filling: Due to reasons such as observation equipment or observation conditions, the phenomenon of discontinuous observation data is obvious. Therefore, the supplement of missing values is very important. The data filling method adopted in the present invention is the filling method based on statistical value interpolation of different time scales:
[0016] V ij =|V monthij +V weekdayij +V hourij -2×V yearij |
[0017] Among them, V ij is the filling value of pollutant i at the missing time j, and Vmonth ijis the average pollutant concentration of pollutant i in the month where the missing time j is located, Vweekday ij is the average pollutant concentration of pollutant i in the week where the missing time j is located, Vyeay ij is the average pollutant concentration of pollutant i in the year where the missing time j is located;
[0018] Data normalization: Data normalization is beneficial to improving the stability of training and the convergence speed of model training. The method adopted in the present invention is normal distribution normalization, specifically subtracting the mean from the characteristic attributes of the data and dividing by the variance; converting it into a standard normal distribution with a mean of 0 and a variance of 1:
[0019] x = (x - v) / σ
[0020] In the formula, μ is the data mean; σ is the data variance; record the mean and variance on the training set. After standardizing the training set, when using the model for subsequent testing, use the variance and mean on the training set to standardize the test data;
[0021] Dataset division: The data used is from 2015. More than a hundred ground observation stations in the monitoring network in Yunnan region began to conduct O 3 , NO 2 , CO, SO 2 , PM 2.5 , PM 10 hourly continuous observation data and ECMWF meteorological forecast products provide meteorological numerical simulation data at the hourly scale of boundary layer height, radiation, sea level pressure, and vertical wind speed; select the data of the recent year as the test set to evaluate the model ability, and divide the other data into the training set and the validation set according to the data volume ratio of 4:1; among them, the training set is used to train the model parameters, and the validation set is used to select the model with the best number of training rounds.
[0022] Further, in S2, based on big data deep learning, the ozone forecasting model in Yunnan region is optimized through the methods of peak perception and influencing factor analysis as follows:
[0023] S2.1. Use a Transformer model based on influencing factor analysis to train the historical dataset. Use the meteorological and atmospheric pollutant monitoring data of more than ten years and the numerical simulation data provided by ECMWF meteorological forecast products as the input of the neural network model, and the predicted ozone concentration as the output. Establish ozone forecasting models in Yunnan region at the hourly scale and daily scale respectively; and output the perception scores, select the largest n perception scores, and take the corresponding influencing factors as the n - main control factors;
[0024] S2.2. Meanwhile, apply the peak perception method, and respectively find the corresponding peaks in the daily-scale forecast and the hourly-scale forecast results by using the peak-finding algorithm, specifically including:
[0025] PeakIndex = PeakDetection(Target)
[0026] where Target is the true value of the predicted ozone concentration, Target ∈ R B×L‘ , PeakIndex is the index of the peak, PeakIndex ∈ {1, 2, 3,..., L‘}; PeakDetection is the peak-finding algorithm, which is the Automaticmultiscale-based peak detection (AMPD) algorithm here;
[0027] Calculate the peak perception loss function and perform gradient backpropagation
[0028] Loss peak = MSELoss(Targe PeakIndex ,FinalOut PeakIndex )
[0029] where Target PeakIndex , represents the sequence composed of Target i , where i ∈ PeakIndex; similarly, FinalOut PeakIndex , represents the sequence composed of FinalOut i , where i ∈ PeakIndex;
[0030] Loss MSE = MSELoss(Target,FinalOut)
[0031] Loss total = Loss MSE + αLoss peak
[0032] In the formula, α takes 1;
[0033] Then increase the penalty degree of the model at the corresponding peaks, so that the model can better perceive the time and corresponding concentration size of the peaks, improve the ozone peak forecast effect, and achieve rapid and accurate forecasting of future ozone heavy pollution processes in Yunnan region; thus, the optimization of the ozone forecasting model for Yunnan region can be completed.
[0034] Furthermore, the meteorological data set includes temperature, wind speed, wind direction, total cloud cover, relative humidity, atmospheric pressure, boundary layer height, radiation, sea-level pressure, and vertical wind speed.
[0035] The beneficial effects of the present invention are as follows:
[0036] The present invention utilizes continuous online observation data of pollutants and meteorological observation data at multiple sites. By processing abnormal and missing data, and combining ECMWF meteorological forecast products to provide hourly-scale meteorological numerical simulation data of boundary layer height, radiation, sea-level pressure, and vertical wind speed, a high-quality large database for Yunnan region to support deep learning is established. An ozone forecasting model based on impact factor analysis and peak perception is established to capture the global dependence relationships in historical meteorological, pollutant data and long time series of ozone, as well as the non-linear relationships between ozone and impact factors, enabling the model to better understand context information and relevant mechanisms affecting ozone generation, solving the defects of long time series dependence of traditional deep learning methods and ignoring the influence of impact factors on predicting ozone concentration, and improving the accuracy of ozone forecasting; by introducing the peak perception method, the accuracy of forecasting the occurrence time and concentration magnitude of ozone peaks is improved, the forecasting effect of ozone heavy pollution processes is optimized, and a fast and accurate forecasting model applicable to ozone heavy pollution processes in Yunnan region is established, providing more detailed scientific support for formulating dynamic and refined ozone peak reduction prevention measures in Yunnan Province. Brief Description of the Drawings
[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0038] Figure 1 It is a flowchart of the ozone concentration forecasting method based on peak perception and impact factor analysis according to the present invention. Detailed Embodiments
[0039] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] Embodiment 1
[0041] An ozone concentration forecasting method based on peak perception and impact factor analysis includes the following steps:
[0042] Step 1: Collect and organize pollutants and their meteorological monitoring data in Yunnan region to establish a large database of ozone pollution, specifically as follows:
[0043] The accumulated data on O in Yunnan since 2015 were obtained by downloading online data. 3 , NO 2 , CO, SO 2 , PM 2.5 , PM 10 Hourly continuous observation data. Currently, hourly observation data from 86 meteorological observation stations can be obtained. The observed meteorological elements include temperature, wind speed, wind direction, total cloud cover, relative humidity and atmospheric pressure.
[0044] In addition, the present invention selects ECMWF numerical forecast products with a resolution of 0.125°×0.125°, and the meteorological data set includes temperature, wind speed, wind direction, total cloud cover, relative humidity and atmospheric pressure. On the basis of collecting and collating historical ground observation data and numerical forecast product output boundary layer height, radiation, sea level pressure, vertical wind speed meteorological numerical simulation data, through data outlier processing, null value filling, data normalization and data set division, so that it meets the data format required for deep learning, a large database of pollutants and meteorological data in Yunnan area for supporting deep learning is established.
[0045] Data outlier processing: Abnormal equipment damage and other reasons may cause the monitoring data to be too large, or equipment insensitivity may cause the monitoring data to be continuously 0 for a period of time. This part of the data content in the conventional monitoring pollutant data and meteorological data is eliminated.
[0046] Data completion: Due to reasons such as observation equipment or observation conditions, the phenomenon of discontinuity of observation data is obvious, so it is very important to supplement missing values. The data completion method adopted in the present invention is a completion method based on interpolation of statistical values at different time scales.
[0047] V ij =|V monthij +V weekdayij +V hourij -2×V yearij |
[0048] Among them, V ij is the filling value of pollutant i at missing time j, Vmonth ij is the mean concentration of pollutant i in the month of missing time j, Vweekday ij is the mean concentration of pollutant i in the week of missing time j, Vyeay ij is the mean pollutant concentration of pollutant i in the year of missing time j.
[0049] Data normalization: Data normalization is beneficial to improving the stability of training and the convergence speed of model training. The normal distribution normalization method is adopted in the present invention. Specifically, the characteristic attributes of the data are subtracted by the mean value and then divided by the variance, which is transformed into a standard normal distribution with a mean value of 0 and a variance of 1.
[0050] x = (x - μ) / σ
[0051] Where μ is the data mean value, σ is the data variance. The mean value and variance on the training set are recorded. After standardizing the training set, when using the model for subsequent testing, the variance and mean value on the training set are used to standardize the test data.
[0052] Dataset division: The data adopted in the present invention are the hourly continuous observation data of more than a hundred ground observation stations in the monitoring network in Yunnan region starting from 2015 for O 3 , NO 2 , CO, SO 2 , PM 2.5 , PM 10 .
[0053] Select the data of the recent one year as the test set to evaluate the model ability. The other data are divided into the training set and the validation set according to the data volume ratio of 4:1. The training set is used to train the model parameters, and the validation set is used to select the model with the best number of training rounds.
[0054] Step 2: Use the Transformer model based on impact factor analysis and peak perception optimization to train the historical observation dataset in the large database of pollutants and meteorological data in Yunnan region, establish the ozone prediction model for Yunnan region of historical meteorology, pollutants and future ozone, and optimize the ozone prediction model for Yunnan region based on deep learning, specifically as follows:
[0055] Use the Transformer model based on impact factor analysis to train the historical observation dataset. Use the meteorological and air pollutant monitoring data of more than ten years as the input of the neural network model, and the predicted ozone concentration as the output. Respectively establish the ozone prediction models for Yunnan region on the hourly scale and the daily scale, and output the influence degree scores of the impact factors on the predicted ozone concentration of the output. Select the largest n influence degree scores, and take the corresponding impact factors as the n - main control factors.
[0056] At the same time, apply the peak perception method to the ozone concentration on the daily scale and the hourly scale respectively, and increase the penalty degree of the model at the corresponding peak values, so that the model can better perceive the peak time and the corresponding concentration size, improve the ozone peak prediction effect, and realize the rapid and accurate prediction of the future ozone heavy pollution process in Yunnan region; thus, the optimization of the ozone prediction model for Yunnan region can be completed.
[0057] A. Principle of Model Selection
[0058] The self-attention mechanism is used to train the historical observation dataset, accurately extract the context features between time series, establish the corresponding temporal context dependencies, and through channel information interaction, accurately establish the temporal dependencies between ozone, meteorology, and other pollutants.
[0059] B. Model to be Established
[0060] Build a scientific and effective neural network model. Set the number of encoding layers and decoding layers to 1, 2, 3, 4..., the number of multi-head attention heads to 8, the training batch size to 256, select the L2 regularization algorithm, and the input data lengths to 120, 96, 72, 48, 24. Continuously reduce the learning rate of the optimizer algorithm as the number of training rounds increases; on this basis, optimize and adjust the model parameters to improve the accuracy of ozone forecasting;
[0061] The backbone network of the encoding layer is the key network for extracting time series features, and its performance plays an important role in the network's ability to capture time series features; the preset backbone network in the present invention is the encoding network of the Transformer model, which accurately extracts features related to time series, and the decoder provides effective information for sequence prediction.
[0062] C. Model Training Optimization
[0063] The present invention will optimize the model training through the following process:
[0064] I) First, divide the ozone monitoring dataset into a training set, a validation set, and a test set. To test the applicability of the model, select the months with heavy ozone pollution in Kunming, Yunnan Province (March - May) as the test set of the model;
[0065] II) Since the units of meteorological data and air quality data are different and the magnitude differences are large, in order to reduce the model training error, the present invention will use the normal distribution normalization method to eliminate the influence brought by the dimension difference of pollutants and meteorological data;
[0066] III) Select a batch size of 256 and use the mini-batch gradient descent method to optimize the model training. In addition, to ensure the randomness of each batch of data selection, when using the mini-batch gradient descent method, shuffle the training set;
[0067] IV) The model training determination condition set by the present invention is: when the maximum number of training rounds is reached or the loss function value of the validation set has not decreased for 10 consecutive rounds, it is considered that the condition is met and training stops to save the parameters, otherwise, return to continue the next round of training;
[0068] V) Finally, to scientifically evaluate the prediction accuracy of the model, the root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R 2 ) are selected as the evaluation parameters of the model.
[0069] Step 3: Use the evaluation results of the optimized ozone prediction model for the Yunnan region in the validation set as an indicator to evaluate the quality of the deep learning model, and conduct ozone concentration evaluation;
[0070] The evaluation indicators include mean square error MSE, mean absolute error MAE, and coefficient of determination R 2 , and select the evaluation index
[0071] Metric = (MSE + MAE) / 2 - R 2 as the evaluation index.
[0072] Select the model with the lowest Metric on the validation set as the optimal model and use it for fast and accurate prediction.
[0073] Example 2
[0074] Specific implementation steps of the Transformer model optimized based on impact factor analysis and peak perception:
[0075] 1. Calculate channel features
[0076] First, perform preprocessing and embedding operations on the multivariate time series data of the input combined climate data and pollutant data; specifically:
[0077] The multivariate time series data of the input combined climate data and pollutant data is denoted as where: B represents the batch size; L represents the length of the input time series; num factors represents the number of types of climate and pollutants. The variables here are meteorological observation data, meteorological numerical simulation data, and pollutant observation data. Specifically, they are: historical O 3 , CO, SO 2 , NO 2 , PM 2.5 , PM 10、 height, radiation, sea level pressure, vertical wind speed, temperature, wind speed, wind direction, total cloud cover, relative humidity, and atmospheric pressure.
[0078] Map the data to a high-dimensional feature space through the embedding layer, and calculate the channel features:
[0079] Feature channel = Embedding value (x enc ) + Embedding position (xenc );
[0080] in, Embedding value Map the original numerical data x to a high-dimensional feature space to capture the numerical information of the data; Embedding position Captures position information of each position in the input sequence, helping the model understand the relative position of elements in the sequence.
[0081] 2. Extract global features of channels
[0082] 2.1 Perform average pooling on the features of each channel in the time dimension to obtain the global feature representation of each channel: MFeature channel =Mean(Feature channel , dim=-1) shape is
[0083] This process compresses the time series information into a fixed-length vector, thereby providing a simplified feature representation for subsequent impact factor analysis attention calculation.
[0084] 3. Conduct impact factor analysis and calculate impact scores
[0085] 3.1 Using the channel attention mechanism, calculate the attention score of each channel. The specific steps are:
[0086] Score channel =Attention channel (MFeature channel ),
[0087] Among them, the impact score is Here, Attention channel It consists of two fully connected layers. The first fully connected layer maps the input, then passes through the activation function layer (ReLU), then passes through the second fully connected layer, and finally uses Softmax for normalization to output the attention score of each channel. channel The attention score here is the impact score.
[0088] 3.2 To identify the quantitative impact of different factors on ozone forecasting, select Score channel The largest n impact scores have their corresponding impact factors taken as n-main controlling factors.
[0089] 4. Expand channel scores to match input dimensions
[0090] The channel scores are dimensionalized so that they are weighted with the input data.
[0091] ExpandedScore channel = Score channel .unsqueeze(1).repeat(1,L,1)
[0092] The influence degree score after expanding the dimension is Use unsqueeze(1) to add a dimension after the first dimension of the channel score; use repeat(1,L,1) to duplicate the channel score in the time dimension to ensure it matches the time dimension of the input data.
[0093] 5. Weight each channel of the input data
[0094] Perform weighted processing on each channel of the input data:
[0095] x enc_weighted = Encoder data (x enc ,x mark ) × ExpandedScore channel ,
[0096] where the embedding layer Encoder data (x enc ,x mark ) is calculated as:
[0097] Encoder data (x enc ,x mark ) = Embedding value (x enc ) +
[0098] Embedding position (x enc ) + Embedding temporal (x mark );
[0099] Embedding value maps the original numerical data x enc to a high-dimensional feature space to capture the numerical information of the data.
[0100] Embedding temporal extracts time-related information, such as timestamps, dates, or periodic features.
[0101] Embedding position captures the positional information of each position in the input sequence to help the model understand the relative positions of the elements in the sequence.
[0102] Multiply the weighted channel scores element-wise with the original input data. Channels with high importance have increased influence on the model due to their higher attention scores, while channels with lower attention scores have reduced influence.
[0103] 6. Encoder part
[0104] 6.1. Embed the weighted input data and prepare to enter the encoder:
[0105] ENc out = Encoder input (x enc , x mark )
[0106] where the network architecture of Encoder input is the same as Encoder data ; Enc out has a shape of [B, L, D], and x mark is the time stamp of the input sequence;
[0107] 6.2. Input the embedded data into the encoder to extract high-level feature representations:
[0108] TEnc out = TransformerEncoder(Enc out , AttentionMask = None)
[0109] The encoder adopts the Transformer architecture. TransformerEncoder consists of two layers and can effectively extract the context information of the input sequence. Each layer contains a multi-head self-attention mechanism and a feed-forward neural network, and uses residual connections and layer normalization to enhance the training effect of the model.
[0110] 7. Decoder part
[0111] 7.1. Embed the input data of the decoder:
[0112] Dec out = Decoder input (x decode , x markdec )
[0113] where the network architecture of Decoder input is the same as Encoder data ; x decode is a sequence of all zeros, and its dimension is x markdec is the time stamp of the prediction sequence, and Decout Decoding feature Dec of the decoder out ∈R B×L‘×D 。L‘ represents the length of the decoder sequence.
[0114] 7.2. Input the embedded data together with the output of the encoder into the decoder for decoding operations
[0115] TDec out =TransformerDecoder(TEnc out )
[0116] where TDec out ∈R B×L‘×D , the decoder adopts an autoregressive Transformer architecture to generate a prediction sequence. Through the self-attention mechanism, the decoder can utilize the context information output by the encoder to generate the output of the current time step and guide the subsequent generation process through the historical output information.
[0117] 8. Output the prediction result
[0118] 8.1. Convert the output of the decoder into the final prediction result through a linear layer or other mapping
[0119] FinalOut=projection(TDec out ),
[0120] where projection is a linear layer or other mapping function used to map the output of the decoder to the target dimension to obtain the final prediction result FinalOut, FinalOut∈R B×L‘×target_dim 。
[0121] In the above steps, B represents the batch size; L represents the length of the time series; num_factors represents the number of climate and pollutant types; D represents the feature space dimension; target_dim represents the dimension of the prediction target, which is 1 in the ozone prediction process.
[0122] 9. Peak-aware loss function
[0123] During the training process, the peak-aware loss function is used to improve the model's ability to perceive peaks, specifically as follows
[0124] 9.1 Use the peak detection algorithm to find the peak position, specifically including:
[0125] PeakIndex=PeakDetection(Target)
[0126] where Target is the true value of the predicted ozone concentration, Target∈RB×L‘ , PeakIndex is the index of the peak, and PeakIndex ∈ {1, 2, 3,..., L'}; PeakDetection is the peak finding algorithm, here it is the Automatic multiscale - based peak detection (AMPD) algorithm;
[0127] 9.2 Calculate the peak - aware loss function and perform gradient backpropagation
[0128] Loss peak = MSELoss(Target PeakIndex , FinalOut PeakIndex )
[0129] where Target PeakIndex , represents the sequence formed by Target i , where i ∈ PeakIndex. Similarly, FinalOut PeakIndex , represents the sequence formed by FinalOut i , where i ∈ PeakIndex;
[0130] Loss MSE = MSELoss(Target, FinalOut)
[0131] Loss total = Loss MSE + aLoss peak
[0132] α in the formula takes 1.
[0133] To sum up, by increasing the penalty degree of the model at the corresponding peaks, the model can better perceive the time of the peaks and the corresponding concentration levels, improve the ozone peak prediction effect, and achieve rapid and accurate prediction of future ozone heavy - pollution processes in Yunnan region; thus, the optimization of the ozone prediction model in Yunnan region can be completed.
Claims
1. A method for predicting ozone concentration based on peak perception and influencing factor analysis, characterized in that: The following steps are involved: S1. Obtain pollutant and meteorological monitoring data in Yunnan for many years through network data collection, and perform outlier processing, null value filling, data normalization and data set division on the pollutant and meteorological monitoring data to build a large database of pollutant and meteorological data in Yunnan that supports deep learning; S2. The Transformer model based on influencing factor analysis and peak perception optimization is used to train the historical observation data set in the pollutant and meteorological data database in Yunnan, and the ozone forecast model in Yunnan based on historical meteorology, pollutants and future ozone is established. Based on big data deep learning, the ozone forecast model in Yunnan is optimized through peak perception and influencing factor analysis, and the n-master control factors are output; S3. The evaluation results of the optimized Yunnan ozone forecast model in the validation set are used as indicators to evaluate the quality of the deep learning model and to evaluate the ozone concentration. The evaluation indicators include MSE, MAE and R 2 , the evaluation index expression is selected as follows: Metric=(MSE+MAE) / 2-R 2 ; The model with the lowest metric on the validation set is selected as the optimal model and used for fast and accurate forecasting; MSE is mean square error, MAE is absolute error, R 2 is the coefficient of determination; In S2 above, based on big data deep learning, the ozone forecast model in Yunnan is optimized through peak perception and influencing factor analysis as follows: S2.
1. The Transformer model based on influencing factor analysis is used to train the historical data set. The meteorological and atmospheric pollutant monitoring data of more than ten years and the numerical simulation data provided by the ECMWF meteorological forecast products are used as the input of the neural network model, and the predicted ozone concentration is used as the output. The ozone forecast model for the Yunnan region at the hourly scale and the daily scale is established respectively; and the perception score is output, and the largest n perception scores are selected, and the corresponding influencing factors are taken as n-main control factors; S2.
2. Simultaneously apply the peak sensing method to find the corresponding peak values in the daily forecast and hourly forecast results using the peak finding algorithm, including: PeakIndex=PeakDetection(Target) Where Target is the true value of the predicted ozone concentration, Target∈R B×L‘ , PeakIndex is the index of the peak, PeakIndex∈{1,2,3,...,L'}; PeakDetection is the peak search algorithm, here is the Automatic multiscale-based peak detection (AMPD) algorithm; Calculate the peak perceptual loss function and perform gradient backpropagation Loss peak =MSELoss(Target PeakIndex ,FinalOut PeakIndex ) Target PeakIndex , indicating that by Target i The sequence of numbers, where i∈PeakIndex; similarly FinalOut PeakIndex , indicating that FinalOut i The sequence of numbers, where i∈PeakIndex; Loss MSE =MSELoss(Target,FinalOut) Loss total =Loss MSE +αLoss peak In the formula, α is 1; Then, the penalty level of the model is increased at the corresponding peak value, so that the model can better perceive the peak time and the corresponding concentration, improve the ozone peak forecast effect, and realize the rapid and accurate forecast of future severe ozone pollution process in Yunnan Province; thus, the optimization of the ozone forecast model in Yunnan Province can be completed.
2. The ozone concentration prediction method based on peak perception and influencing factor analysis according to claim 1 is characterized in that: In S1, the large database of pollutants and meteorological data in Yunnan is established as follows: Through online data download, the six conventional monitoring pollutants (O3, CO, SO2, NO2, PM 2.5 、PM 10 ) Continuous online observation data and meteorological observation data are combined with ECMWF meteorological forecast products to provide hourly-scale meteorological numerical simulation data on boundary layer height, radiation, sea level pressure, and vertical wind speed to collect and organize, and build a long-term historical data set of ozone, meteorology, and other pollutants; and perform data outlier processing, null value filling, data normalization, and data set division on the data set to meet the data format required for deep learning, and establish a large database of pollutants and meteorological data in Yunnan to support deep learning; During the project implementation period, meteorological and pollutant observation data will be continuously stored and sorted to achieve dynamic updating of the large database.
3. The ozone concentration prediction method based on peak perception and influencing factor analysis according to claim 2 is characterized in that: The data outlier processing, null value filling, data normalization and data set division are as follows: Data outlier processing: abnormal reasons such as equipment damage may cause the monitoring data to be too large or the equipment to be insensitive, resulting in the monitoring data being continuously 0 for a period of time. This part of the data content in the conventional monitoring pollutant data and meteorological data is eliminated; Filling in missing values: Due to observation equipment or observation conditions, the phenomenon of discontinuity of observation data is obvious, so it is very important to supplement missing values. The data filling method used in this project is based on the interpolation of statistical values at different time scales: V ij =|V monthij +V weekdayij +V hourij -2×V yearij | Where V ij is the filling value of pollutant i at missing time j, Vmonth ij is the mean concentration of pollutant i in the month of missing time j, Vweekday ij is the mean concentration of pollutant i in the week of missing time j, Vyeay ij is the mean pollutant concentration of pollutant i in the year of missing time j; Data normalization: Data normalization is conducive to improving the stability of training and the convergence speed of model training. This project adopts the normal distribution normalization method, which is to subtract the mean of the characteristic attributes of the data and divide it by the variance; convert it into a standard normal distribution with a mean of 0 and a variance of 1: x=(x-μ) / σ In the formula, μ is the data mean, σ is the data variance. After standardizing the training set by recording the mean and variance on the training set, when the model is used in subsequent tests, the test data is standardized using the variance and mean on the training set. Dataset division: The data used are from more than 100 ground observation stations in the Yunnan monitoring network that began to monitor O3, NO2, CO, SO2, PM2.5 and 3.5% of the total CO2 emissions in 2015. 2.5 , PM 10 Observe data continuously hourly; select data from the past year as the test set to evaluate the model capability, and divide the other data into training set and validation set according to the year in a 4:1 ratio; the training set is used to train the model parameters, and the validation set is used to select the model with the best number of training rounds.
4. The ozone concentration prediction method based on peak perception and influencing factor analysis according to claim 1 is characterized in that: The meteorological data set includes temperature, wind speed, wind direction, total cloud cover, relative humidity, atmospheric pressure, boundary layer height, radiation, sea level pressure, and vertical wind speed.
Citation Information
Patent Citations
Ozone concentration prediction method and device, electronic equipment and medium
CN116702981A
Ozone pollution potential forecasting method based on machine learning
CN118194019A