O3 pollution forecasting method and system based on multiple linear regression

Through the method based on multivariate linear regression, an ozone pollution case library is constructed and a stepwise multivariate linear regression model is established. Combined with the numerical model of atmospheric chemicals, the problem of insufficient accuracy and timelinearity in traditional methods when dealing with nonlinear problems is solved, and ozone pollution forecast with higher accuracy and timeliness is achieved.

CN120124008AActive Publication Date: 2025-06-10SUZHOU METEOROLOGICAL BUREAU

Patent Information

Application Number
CN202510187548.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-10
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The traditional ozone pollution forecasting methods have limitations such as high computational complexity, poor timeliness, and incomplete physical and chemical processes. Especially when dealing with nonlinear problems, the prediction accuracy and applicability are limited.

Method used

The O3 pollution forecast method based on multivariate linear regression is adopted, and the O3 pollution case database is constructed by obtaining and pretreating pollutant concentration site data and reanalyzing data grid data, and analyzing the data, and typing and predicting the cases based on meteorological factors, a step-by-step multivariate linear regression model is established, and the future O3 pollution situation is predicted based on high-resolution atmospheric chemical numerical model.

Benefits of technology

It improves the ability to capture complex relationships between ozone concentration and various influencing factors, effectively supplements the characteristics of ozone generation and emission and transportation, thereby improving the accuracy and timeliness of prediction, and providing strong support for the accurate forecasting and prevention and control of ozone pollution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124008A_ABST
    Figure CN120124008A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of ambient air quality forecasting, in particular to an O3 pollution forecasting method and system based on multiple linear regression. The O3 pollution forecasting method based on multiple linear regression comprises the following steps: S1, acquiring data; s2, preprocessing the data, and constructing an O3 pollution case library; s3, typing the O3 pollution cases according to the meteorological elements; s4, screening proper meteorological elements as forecasting factors for each type of O3 pollution case; s5, establishing a step-by-step multiple linear regression model of meteorological data and pollutant concentration for each type of O3 pollution case, and evaluating a regression effect; and S6, obtaining a future day-by-day weather factor forecast result of the high-resolution atmospheric chemical numerical model, and predicting a future O3 pollution condition by using a multiple linear regression model. According to the O3 pollution event forecasting method provided by the invention, the calculation method is simple, the accuracy of the output forecasting result is high, the O3 pollution forecasting effect is effectively improved, an open research framework is provided for O3 pollution forecasting, and powerful support is provided for related researches such as forecasting and evaluation of O3 combined pollution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ambient air quality forecasting, and specifically to an O 3 pollution forecasting method and system based on multiple linear regression. Background Technique

[0002] With the accelerating advancement of industrialization and urbanization, the problem of air pollution has become increasingly severe, becoming a major obstacle restricting sustainable development. Among them, ozone (O 3 ) as an important air pollutant, its excessive concentration not only poses a threat to human health, but also causes damage to the ecosystem and crops. Therefore, accurately predicting ozone concentration and studying the ozone regional transmission mechanism are of great significance for formulating effective air pollution prevention and control measures and improving air quality.

[0003] Traditional ozone pollution forecasting methods mainly rely on chemical transport models such as WRF-CMAQ. Although these methods can simulate and predict ozone pollution to a certain extent, they have limitations such as high computational complexity, poor timeliness, and incomplete physical and chemical processes. Especially when dealing with nonlinear problems, the prediction accuracy and applicability of these traditional models are limited.

[0004] In recent years, with the rapid development of big data and machine learning technologies, new ideas and methods have been provided for ozone pollution forecasting. Machine learning models, with their advantages such as self-adaptability, high accuracy, and low computational complexity, have shown great potential in dealing with nonlinear problems and complex system predictions. By using a large amount of historical meteorological data and air quality data, machine learning models can more accurately capture the nonlinear characteristics of ozone concentration changes, improving the prediction accuracy and timeliness.

[0005] Multiple linear regression, as a classical statistical analysis method, also has certain application value in the field of ozone pollution forecasting. It analyzes the linear relationship between ozone concentration and meteorological parameters, establishes a prediction model, and realizes the prediction of future ozone concentration. However, multiple linear regression has limitations in dealing with nonlinear problems and complex system predictions, and its prediction accuracy and generalization ability need to be improved. Summary of the Invention

[0006] The purpose of the present invention is to provide an O 3 pollution forecasting method and system based on multiple linear regression.

[0007] In the first aspect, the present invention provides an O 3 pollution forecasting method based on multiple linear regression, specifically including the following steps:

[0008] S1. Obtain data;

[0009] S2. Preprocess the data and construct an O3 pollution case database; 3 Pollution case database;

[0010] S3. Classify the O3 pollution cases according to meteorological elements; 3 Pollution cases are classified;

[0011] S4. Select appropriate meteorological elements as forecasting factors for each type of O3 pollution case; 3 Pollution cases are screened for appropriate meteorological elements as forecasting factors;

[0012] S5. Establish a stepwise multiple linear regression model between meteorological data and pollutant concentration for each type of O3 pollution case, and evaluate the regression effect; 3 Pollution cases are used to establish a stepwise multiple linear regression model between meteorological data and pollutant concentration to evaluate the regression effect;

[0013] S6. Obtain the forecasting results of future daily weather elements from a high-resolution atmospheric chemistry numerical model, and use the multiple linear regression model to predict the future O3 pollution situation. 3 Pollution situation.

[0014] In an embodiment of the present application, in S1, the obtained data includes pollutant concentration station data and reanalysis data grid data; wherein the pollutant concentration station data includes: O3 concentration station data with a time resolution of 1 h; and the reanalysis data grid data includes: boundary layer height, 2 m dew point temperature, 2 m temperature, precipitation, sea level pressure, surface shortwave radiation, cloud cover, geopotential height, relative humidity, specific humidity, radial wind field, zonal wind field, vertical wind field, with a spatial resolution of 0.25°×0.25° and a time resolution of 1 h. 3 Concentration station data, time resolution 1 h; and the reanalysis data grid data includes: boundary layer height, 2 m dew point temperature, 2 m temperature, precipitation, sea level pressure, surface shortwave radiation, cloud cover, geopotential height, relative humidity, specific humidity, radial wind field, zonal wind field, vertical wind field, spatial resolution 0.25°×0.25°, time resolution 1 h.

[0015] In an embodiment of the present application, in S2, the steps of preprocessing the data and constructing an O3 pollution case database include:

[0016] S2.1. Calculate the daily maximum 8-hour moving average of O3 concentration: 3 Daily maximum 8-hour moving average of concentration:

[0017]

[0018] Wherein, is the 8-hour average of O3 concentration at a certain station starting from the i-th time of the current day; O3(i) is the O3 concentration value at the i-th time of the current day at a certain station; n1 is the amount of valid data. When the number of valid data within this 8-hour period is less than 6, it is not included in the case; MDA8O3 is the daily maximum 8-hour moving average of O3 concentration at a certain station, and n2 is the amount of valid data. 3 8-hour average starting from the i-th time of the current day; O3 3 (i) is the O3 concentration at a certain station 3 concentration value at the i-th time of the current day; n1 is the amount of valid data. When the number of valid data within this 8-hour period is less than 6, it is not included in the case; MDA8O3 3 is the daily maximum 8-hour moving average of O3 concentration at a certain station, and n2 is the amount of valid data; 3 Daily maximum 8-hour moving average of concentration;

[0019] S2.2. Calculate the daily maximum 2m temperature, daytime boundary layer height, daytime average temperature, daytime surface shortwave radiation, and daytime cloud cover;

[0020] S2.3. Screen for cases where the daily maximum 8-hour moving average of O 3 concentration exceeds 160 μg / m -3 to form a pollution case library. 3 Pollution cases

[0021] In an embodiment of the present application, in S3, the self-organizing mapping method is used for classification, and the steps include:

[0022] S3.1. For a given region, standardize the 850hPa height field grid data:

[0023]

[0024] where Z is the original data of the 850hPa height field; μ is the average value of the training data; σ is the standard deviation of the training data; is the standardized data of the 850hPa height field;

[0025] S3.2. Establish and initialize the SOM network, including: determining the number of weather types as needed to set the number of nodes m in the output layer, establishing the initial winning neighborhood and the initial value of the learning rate α, and randomly assigning a weight vector to each node in the output layer, that is, each grid node j has a weight vector with the same dimension as the input feature vector;

[0026] S3.3. Find the winning node, including: for the randomly sampled input sample Z, find the node with the closest distance by calculating the difference between the input sample and the weight vectors of all nodes. This node is the winning node, that is, the "best matching unit" BMU, and its index is b. The BMU needs to satisfy:

[0027] BMU = argmin j ||Z - W j ||

[0028] where Z is the current sample vector, that is, the standardized data of the 850hPa height field, W i is the weight vector of each node, ||·|| is the Euclidean distance EDU, and the EDU needs to satisfy:

[0029]

[0030] where n is the number of neurons in the input layer, that is, the sample size of the 850hPa height field;

[0031] S3.4. Update the weight vectors of the BMU and each node in its neighborhood to make them approach the sample Z. The weight update needs to satisfy:

[0032]

[0033] where t is the current iteration variable, and W j (t) is the weight of node j at the t-th iteration, α(t) is the learning rate at the t-th iteration, h(t) is the neighborhood function value. When choosing the Gaussian neighborhood function, the following conditions need to be satisfied:

[0034]

[0035] where h is the neighborhood function value, representing the updated weight of the nodes near the BMU, ║r b -r j ║ is the Euclidean distance of the grid coordinates between the BMU and the surrounding nodes, and σ is the neighborhood width.

[0036] S3.5. Repeat the above process several times until the training is completed.

[0037] In an embodiment of the present application, in S4, the steps of screening appropriate meteorological elements as forecasting factors for each type of O 3 pollution case include:

[0038] S4.1. Calculate the Pearson correlation coefficient R between the MDA8 O 3 concentration of each type of pollution case and the candidate meteorological variables in the reanalysis data grid data. The meteorological variables that are statistically significant at the 95% confidence level are retained as candidate meteorological characteristic factors for the next selection step, and the following conditions need to be satisfied:

[0039]

[0040] where and are the means of the related time series x and y;

[0041] S4.2. Further screen the meteorological characteristic factors according to the eigenvalue importance using the gradient boosting decision tree algorithm.

[0042] In an embodiment of the present application, in S5, the steps of establishing a stepwise multiple linear regression model of meteorological data and pollutant concentration for each type of compound pollution case and evaluating the regression effect include:

[0043] S5.1. Construct a stepwise multiple linear regression model MLR, which needs to satisfy:

[0044]

[0045] where is the pollutant concentration predicted by the MLR, and b 0 represents the intercept term, bk denotes the regression coefficient, M k denotes the predictor, and the total number is N;

[0046] According to the Akaike information criterion, add or delete predictors to perform stepwise regression, which needs to satisfy:

[0047]

[0048] where SSE is the sum of squared errors between the pollutant concentration C and the MLR predicted pollution concentration ; T is the number of days, and N is the number of predictors used in the regression;

[0049] The relative weight of each predictor on the pollutant concentration is measured by its relative contribution to the total explained variance of the multiple linear regression, which needs to satisfy:

[0050]

[0051] where w k is the weight of the predictor, N is the number of predictors, b k denotes the regression coefficient, S k denotes the standard deviation of the predictor i, and S 0 denotes the standard deviation of the pollutant concentration;

[0052] S5.2. Evaluate the regression effect; conduct a significance test on the overall regression equation to determine whether all independent variables have a significant linear relationship with the dependent variable, which is completed using the F-test. The F statistic is the ratio of the regression sum of squares to the error sum of squares, and it needs to satisfy a 95% significance level:

[0053]

[0054] In an embodiment of the present application, in S6, the method for obtaining the forecast results of future daily weather elements from a high-resolution atmospheric chemical numerical model and predicting the future O 3 pollution situation includes: obtaining the daily meteorological data of the required predictors and inputting them into the determined multiple linear regression equation to calculate the future pollutant concentration of the city.

[0055] In an embodiment of the present application, the O 3 pollution forecasting method further includes: S7. Conduct a review and summary of the O 3 pollution process and improve the case library.

[0056] In an embodiment of the present application, in step S7, the content of the review and summary includes: pollution facts, pollution forecasts, and forecast verification results.

[0057] Second aspect, the present invention provides an O based on the multiple linear regression as described above 3 pollution forecasting method-based multiple linear regression O 3 pollution forecasting system, including:

[0058] Data acquisition module, used to acquire data;

[0059] Pollution case library construction module, used to preprocess data and construct an O 3 pollution case library;

[0060] Classification module, used to classify O 3 pollution cases according to meteorological elements;

[0061] Forecasting factor screening module, used to screen appropriate meteorological elements as forecasting factors for each type of O 3 pollution case;

[0062] Model construction module, used to establish a stepwise multiple linear regression model between meteorological data and pollutant concentration for each type of O 3 pollution case and evaluate the regression effect;

[0063] O 3 Pollution situation forecasting module, used to obtain the forecasting results of future daily weather elements by a high-resolution atmospheric chemistry numerical model, and use the multiple linear regression model to predict the future O 3 pollution situation;

[0064] Review and improvement module, used to conduct O 3 review and summary of pollution processes and improve the case library.

[0065] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0066] The O based on multiple linear regression of the present invention 3 pollution forecasting method uses long-term station data and reanalysis data with high spatio-temporal resolution, combines the non-linear processing ability of machine learning and the stability of multiple linear regression, can more accurately capture the complex relationship between ozone concentration and various influencing factors, can effectively supplement the ozone generation, consumption and transport characteristics, thereby improving the prediction accuracy and timeliness; at the same time, this method can also comprehensively consider more types of data, providing strong support for the accurate forecasting and prevention and control of ozone pollution. Description of the Drawings

[0067] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the accompanying drawings required for the description of the specific embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0068] Figure 1 is a schematic diagram of the O pollution prediction method based on multiple linear regression of the present invention; 3

[0069] Figure 2 is the actual situation result of MDA8 O and the number of polluted days in the Yangtze River Delta region from 2015 to 2023 in an embodiment of the present invention; 3

[0070] Figure 3 is the weather typing result of the Yangtze River Delta region in 2019 in an embodiment of the present invention. The filled color is the 850hPa height field, and the arrow is the wind field, and its color represents temperature;

[0071] Figure 4 is a line chart showing the change of MDA8 O over time from July 24 to July 31, 2020 in the Yangtze River Delta region in an embodiment of the present invention. The solid line represents the observed value, and the dotted line represents the predicted value. 3 Specific Embodiments

[0072] The following will elaborate on the technical solutions of the present invention in detail in combination with the accompanying drawings and specific implementation cases. It should be noted that the embodiments described here are only for illustrative purposes and do not limit the present invention.

[0073] In one embodiment, as Figure 1 shown, the O pollution prediction method based on multiple linear regression includes the following steps: 3

[0074] S1. Obtain data;

[0075] S2. Preprocess the data and construct an O pollution case library; 3

[0076] S3. Classify the O pollution cases according to meteorological elements; 3

[0077] S4. Select appropriate meteorological elements as prediction factors for each type of O pollution case; 3

[0078] S5. For each type of O 3 ​​​​​​​For pollution cases, a stepwise multiple linear regression model of meteorological data and pollutant concentration is established to evaluate the regression effect;

[0079] S6. Obtain the forecast results of future daily weather elements from a high-resolution atmospheric chemistry numerical model, and use the multiple linear regression model to predict the future O 3 pollution situation.

[0080] Specifically, in the above S1, the data obtained includes pollutant concentration station data and reanalysis data grid data. The pollutant concentration station data includes: O 3 concentration station data. The stations can include 391 stations across the country, with a time resolution of 1 h; the reanalysis data grid data can include: boundary layer height, 2 m dew point temperature, 2 m temperature, precipitation, sea level pressure, surface shortwave radiation, cloud amount, geopotential height, relative humidity, specific humidity, radial wind field, zonal wind field, vertical wind field, with a spatial resolution of 0.25°×0.25° and a time resolution of 1 h. Optionally, the data can be the fifth-generation global weather and climate reanalysis data ERA5 from the European Centre for Medium-Range Weather Forecasts (ECMWF).

[0081] Furthermore, in the above S2, the steps of preprocessing the data and constructing the O 3 pollution case library include:

[0082] O 3 The daily maximum 8-hour moving average (MDA8 O 3 ) of the concentration exceeds 160 μg m -3 .

[0083] S2.1. Calculate the O 3 daily maximum 8-hour moving average of the concentration. The formula is as follows:

[0084]

[0085] Among them, is the 8-hour average value of the O 3 concentration starting from the i-th time of the day at a certain station. O 3 (i) is the O 3 concentration value at the i-th time of the day at a certain station. n1 is the amount of valid data. When the amount of valid data within this 8-hour period is less than 6 times (n1 < 6), it is not included in the case. MDA8 O 3 is the daily maximum 8-hour moving average of the O 3 concentration at a certain station, and n2 is the amount of valid data.

[0086] S2.2. Calculate the daily maximum 2m temperature, boundary layer height during the day (8:00-17:00), average temperature during the day (8:00-17:00), surface shortwave radiation during the day (8:00-17:00), and cloud cover during the day (8:00-17:00);

[0087] S2.3, Screening O 3 The maximum 8-hour moving average of the daily concentration exceeds 160 μg m -3 O 3 Pollution cases constitute the pollution case library.

[0088] Output O 3 Pollution actual data set and forecast data set, including O 3 Location, duration and process of the pollution incident MDA8 O 3 See also Figure 2 , which is an example of the MDA8 in the Yangtze River Delta region from 2015 to 2023. 3 The actual results of pollution days (OPdays) are shown in Figure 2. MAM: spring; JJA: summer; SON: autumn; DJF: winter.

[0089] Furthermore, in S3, a self-organizing map (SOM) method is used to perform weather classification, and the steps include:

[0090] S3.1. For a given area, the 850 hPa height field grid data is standardized using the following formula:

[0091]

[0092] Where Z is the original data of the 850 hPa height field (the eigenvalue of the individual sample); μ is the mean value of the training data (each column of eigenvalues); σ is the standard deviation of the training data (each column of eigenvalues); It is the standardized data of 850hPa height field;

[0093] S3.2, establish and initialize the SOM network, including: set the number of nodes m of the output layer (i.e., determine the number of weather types as needed), establish the initial winning neighborhood and the initial value of the learning rate α, and randomly assign a weight vector to each node of the output layer, i.e., each grid node j has a weight vector with the same dimension as the input feature vector;

[0094] S3.3, finding the winning node, including: for a randomly sampled input sample Z, by calculating the weight vector difference between the input sample and all nodes, find the node closest to it, which is the winning node, that is, the "best matching unit" BMU, whose index is b, and the BMU calculation formula is as follows:

[0095] BMU = argmin j ║Z - W j ║

[0096] where Z is the current sample vector, i.e., the standardized data of the 850 hPa height field, and W i is the weight vector of each node, and ||·|| is the Euclidean distance EDU. The calculation formula of EDU is as follows:

[0097]

[0098] where n is the number of neurons in the input layer, i.e., the sample size of the 850 hPa height field.

[0099] S3.4. Update the weight vectors of the BMU and each node in its neighborhood to make them approach the sample Z. The weight update needs to satisfy:

[0100]

[0101] where t is the current iteration variable, and W j (t) is the weight of node j at the t-th iteration, α(t) is the learning rate at the t-th iteration, and h(t) is the value of the neighborhood function. Selecting the Gaussian neighborhood function, it needs to satisfy:

[0102]

[0103] where h is the value of the neighborhood function, representing the updated weight of the nodes near the BMU, and ║r b - r j ║ is the Euclidean distance of the grid coordinates between the BMU and the surrounding nodes, and σ is the neighborhood width.

[0104] S3.5. Repeat the above process (random sampling, winning node search, weight update) several times until the training is completed. Generally, the termination condition of training: the learning rate is less than 0 or the specified number of iterations is reached.

[0105] Figure 3 is the weather classification result of the Yangtze River Delta region in 2019 for an embodiment. The coloring is the 850 hPa height field, the arrows are the wind fields, and their colors represent the temperature.

[0106] Furthermore, in the above S4, the steps of screening appropriate meteorological elements as forecasting factors for each type of O 3 pollution case include:

[0107] S4.1. Calculate the MDA8 O of each type of pollution case 3The Pearson correlation coefficient R between the concentration and the candidate meteorological variables in the above reanalysis data grid points. Meteorological variables with statistical significance at the 95% confidence level are retained as candidate meteorological characteristic factors for the next selection step. The formula is as follows:

[0108]

[0109] Where, and are the means of the correlated time series x and y.

[0110] S4.2. Further screen the meteorological characteristic factors according to the eigenvalue importance using the gradient boosting decision tree algorithm method.

[0111] In S5, for each type of compound pollution case, a stepwise multiple linear regression model of meteorological data and pollutant concentration is established. The steps to evaluate the regression effect include:

[0112] S5.1. Construct a stepwise multiple linear regression model MLR, the formula is as follows:

[0113]

[0114] Where, is the MLR predicted pollutant concentration, b 0 represents the intercept term, b k represents the regression coefficient, M k represents the predictor, and the total number is N.

[0115] Add or delete predictors according to the Akaike information classification statistic for stepwise regression. The formula is as follows:

[0116]

[0117] Where, SSE is the sum of squared errors between the pollutant concentration C and the MLR predicted pollution concentration T is the number of days, and N is the number of predictors used in the regression.

[0118] The relative weight of each predictor on the pollutant concentration is measured by its relative contribution to the total explained variance of the multiple linear regression. The formula is as follows:

[0119]

[0120] Where, w k is the weight of the predictor, N is the number of predictors, b k represents the regression coefficient, S k represents the standard deviation of the predictor i, and S 0 represents the standard deviation of the pollutant concentration.

[0121] S5.2. The significance test of the overall regression equation is used to determine whether all independent variables have a significant linear relationship with the dependent variable. It is usually completed using the F test. The F statistic is the ratio of the regression sum of squares to the error sum of squares and must meet the 95% significance level. The formula is as follows:

[0122]

[0123] Furthermore, in S6, the forecast results of future daily weather elements using high-resolution atmospheric chemistry numerical models are obtained, and the future O 3 Pollution situation, obtain daily meteorological data of the required predictive factors, and input them into the determined multiple linear regression equation to calculate the future pollutant concentrations in the city.

[0124] Figure 4 For example, the MDA8 in the Yangtze River Delta region from July 24, 2020 to July 31, 2020 3 A line graph showing changes over time, with the solid line representing the observed values ​​and the dotted line representing the predicted values.

[0125] Furthermore, the O 3 The pollution forecasting method further includes: S7, performing 3 Review and summarize the pollution process and improve the case library. Specifically, combine the prediction results with the actual situation to conduct 3 Review and summarize the pollution process and improve the case library. The review and summary content includes: actual pollution situation (temporal and spatial distribution characteristics); pollution forecast (temporal and spatial distribution characteristics); forecast verification results.

[0126] In summary, the present invention adopts long-term site data and reanalysis data with high temporal and spatial resolution, combines the nonlinear processing capabilities of machine learning and the stability of multivariate linear regression, and can more accurately capture the complex relationship between ozone concentration and various influencing factors, and can effectively supplement the ozone generation, elimination and transport characteristics, thereby improving the accuracy and timeliness of the prediction; at the same time, this method can also comprehensively consider more types of data, and provide strong support for the accurate prediction and prevention and control of ozone pollution.

Claims

1. A method for predicting O3 pollution based on multiple linear regression, characterized in that: The following steps are involved: S1. Obtain information; S2. Preprocess the data and build an O3 pollution case database; S3. Classify O3 pollution cases based on meteorological factors; S4. Select appropriate meteorological elements as predictors for each type of O3 pollution case; S5. Establish a stepwise multiple linear regression model between meteorological data and pollutant concentration for each type of O3 pollution case and evaluate the regression effect; S6. Obtain the forecast results of future daily weather elements using high-resolution atmospheric chemistry numerical models, and use multivariate linear regression models to predict future O3 pollution.

2. The O3 pollution forecasting method according to claim 1, characterized in that: In said S1, the acquired data include pollutant concentration station data and reanalysis data grid data; in The pollutant concentration site data includes: O3 concentration site data, with a time resolution of 1h; and The grid data of the reanalysis data include: boundary layer height, 2m dew point temperature, 2m temperature, precipitation, sea level pressure, surface shortwave radiation, cloud cover, potential height, relative humidity, specific humidity, radial wind field, zonal wind field, vertical wind field, with a spatial resolution of 0.25°×0.25° and a temporal resolution of 1h.

3. The O3 pollution forecasting method according to claim 1, characterized in that: In S2, the steps of preprocessing the data and building an O3 pollution case library include: S2.

1. Calculate the maximum 8-hour sliding mean of the daily O3 concentration: in, is the 8-hour average of the O3 concentration at a certain station from the i-th time of the day; O3(i) is the O3 concentration at a certain station at the i-th time of the day; n1 is the amount of valid data. When the valid data within the 8-hour period is less than 6 times, it is not counted in the case; MDA8 O3 is the maximum 8-hour sliding mean of the O3 concentration at a certain station, and n2 is the amount of valid data; S2.

2. Calculate the daily maximum 2m temperature, daytime boundary layer height, daytime average temperature, daytime surface shortwave radiation, and daytime cloud cover; S2.

3. Screening O3 concentration with a maximum 8-hour sliding mean value exceeding 160 μg m -3 The O3 pollution cases constitute the pollution case library.

4. The O3 pollution forecasting method according to claim 1, characterized in that: In S3, the self-organizing mapping method is used for typing, and the steps include: S3.

1. For a given area, standardize the 850 hPa height field grid data: Where Z is the original data of the 850hPa height field; μ is the mean value of the training data; σ is the standard deviation of the training data; It is the standardized data of 850hPa height field; S3.2, establish and initialize the SOM network, including: determine the number of weather types according to the need to set the number of nodes m in the output layer, establish the initial winning neighborhood and the initial value of the learning rate α, and randomly assign a weight vector to each node in the output layer, that is, each grid node j has a weight vector with the same dimension as the input feature vector; S3.3, finding the winning node, including: for a randomly sampled input sample Z, by calculating the weight vector difference between the input sample and all nodes, find the node closest to it, which is the winning node, that is, the "best matching unit" BMU, whose index is b, and the BMU needs to meet the following requirements: BMU=argmin j ||ZW j || Among them, Z is the current sample vector, that is, the standardized data of the 850hPa height field, and W i is the weight vector of each node, ||·|| is the Euclidean distance EDU, and EDU must satisfy: Where n is the number of neurons in the input layer, i.e., the sample size of the 850 hPa height field; S3.

4. Update the weight vectors of the BMU and each node in its domain to make them close to sample Z. The weight update must satisfy: Among them, t is the current iteration variable, W j (t) is the weight of node j at the tth iteration, α(t) is the learning rate at the tth iteration, h(t) is the neighborhood function value, and the Gaussian neighborhood function is selected, which must satisfy: Among them, h is the neighborhood function value, which represents the update weight of the nodes near the BMU, ||r b -r j || is the grid coordinate Euclidean distance between BMU and surrounding nodes, and σ is the neighborhood width. S3.

5. Repeat the above process several times until the training is completed.

5. The O3 pollution forecasting method according to claim 2, characterized in that: In S4, the step of selecting appropriate meteorological elements as forecasting factors for each type of O3 pollution case includes: S4.

1. Calculate the Pearson correlation coefficient R between the MDA8 O3 concentration of each type of pollution case and the candidate meteorological variables in the reanalysis data grid data. Meteorological variables with statistical significance at the 95% confidence level are retained as candidate meteorological characteristic factors for the next selection step, and must meet the following requirements: in, and is the average value of the related time series x and y; S4.

2. Use the gradient boosting decision tree algorithm to further screen meteorological characteristic factors based on the importance of eigenvalues.

6. The O3 pollution forecasting method according to claim 1, characterized in that: In S5, a stepwise multiple linear regression model of meteorological data and pollutant concentration is established for each type of composite pollution case, and the step of evaluating the regression effect includes: S5.1 Constructing a stepwise multiple linear regression model MLR requires the following: in, is the pollutant concentration predicted by MLR, b0 represents the intercept term, b k represents the regression coefficient, M k represents predictors, the total number is N; Add or delete predictors based on the Akaike information classification statistic and perform stepwise regression, which requires: Among them, SSE is the pollutant concentration C and the MLR predicted pollution concentration The sum of squared errors between ; T is the number of days, and N is the number of predictors used in the regression; The relative weight of each predictor on pollutant concentration is measured by its relative contribution to the total explained variance of the multivariate linear regression, which must satisfy: Among them, w k is the weight of the predictor, N is the number of predictors, b k represents the regression coefficient, S k represents the standard deviation of the prediction factor i, S0 represents the standard deviation of the pollutant concentration; S5.

2. Evaluate the regression effect; the significance test of the overall regression equation is used to determine whether all independent variables have a significant linear relationship with the dependent variable. This is done using the F test. The F statistic is the ratio of the regression sum of squares to the error sum of squares and must meet a 95% significance level:

7. The O3 pollution forecasting method according to claim 1, characterized in that: In S6, the method of obtaining the forecast results of future daily weather elements by a high-resolution atmospheric chemistry numerical model and predicting future O3 pollution using a multivariate linear regression model includes: obtaining daily meteorological data of required prediction factors and inputting them into a determined multivariate linear regression equation to calculate the future pollutant concentration of the city.

8. The O3 pollution forecasting method according to claim 1, characterized in that: Also includes: S7. Review and summarize the O3 pollution process and improve the case library.

9. An O3 pollution forecasting system based on multiple linear regression using the O3 pollution forecasting method based on multiple linear regression as claimed in any one of claims 1 to 8, characterized in that: include: A data acquisition module, used for acquiring data; The pollution case library construction module is used to pre-process the data and build an O3 pollution case library; Classification module, used to classify O3 pollution cases according to meteorological factors; Prediction factor screening module, used to screen appropriate meteorological elements as prediction factors for each type of O3 pollution case; Model building module, used to establish a stepwise multivariate linear regression model of meteorological data and pollutant concentration for each type of O3 pollution case and evaluate the regression effect; O3 pollution forecast module, which is used to obtain the forecast results of future daily weather elements using high-resolution atmospheric chemistry numerical models, and use multivariate linear regression models to predict future O3 pollution; The review and improvement module is used to review and summarize the O3 pollution process and improve the case library.

Citation Information

Patent Citations

  • Regional air pollutant concentration prediction method, terminal and readable storage medium

    CN108053071A

  • Air quality secondary forecasting model construction method based on LSTM neural network

    CN114912343A

  • Method for predicting PM2.5 and ozone mixed pollution under high temporal-spatial resolution

    CN116679356A

  • Meteorological dominant-based multi-factor future ozone prediction multi-target super integrated learning method

    CN117933430A

  • Environmental NO2 photolysis rate estimation method based on multivariate data

    CN118351964A

Cited By

  • Ozone concentration prediction method and detection system based on relevance multiple linear regression

    CN121049463A