PDF-based air quality forecast correction method
Through the PDF-based air quality forecast correction method, the problem of large pollutant forecast error in the prior art is solved, and higher forecast accuracy and flexibility are achieved, and a new correction method is provided, suitable for air quality management and emergency response.
Patent Information
- Application Number
- CN202510121607.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-25
- Publication Date
- 2025-05-23
AI Technical Summary
The existing air quality forecasting model has systematic deviations in pollutant forecasting, resulting in large forecast errors and is difficult to meet the actual application needs.
The PDF-based air quality forecast correction method is used to determine the threshold interval of pollutants, calculate the probability density function of the live and forecast, and correct the concentration of the numerical model forecast, and calculate it for different start-up times, forecast timeliness and sites to obtain the final correction result.
It improves the accuracy of air quality forecasting, enhances the flexibility and adaptability of forecasting, optimizes the stability and reliability of forecast results, and provides a new, more accurate and effective correction method.
Smart Images

Figure CN120030271A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of environmental science and technology, and in particular to an air quality forecast correction method based on PDF. Background Art
[0002] In the current field of air quality forecasting, haze forecasting models, mainly CMA-CUACE, play an important role. However, these models face significant challenges in practical applications, especially the large forecast errors for various pollutants. This error not only affects the accuracy of the forecast, but also limits the effectiveness of air quality management and emergency response.
[0003] In order to improve the forecast effect, researchers have been exploring ways to eliminate the systematic bias of air quality forecast models. The traditional approach is to correct numerical forecast products through objective forecast products in order to improve the accuracy of forecasts. However, the application of existing technologies in this field still has many shortcomings.
[0004] The current correction methods mainly focus on the correction of conventional meteorological elements such as temperature and precipitation, while the correction technology for key pollutants in air quality forecasts, such as PM2.5, PM10, sulfur dioxide, nitrogen oxides, etc., is not mature enough. This leads to large errors in air quality forecasts and is difficult to meet the needs of practical applications. The existing correction models are relatively fixed and lack sufficient flexibility and adaptability. With the increasing complexity of atmospheric pollution sources and the continuous changes in meteorological conditions, fixed correction models are often difficult to accurately reflect the actual situation, resulting in inaccurate forecast results. In this regard, we propose an air quality forecast correction method based on PDF. Summary of the invention
[0005] In order to solve the above technical problems, a PDF-based air quality forecast correction method is provided. This technical solution solves the above problems.
[0006] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0007] A PDF-based air quality forecast correction method, the air quality forecast correction method is:
[0008] According to the ambient air quality index, determine the actual and forecast threshold ranges of the pollutants, including PM2.5, PM10 and O3;
[0009] Based on the determined pollutant threshold interval, the frequency of different thresholds of the pollutant observed in the rolling period is calculated, and the actual probability under each pollutant threshold is calculated to obtain several probability points;
[0010] Establish the actual probability density function obs(x) and the predicted probability density function fore(x);
[0011] Substitute the pollutant concentration predicted by the numerical model into fore(x) and calculate the corresponding probability value under this function;
[0012] Based on the calculated probability value, substitute the probability value into obs(x) to find the corresponding concentration interval, and obtain the corrected value by downward interpolation;
[0013] The final correction results are calculated for different reporting times, forecast time limits and stations.
[0014] Preferably, the threshold interval for determining the actual situation and forecast of the pollutant specifically includes:
[0015] The threshold ranges of actual and forecast conditions correspond to the eight levels specified in the Ambient Air Quality Index. The rolling period is determined as the past 20 days, with two time periods, 8:00 and 20:00. For all stations within the forecast range, each station is modeled separately.
[0016] Preferably, the formula for calculating the frequency of occurrence of different threshold values of the pollutant observed in real time within the rolling period is:
[0017]
[0018] Where P i Indicates the frequency of a certain pollutant threshold, A i It indicates the number of stations with a certain level of pollutant threshold, and B indicates the total number of stations.
[0019] Preferably, the actual probability density function obs(x) is established by nonlinear fitting and piecewise function fitting.
[0020] Preferably, the method for establishing the actual probability density function obs(x) by nonlinear fitting and piecewise function fitting is:
[0021] Among them, the probability density function form of the nonlinear function is:
[0022]
[0023] In the formula, p(x) represents the probability density function, K represents the number of Gaussian distributions, and w i represents the weight of the i-th Gaussian distribution, satisfying μ i represents the mean of the i-th Gaussian distribution, represents the standard deviation of the i-th Gaussian distribution;
[0024] Use maximum likelihood estimation to determine the parameter w in the above functioni , μ i and σ i :
[0025] Let the observed data be x 1 ,x 2 ,…,x N , the likelihood function L is:
[0026]
[0027] The parameters are solved by maximizing the likelihood function L and using the EM algorithm to iteratively solve:
[0028] In which, step E calculates each data point x j The probability of belonging to the i-th Gaussian distribution is γ ij :
[0029]
[0030] Among them, the M-step update parameter w i , μ i and σ i :
[0031]
[0032] Repeat steps E and M until the parameters converge. The function p(x) obtained at this time is the actual probability density function obs(x);
[0033] Among them, the piecewise function fitting method is:
[0034] According to the previously determined pollutant threshold range, the concentration range is divided into M intervals i 1 ,i 2 ,…,I M ;
[0035] For each interval I M , using a linear function for fitting:
[0036] Let interval I m =[a m ,b m ], the linear function form of the fitting in this interval is:
[0037] y=c m x+d m
[0038] Determine the coefficient c by the least squares method m and d m , let the observed data points in this interval be The goal is to minimize the sum of squared errors:
[0039]
[0040] c m and d m Find the partial derivative and set it to 0, and we get the system of equations:
[0041]
[0042] Solve the system of equations to get c m and d m The value of , thus determining the interval I m The fitting function within ;
[0043] Let interval I m The fitting function in is y = e m , determine e by calculating the average probability value of the data points in the interval m ,Right now:
[0044]
[0045] The final actual probability density function obs(x) is:
[0046]
[0047] Through the above nonlinear fitting or piecewise function fitting method, the actual probability density function obs(x) is established.
[0048] Preferably, the method for establishing the forecast probability density function fore(x) is:
[0049] Collect data on several pollutant forecasts and their corresponding frequencies and organize them into data pairs (x i ,y i );
[0050] For nonlinear fitting, the function form is chosen based on preliminary observations of the data, and a least squares problem is solved using a numerical optimization algorithm to determine the parameters of the function;
[0051] For piecewise function fitting, divide the data into segments, determine the function form of each segment, solve the least squares problem for each segment, and ensure that the continuity condition is met at the connection points;
[0052] The goodness of fit index is used to evaluate the fitting effect, and the fitting result is obtained through parameter estimation method based on the evaluation result.
[0053] Preferably, the nonlinear fitting and piecewise function fitting specifically include:
[0054] Among them, the nonlinear fitting method is:
[0055] For data that exhibit a bell-shaped curve, the fitted function takes the form:
[0056]
[0057] In the formula, μ represents the mean and σ represents the standard deviation;
[0058] For data with positive skewness, the fitted function takes the form:
[0059]
[0060] Use the least squares method to estimate parameters:
[0061] Assume there are n data points (x i ,y i ), where x i is the predicted pollutant value, y i is the corresponding pollutant forecast frequency, and the sum of squared errors is minimized by the least squares method:
[0062]
[0063] In the formula, θ represents the parameter vector of the function;
[0064] The coefficient of determination R 2 Calculate goodness of fit:
[0065]
[0066] in, R 2 The closer the value is to 1, the better the fitting effect is;
[0067] Among them, the piecewise function fitting method is:
[0068] Divide the data into several non-overlapping segments according to the range of predicted pollutants;
[0069] Within each data segment, the function form is:
[0070] y=f k (x) = a k x+b k
[0071] In the formula, k represents the kth segment;
[0072] For each segment, estimate the parameters using the least squares method:
[0073] Given a data point (x i,k ,y i,k ) In the kth segment, minimize the sum of squared errors:
[0074]
[0075] Ensure that the functions of adjacent segments are continuous at the connection point, that is, for the adjacent kth segment and k+1th segment, at the connection point x = x c,k Where f k (x c,k )=f k+1 (x c,k ).
[0076] Preferably, after calculating the correction value, an uncertainty assessment is performed on the correction result;
[0077] The uncertainty is measured by calculating the standard deviation of the correction results for different onset times, forecast ages and stations:
[0078] Assume that for a particular pollutant, under different reporting times, forecast time periods and stations, the corrected concentration value is C ij , where i represents the start time of the report, j represents the combination identifier of the forecast time and the station, and there are n combinations in total;
[0079] First calculate the average of the corrected concentration values:
[0080]
[0081] In the formula, represents the average value of the corrected concentration values;
[0082]
[0083] The larger the standard deviation, the greater the dispersion of the correction results and the higher the uncertainty; conversely, the smaller the standard deviation, the more stable the correction results and the lower the uncertainty.
[0084] Preferably, time series analysis is introduced to dynamically update the correction results;
[0085] Let y t is the corrected air quality data at time t, which is modeled by the prediction algorithm, where the expression of the prediction algorithm model is:
[0086]
[0087] In the formula, and θ j Represent the autoregressive coefficient and the moving average coefficient, ∈ t represents a white noise sequence, p and q represent the autoregressive order and the moving average order respectively;
[0088] The model parameters are determined by minimizing the residual sum of squares. and θ j , when there is new observation data y t+1When y is predicted based on the established algorithm model t+1 The predicted value of And compare it with the actual observation value and calculate the residual:
[0089]
[0090] In the formula, e t+1 represents the residual value;
[0091] According to the residual size and change trend, the model parameters are dynamically adjusted to update the subsequent correction results.
[0092] Preferably, the model parameters are determined by minimizing the residual sum of squares and θ j The method is:
[0093] Based on the prediction algorithm model, the residual sum of squares function is constructed:
[0094]
[0095] In the formula, S represents the residual sum of squares, represents the predicted value;
[0096] Since the predicted value It's about parameters and θ j function, so S is the parameter and θ j The function of
[0097] To find the parameter value that minimizes S, take the partial derivatives of S with respect to the autoregressive coefficient and the moving average coefficient respectively;
[0098] The gradient descent method is used to iteratively solve the problem until the residual square sum S converges to obtain the model parameters and θ j .
[0099] Compared with the prior art, the present invention has the following beneficial effects:
[0100] The air quality forecast correction method proposed in the present invention realizes the effective correction of the systematic deviation of the air quality forecast by constructing the probability density function of the actual situation and the forecast. Compared with the traditional method, the method not only improves the accuracy of the forecast, but also enhances the flexibility and adaptability of the forecast. Through the in-depth analysis of the actual observation data within the rolling period, the method can more accurately reflect the actual distribution characteristics of pollutants, thereby optimizing the forecast results. It also introduces uncertainty assessment and time series analysis, further improving the stability and reliability of the forecast, and providing a new, more accurate and effective correction method for air quality forecasting, which has important practical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0101] Figure 1 Flowchart of the PDF-based air quality forecast correction method. DETAILED DESCRIPTION
[0102] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are only examples, and those skilled in the art may think of other obvious variations.
[0103] Reference Figure 1 As shown in FIG. 1 , an air quality forecast correction method based on PDF is shown in FIG. 1 . The air quality forecast correction method is:
[0104] According to the ambient air quality index, determine the actual and forecast threshold ranges of the pollutants, including PM2.5, PM10 and O3;
[0105] Based on the determined pollutant threshold interval, the frequency of different thresholds of the pollutant observed in the rolling period is calculated, and the actual probability under each pollutant threshold is calculated to obtain several probability points;
[0106] Establish the actual probability density function obs(x) and the predicted probability density function fore(x);
[0107] Substitute the pollutant concentration predicted by the numerical model into fore(x) and calculate the corresponding probability value under this function;
[0108] Based on the calculated probability value, substitute the probability value into obs(x) to find the corresponding concentration interval, and obtain the corrected value by downward interpolation;
[0109] The final correction results are calculated for different reporting times, forecast time limits and stations.
[0110] It determines the threshold interval according to the provisions of the ambient air quality index, ensuring the scientificity and standardization of the method. By calculating the probability points of different threshold frequencies and probabilities, it can accurately reflect the actual distribution of pollutants under different conditions. The established actual and forecast probability density functions provide intuitive and effective tools for subsequent analysis. The probability values of the predicted concentrations are calculated using numerical models, and then the corrected values are obtained by downward interpolation, making the correction process rigorous and in line with the actual law of change. Calculations are made separately for different reporting times, forecast timeliness and sites, taking into full account the differences in time and space, greatly improving the accuracy and reliability of the forecast, and providing strong support for public life and environmental decision-making.
[0111] The threshold intervals for determining the actual situation and forecast of the pollutant specifically include:
[0112] The threshold intervals of the actual situation and forecast correspond to the eight levels specified in the ambient air quality index. The rolling period is determined to be the past 20 days, with two times at 8:00 and 20:00. For all stations within the forecast range, each station adopts a separate modeling method, corresponding to the eight levels specified in the ambient air quality index, to ensure that it is in line with the general standards and easy to understand and apply. Setting the rolling period to the past 20 days can ensure the timeliness of the data while accumulating enough samples to reflect the changing laws of pollutants. Selecting two times at 8:00 and 20:00 can capture typical conditions at different times of the day and improve the grasp of pollution characteristics. Separate modeling is carried out for all stations within the forecast range, taking into full account the differences in the geographical environment, pollution sources, etc. of each station to avoid generalizations. This comprehensive and detailed setting greatly improves the pertinence and accuracy of the threshold interval, providing a solid foundation for subsequent air quality forecast corrections.
[0113] Taking the forecast of PM2.5 pollutant concentration as an example, according to the PM2.5 concentration limits corresponding to different levels in my country's "Technical Regulations on Ambient Air Quality Index (AQI) (Trial)", 35, 75, 115, 150, 250, 350 and 500 are first determined as initial thresholds. In order to further divide more interval thresholds and avoid the influence of more thresholds on the calculation time, it is finally determined that the PM2.5 actual and forecast values are divided into 11 threshold intervals, namely 10, 35, 50, 75, 95, 115, 135, 150, 200, 250, 350 and 500.
[0114] The formula for calculating the frequency of different threshold values of the pollutant observed in real time during the rolling period is:
[0115]
[0116] Where P i Indicates the frequency of a certain pollutant threshold, A i It indicates the number of stations with a certain level of pollutant threshold, B indicates the total number of stations, and the pollutant threshold frequency is calculated by the ratio of the number of stations to the total number of stations. It can clearly reflect the frequency of different thresholds at each station, provide a quantitative basis for analyzing the distribution characteristics of pollutants in different regions, facilitate rapid grasp of the actual occurrence of each threshold in the rolling cycle, and help to make subsequent more accurate air quality forecast corrections.
[0117] The actual probability density function obs(x) is established by nonlinear fitting and piecewise function fitting.
[0118] The actual probability density function is established through nonlinear fitting and piecewise function fitting, which can flexibly and accurately describe the actual probability distribution of pollutants. The method of establishing the actual probability density function obs(x) through nonlinear fitting and piecewise function fitting is:
[0119] Among them, the probability density function form of the nonlinear function is:
[0120]
[0121] In the formula, p(x) represents the probability density function, K represents the number of Gaussian distributions, and w i represents the weight of the i-th Gaussian distribution, satisfying μ i represents the mean of the i-th Gaussian distribution, represents the standard deviation of the i-th Gaussian distribution;
[0122] Use maximum likelihood estimation to determine the parameter w in the above function i , μ i and σ i :
[0123] Let the observed data be x 1 ,x 2 ,…,x N , the likelihood function L is:
[0124]
[0125] The parameters are solved by maximizing the likelihood function L and using the EM algorithm to iteratively solve:
[0126] In which, step E calculates each data point x j The probability of belonging to the i-th Gaussian distribution is γ ij :
[0127]
[0128] Among them, M steps update parameter w i , μ i and σ i :
[0129]
[0130] Repeat the E and M steps until the parameters converge. The function p(x) obtained at this time is the actual probability density function obs(x). In nonlinear fitting, the function based on Gaussian distribution can capture the complex probability distribution characteristics, and use the maximum likelihood estimation and EM algorithm to iteratively solve the parameters to make the model fit the actual data.
[0131] Among them, the piecewise function fitting method is:
[0132] According to the previously determined pollutant threshold range, the concentration range is divided into M intervals I 1 ,I 2 ,…,I M ;
[0133] For each interval I M , using a linear function for fitting:
[0134] Let interval I m =[a m ,b m ], the linear function form of the fitting in this interval is:
[0135] y=c m x+d m
[0136] Determine the coefficient c by the least squares method m and d m , let the observed data point in this interval be (x m1 ,y m1 ),(x m2 ,y m2 ),…, The goal is to minimize the sum of squared errors:
[0137]
[0138] c m and d m Find the partial derivative and set it to 0, and we get the system of equations:
[0139]
[0140] Solve the system of equations to get c m and d m The value of , thus determining the interval I m The fitting function within ;
[0141] Let interval I m The fitting function in is y = e m , determine e by calculating the average probability value of the data points in the interval m ,Right now:
[0142]
[0143] The final actual probability density function obs(x) is:
[0144]
[0145] Through the above nonlinear fitting or piecewise function fitting method, the actual probability density function obs(x) is established. The piecewise function fitting is divided according to the pollutant threshold interval, each interval is fitted with a linear function, and the coefficient is determined by the least squares method. It can effectively reflect the probability change law of different concentration intervals. The two methods complement each other. Regardless of the complex distribution of the data or the segmented characteristics, the function can be accurately constructed, providing a reliable probability model basis for air quality forecast correction.
[0146] The method to establish the forecast probability density function fore(x) is:
[0147] Collect data on several pollutant forecasts and their corresponding frequencies and organize them into data pairs (x i ,y i );
[0148] For nonlinear fitting, the function form is chosen based on preliminary observations of the data, and a least squares problem is solved using a numerical optimization algorithm to determine the parameters of the function;
[0149] For piecewise function fitting, divide the data into segments, determine the function form of each segment, solve the least squares problem for each segment, and ensure that the continuity condition is met at the connection points;
[0150] The goodness of fit index is used to evaluate the fitting effect, and the fitting result is obtained through parameter estimation method based on the evaluation result.
[0151] By collecting pollutant forecasts and corresponding frequency data to form data pairs, a solid foundation is provided for subsequent analysis. Nonlinear fitting selects the function form based on preliminary observations of the data and uses numerical optimization algorithms to solve it. It can flexibly adapt to complex data distributions and accurately capture potential patterns. Piecewise function fitting divides the data segments and determines the function form of each segment, ensuring continuity at the connection points, and can carefully characterize the characteristics of different intervals. Use goodness of fit indicators to evaluate and optimize the results through parameter estimation to ensure that the model is accurate and reliable. These methods combined can effectively construct a forecast probability density function that fits the actual situation, provide strong support for air quality forecast corrections, and improve forecast accuracy.
[0152] This nonlinear fitting and piecewise function fitting method has outstanding advantages in establishing the forecast probability density function. Nonlinear fitting uses specific function forms for different data forms. There are adaptation functions for bell curves and positively skewed data. Combined with the least squares method to estimate parameters and determine the coefficient to evaluate the goodness of fit, it can accurately characterize the characteristics of complex data and ensure the fitting effect. The nonlinear fitting and piecewise function fitting specifically include:
[0153] Among them, the nonlinear fitting method is:
[0154] For data that exhibit a bell-shaped curve, the fitted function takes the form:
[0155]
[0156] In the formula, μ represents the mean and σ represents the standard deviation;
[0157] For data with positive skewness, the fitted function takes the form:
[0158]
[0159] Use the least squares method to estimate parameters:
[0160] Assume there are n data points (x i ,y i ), where x i is the predicted pollutant value, y i is the corresponding pollutant forecast frequency, and the sum of squared errors is minimized by the least squares method:
[0161]
[0162] In the formula, θ represents the parameter vector of the function;
[0163] The coefficient of determination R 2 Calculate goodness of fit:
[0164]
[0165] in, R 2 The closer the value is to 1, the better the fitting effect is;
[0166] Among them, the piecewise function fitting method is:
[0167] Divide the data into several non-overlapping segments according to the range of predicted pollutants;
[0168] In each data segment, the function form is:
[0169] y=f k (x) = a k x+b k
[0170] In the formula, k represents the kth segment;
[0171] For each segment, estimate the parameters using the least squares method:
[0172] Given a data point (x i,k ,y i,k ) In the kth segment, minimize the sum of squared errors:
[0173]
[0174] Ensure that the functions of adjacent segments are continuous at the connection point, that is, for the adjacent kth segment and k+1th segment, at the connection point x = x c,k Where f k (x c,k ) = f k+1 (x c,k ). Piecewise function fitting divides the data into segments according to the range of predicted pollutants, and each segment is fitted with a dedicated function. The least squares method is used to determine the parameters, and the connection points of adjacent segments are guaranteed to be continuous, presenting the characteristics of different intervals in detail. The combination of the two comprehensively and accurately processes various types of data, provides a reliable means for the construction of air quality forecast probability density functions, and improves the accuracy and scientificity of forecast corrections.
[0175] After calculating the correction value, the uncertainty of the correction result is evaluated;
[0176] The uncertainty is measured by calculating the standard deviation of the correction results for different onset times, forecast ages and stations:
[0177] Assume that for a particular pollutant, under different reporting times, forecast time periods and stations, the corrected concentration value is C ij , where i represents the start time of the report, j represents the combination identifier of the forecast time and the station, and there are n combinations in total;
[0178] First calculate the average of the corrected concentration values:
[0179]
[0180] In the formula, represents the average value of the corrected concentration values;
[0181]
[0182] The larger the standard deviation, the greater the degree of dispersion of the correction results and the higher the uncertainty; conversely, the smaller the standard deviation, the more stable the correction results and the lower the uncertainty. By calculating the standard deviation of the correction results for different reporting times, forecast time periods and stations, the dispersion of the corrected concentration values can be fully reflected. This quantitative indicator intuitively presents the stability of the correction results under the combination of various factors. When the standard deviation is large, it indicates a high degree of dispersion, which means that the correction results under different conditions vary greatly and the uncertainty is high, reminding users to be cautious about the forecast results; while a small standard deviation shows that the results are stable and the uncertainty is low, which enhances the credibility of the forecast results. This assessment provides more comprehensive information for the application of air quality forecasts, helping decision makers to make rational use of forecast data and better respond to environmental changes.
[0183] Introduce time series analysis to dynamically update the correction results;
[0184] Let y tis the corrected air quality data at time t, which is modeled by the prediction algorithm, where the expression of the prediction algorithm model is:
[0185]
[0186] In the formula, and θ j Represent the autoregressive coefficient and the moving average coefficient, ∈ t represents a white noise sequence, p and q represent the autoregressive order and the moving average order respectively;
[0187] The model parameters are determined by minimizing the residual sum of squares. and θ j , when there is new observation data y t+1 When y is predicted based on the established algorithm model t+1 The predicted value of And compare it with the actual observation value and calculate the residual:
[0188]
[0189] In the formula, e t+1 represents the residual value;
[0190] According to the residual size and change trend, the model parameters are dynamically adjusted to update the subsequent correction results.
[0191] The model parameters are determined by minimizing the residual sum of squares. and θ j The method is:
[0192] Based on the prediction algorithm model, the residual sum of squares function is constructed:
[0193]
[0194] In the formula, S represents the residual sum of squares, represents the predicted value;
[0195] Since the predicted value It's about parameters and θ j function, so S is the parameter and θ j The function of
[0196] To find the parameter value that minimizes S, take the partial derivatives of S with respect to the autoregressive coefficient and the moving average coefficient respectively;
[0197] The gradient descent method is used to iteratively solve the problem until the residual square sum S converges to obtain the model parameters and θ j. With the help of the autoregressive moving average model, the parameters are determined by minimizing the residual sum of squares, which can effectively capture the change pattern of air quality data over time. When new data arrives, the residuals are calculated by comparing the predicted values with the actual observed values, and the model parameters are dynamically adjusted according to the size and trend of the residuals, so as to optimize the correction results in real time. This process enables the forecast model to keep pace with the times, constantly adapt to the dynamic changes in air quality, and significantly improve the accuracy of the forecast. At the same time, based on the iterative solution of the gradient descent method, the efficiency and stability of parameter optimization are ensured, providing more reliable and accurate dynamic corrections for air quality forecasts, and assisting environmental management and decision-making.
[0198] The use process of the present invention is: determine the pollutant threshold interval; calculate the actual observation frequency and probability; establish the actual and predicted probability density function; substitute the predicted concentration to obtain the probability value; interpolate to obtain the correction value; and calculate the final correction result.
[0199] In summary, the advantages of the present invention are: by comprehensively considering the data of different reporting times, forecast timeliness and sites, the probability density function of the actual situation and the forecast is established, and the effective correction of the systematic deviation is achieved. In addition, the method is also flexible and adaptable, and can perform personalized modeling according to the actual conditions of different regions, further improving the accuracy of the forecast. The introduction of this patented technology has brought new breakthroughs in the field of air quality forecasting.
[0200] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions only describe the principles of the present invention. The present invention may be subject to various changes and improvements without departing from the spirit and scope of the present invention. These changes and improvements fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the attached claims and their equivalents.
Claims
1. A PDF-based air quality forecast correction method, characterized in that: The air quality forecast correction method is: According to the ambient air quality index, determine the actual and forecast threshold ranges of the pollutants, including PM2.5, PM10 and O3; Based on the determined pollutant threshold interval, the frequency of different thresholds of the pollutant observed in the rolling period is calculated, and the actual probability under each pollutant threshold is calculated to obtain several probability points; Establish the actual probability density function obs(x) and the predicted probability density function fore(x); Substitute the pollutant concentration predicted by the numerical model into fore(x) and calculate the corresponding probability value under this function; Based on the calculated probability value, substitute the probability value into obs(x) to find the corresponding concentration interval, and obtain the corrected value by downward interpolation; The final correction results are calculated for different reporting times, forecast time limits and stations.
2. The PDF-based air quality forecast correction method according to claim 1, characterized in that: The threshold intervals for determining the actual situation and forecast of the pollutant specifically include: The threshold ranges of actual and forecast conditions correspond to the eight levels specified in the Ambient Air Quality Index. The rolling period is determined as the past 20 days, with two time periods, 8:00 and 20:
00. For all stations within the forecast range, each station is modeled separately.
3. The PDF-based air quality forecast correction method according to claim 1, characterized in that: The formula for calculating the frequency of different threshold values of the pollutant observed in real time during the rolling period is: Where P i Indicates the frequency of a certain pollutant threshold, A i It indicates the number of stations with a certain level of pollutant threshold, and B indicates the total number of stations.
4. The PDF-based air quality forecast correction method according to claim 1, characterized in that: The actual probability density function obs(x) is established by nonlinear fitting and piecewise function fitting.
5. The PDF-based air quality forecast correction method according to claim 4, characterized in that: The method of establishing the actual probability density function obs(x) through nonlinear fitting and piecewise function fitting is: Among them, the probability density function form of the nonlinear function is: In the formula, p(x) represents the probability density function, K represents the number of Gaussian distributions, and w i represents the weight of the i-th Gaussian distribution, satisfying μ i represents the mean of the i-th Gaussian distribution, represents the standard deviation of the i-th Gaussian distribution; Use maximum likelihood estimation to determine the parameter w in the above function i , μ i and σ i : Assume that the observed data are x1, x2, …, x N , the likelihood function L is: The parameters are solved by maximizing the likelihood function L and using the EM algorithm to iteratively solve: In which, step E calculates each data point x j The probability of belonging to the i-th Gaussian distribution is γ ij : Among them, M steps update parameter w i , μ i and σ i : Repeat steps E and M until the parameters converge. The function p(x) obtained at this time is the actual probability density function obs(x); Among them, the piecewise function fitting method is: According to the previously determined pollutant threshold range, the concentration range is divided into M intervals I1, I2, …, I M ; For each interval I M , using a linear function for fitting: Let interval I m =[a m ,b m ], the linear function form of the fitting in this interval is: y=c m x+d m Determine the coefficient c by the least squares method m and d m , let the observed data points in this interval be The goal is to minimize the sum of squared errors: c m and d m Find the partial derivative and set it to 0, and we get the system of equations: Solve the system of equations to get c m and d m The value of , thus determining the interval I m The fitting function within ; Let interval I m The fitting function in is y = e m , determine e by calculating the average probability value of the data points in the interval m ,Right now: The final actual probability density function obs(x) is: Through the above nonlinear fitting or piecewise function fitting method, the actual probability density function obs(x) is established.
6. The PDF-based air quality forecast correction method according to claim 1, characterized in that: The method to establish the forecast probability density function fore(x) is: Collect data on several pollutant forecasts and their corresponding frequencies and organize them into data pairs (x i ,y i ); For nonlinear fitting, the function form is chosen based on preliminary observations of the data, and a least squares problem is solved using a numerical optimization algorithm to determine the parameters of the function; For piecewise function fitting, divide the data into segments, determine the function form of each segment, solve the least squares problem for each segment, and ensure that the continuity condition is met at the connection points; The goodness of fit index is used to evaluate the fitting effect, and the fitting result is obtained through parameter estimation method based on the evaluation result.
7. The PDF-based air quality forecast correction method according to claim 6, characterized in that: The nonlinear fitting and piecewise function fitting are specifically include: Among them, the nonlinear fitting method is: For data that exhibit a bell-shaped curve, the fitted function takes the form: In the formula, μ represents the mean and σ represents the standard deviation; For data with positive skewness, the fitted function takes the form: Use the least squares method to estimate parameters: Assume there are n data points (x i ,y i ), where x i is the predicted pollutant value, y i is the corresponding pollutant forecast frequency, and the sum of squared errors is minimized by the least squares method: In the formula, θ represents the parameter vector of the function; The coefficient of determination R 2 Calculate goodness of fit: in, R 2 The closer the value is to 1, the better the fitting effect is; Among them, the piecewise function fitting method is: Divide the data into several non-overlapping segments according to the range of predicted pollutants; In each data segment, the function form is: y=f k (x)=a k x+b k In the formula, k represents the kth segment; For each segment, estimate the parameters using the least squares method: Given a data point (x i,k ,y i,k ) In the kth segment, minimize the sum of squared errors: Ensure that the functions of adjacent segments are continuous at the connection point, that is, for the adjacent kth segment and k+1th segment, at the connection point x = x c,k Where f k (x c,k ) = f k+1 (x c,k ).
8. The PDF-based air quality forecast correction method according to claim 1, characterized in that: After calculating the correction value, the uncertainty of the correction result is evaluated; The uncertainty is measured by calculating the standard deviation of the correction results for different onset times, forecast ages and stations: Assume that for a particular pollutant, under different reporting times, forecast time periods and stations, the corrected concentration value is C ij , where i represents the start time of the report, j represents the combination identifier of the forecast time and the station, and there are n combinations in total; First calculate the average of the corrected concentration values: In the formula, represents the average value of the corrected concentration values; The larger the standard deviation, the greater the dispersion of the correction results and the higher the uncertainty; conversely, the smaller the standard deviation, the more stable the correction results and the lower the uncertainty.
9. The PDF-based air quality forecast correction method according to claim 1, characterized in that: Introduce time series analysis to dynamically update the correction results; Let y t is the corrected air quality data at time t, which is modeled by the prediction algorithm, where the expression of the prediction algorithm model is: In the formula, and θ j Represent the autoregressive coefficient and the moving average coefficient, ∈ t represents a white noise sequence, p and q represent the autoregressive order and the moving average order respectively; The model parameters are determined by minimizing the residual sum of squares. and θ j , when there is new observation data y t+1 When y is predicted based on the established algorithm model t+1 The predicted value of And compare it with the actual observation value and calculate the residual: In the formula, e t+1 represents the residual value; According to the residual size and change trend, the model parameters are dynamically adjusted to update the subsequent correction results.
10. The PDF-based air quality forecast correction method according to claim 9, characterized in that: The model parameters are determined by minimizing the residual sum of squares. and θ j The method is: Based on the prediction algorithm model, the residual sum of squares function is constructed: In the formula, S represents the residual sum of squares, represents the predicted value; Since the predicted value It's about parameters and θ j function, so S is the parameter and θ j The function of To find the parameter value that minimizes S, take the partial derivatives of S with respect to the autoregressive coefficient and the moving average coefficient respectively; The gradient descent method is used to iteratively solve the problem until the residual square sum S converges to obtain the model parameters and θ j .