PM2.5 secondary seasonal prediction method fusing circulation evolution and overlapped time window
By integrating circulation evolution and overlapping time windows, a PM2.5 subseasonal prediction model was constructed, which solves the problem of insufficient PM2.5 concentration subseasonal prediction capability in existing technologies and achieves efficient prediction on the subseasonal time scale. By utilizing the coupling relationship between meteorological circulation factors and PM2.5, the prediction accuracy and stability are improved.
Patent Information
- Application Number
- CN202511485776.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-01-13
AI Technical Summary
Existing technologies have weak predictive capabilities for PM2.5 concentrations on the sub-seasonal timescale. Numerical model prediction errors accumulate rapidly, statistical prediction methods cannot reflect dynamic processes, and machine learning lacks physical constraints, resulting in unstable prediction results. There is a lack of an effective sub-seasonal prediction system for pollutant concentrations.
The PM2.5 subseasonal prediction method, which integrates circulation evolution and overlapping time windows, collects three-dimensional meteorological circulation factors and PM2.5 concentration observation data, performs data preprocessing and empirical orthogonal function decomposition, constructs an SVD model, combines S2S model forecast data, performs cross-covariance matrix decomposition and multi-factor integration, performs order-of-magnitude modulation, and finally generates PM2.5 concentration forecast values.
It improves the accuracy of PM2.5 concentration prediction on the sub-seasonal timescale, fills the prediction gap between synoptic and short-term climate prediction, improves prediction skills by utilizing the coupling relationship between meteorological circulation factors and PM2.5, and establishes a complete forecasting system.
Smart Images

Figure CN121325291A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the field of PM2.5 sub-seasonal prediction method combining circulation evolution and overlapping time window. BACKGROUND
[0002] At present, the prediction methods for PM2.5 concentration mainly focus on the prediction of weather scale (within 10 days), including numerical dynamic models such as WRF-Chem, machine learning prediction models, numerical model-machine learning hybrid models, and statistical prediction models such as multiple linear regression for monthly average, seasonal average and annual average; the sub-seasonal prediction (10-30 days) between weather scale (within 10 days) and short-term climate prediction (1 month to 1 year) is the key window period for responding to major meteorological disasters and implementing disaster prevention and mitigation decisions, which has important social and economic value for people's life safety and public property protection, but the current various prediction models have weak sub-seasonal prediction ability for PM2.5 concentration, the main reason is that the understanding of the sub-seasonal predictability source of PM2.5 is limited, and there is a lack of effective sub-seasonal prediction system for pollutant concentration in business; At present, in the prediction method of PM2.5 concentration, due to the characteristics of rapid formation and rapid diffusion of PM2.5 particulate pollutants, the error of numerical model prediction method accumulates rapidly when the prediction time is more than 10 days, and the prediction accuracy decreases significantly, which cannot realize long-term prediction; the statistical prediction method cannot reflect the dynamic process of sub-seasonal scale and the occurrence of heavy pollution events; the machine learning and hybrid prediction model method lack physical constraints, resulting in unstable prediction results, which are not suitable for predicting the concentration of PM2.5 at the sub-seasonal time scale; At present, in the method of sub-seasonal prediction, the S2S dynamic model mainly predicts some meteorological large-scale circulation field and meteorological element field, and is not applied to the prediction of particulate pollutant concentration; the pure statistical prediction method uses the lead-lag statistical relationship between the early prediction factor information and the prediction quantity to model and predict, and the statistical relationship between the early meteorological circulation factor and the PM2.5 concentration is relatively weak, so the prediction skill of the prediction model established by using the pure statistical method is low, the error is large, and the advantage of the mode prediction circulation factor is not utilized. The current dynamic-statistical combined method is only applied to the prediction of temperature, precipitation and other meteorological elements, which can utilize the strong correlation between the same period large-scale circulation and local meteorological elements (temperature, precipitation, etc.), improve the prediction skill of temperature, precipitation and other meteorological elements in the extended period, but is not used for sub-seasonal prediction of PM2.5 and other particulate pollutant concentrations. Most machine learning models are mainly driven by data, and it is difficult to understand and truly explain the physical mechanism behind them.
[0003] Therefore, we build a subseasonal prediction model of PM2.5 concentration based on meteorological circulation factors for 10-30 day scale from the seasonal evolution mechanism and prediction method of PM2.5, which fills the "prediction blank area" between weather scale and short-term climate prediction scale in China's pollution prevention and control. SUMMARY
[0004] The technical problem solved by the present application is to provide a PM2.5 subseasonal prediction method combining circulation evolution and overlapping time window, which can effectively solve the problems raised in the above background.
[0005] To solve the above problems, the technical scheme adopted by the present application is: a PM2.5 subseasonal prediction method combining circulation evolution and overlapping time window, comprising the following steps: S1, data preparation: collecting observed data of three-dimensional meteorological circulation factors, S2S (subseasonal-to-seasonal) model prediction data, and PM2.5 concentration observation data; the three-dimensional meteorological circulation factors include pressure field, potential height field, and wind field variables; S2, data preprocessing: removing the climatic cycle of all data in step S1 to extract the abnormal signal, using a 10-30 day band-pass filter to extract the 10-30 day scale subseasonal signal; determining the prediction factor region according to the lead-lag correlation between the three-dimensional meteorological circulation factors and the PM2.5 concentration, performing empirical orthogonal function decomposition on the subseasonal signal of the three-dimensional meteorological circulation factors in the region, and selecting the mode with cumulative variance contribution rate ≥95% to reconstruct the predictable signal; S3, model building: S3.1, for each prediction starting day, a plurality of continuous 30-day prediction factor evolution sequences are constructed: each sequence is composed of m-day three-dimensional meteorological circulation factor observation data before the prediction starting day and n-day three-dimensional meteorological circulation factor data after the prediction starting day, when constructing SVD, n-day three-dimensional meteorological circulation factor observation data is used, and when predicting projection, n-day three-dimensional meteorological circulation factor is provided by S2S model prediction data, wherein m+n=30; S3.2, define left field matrix X and right field matrix Y: The dimension of the left field matrix X is s1xt, wherein s1=x1x y1x(m+n) is the combination of spatial and temporal dimensions, x1 and y1 are the longitude and latitude grid points of the prediction factor space range determined in step S2, representing the evolution of the prediction factor field for 30 consecutive days, and t is the sample amount of the prediction starting day; The dimension of the right field matrix Y is s2xt, wherein s2=x2x y2x30, x2 and y2 are the number of longitude and latitude grid points of PM2.5 concentration; calculate the cross-covariance matrix of the left field matrix X and the right field matrix Y in It is an s1×s2 matrix. medium elements This represents the covariance between the i-th spatial-temporal variable of the left-field matrix X and the j-th spatial-temporal variable of the right-field matrix Y; S3.3, Regarding the cross-covariance matrix Perform Singular Value Decomposition (SVD). The decomposition formula is as follows: Where U is the left singular vector matrix, and each column of it... This represents a spatial distribution type of the forecast factor field; It is a right singular vector matrix, and each column of it... This represents a spatial distribution type of the predicted object field; For a diagonal singular value matrix, the non-negative elements on its diagonal are... , that is, singular values, which represent the first singular value. The magnitude of the covariance of the coupled modes; after decomposition, the original field is represented as a linear combination of all modes: ; The cumulative variance contribution rate of 95% was used to obtain the evolution coupling modes and time coefficients of the top N forecast factors and PM2.5: Where t is the forecast start date, X is the SVD left field, which consists of the forecast factor sequence for 30 days before and after the forecast start date (m+n=30); Y is the SVD right field, which consists of the forecast object sequence for 30 days after the forecast start date. Let i be the left singular eigenvector of the forecast factor after SVD decomposition. For the corresponding i-th time coefficient; Let be the right singular eigenvector of the predicted object after SVD decomposition. y1 and y2 are the latitude and longitude grid points of the spatial range of the forecast factor, respectively, and x2 and y2 are the latitude and longitude grid points of the spatial range of the forecast PM2.5. S4. Model Training: S4.1. A k-fold cross-validation test is used, where the value of k equals the total number of years M in the modeling period. Each test uses one year as the target year, and the remaining M-1 years are used as modeling data. The m-day observation data of the three-dimensional meteorological circulation factors of the target year and the n-day forecast data of the three-dimensional meteorological circulation factors from the S2S model (m+n=30 days) are projected onto the left singular vector matrix U obtained based on the M-1 year modeling data using the least squares method to obtain the projection coefficients: wherein, is the forecast target day of the reconstruction year, is the reconstructed forecast quantity; the projection coefficient is multiplied by the right singular vector matrix , to obtain the PM2.5 concentration back-calculation sequence ; S4.2, repeat step S4.1 process M times, to obtain the PM2.5 concentration back-calculation data of complete M years; S4.3, calculate the time correlation coefficient of M years of PM2.5 concentration back-calculation data and the same period PM2.5 concentration observation data, the calculation formula is wherein, is the PM2.5 concentration back-calculation value of the i-th reporting day to the j-th target day, is the PM2.5 concentration back-calculation average of the i-th reporting day, is the PM2.5 concentration observation anomaly value of the j-th target day, is the average of the PM2.5 concentration observation anomaly value; M represents the number of time series, the range of TCC is between-1 to 1, the closer TCC is to 1, the higher the model prediction skill is, the optimal three-dimensional meteorological circulation factor used for prediction and the corresponding m, n combination are screened out; S4.4, magnitude modulation is carried out on the screened PM2.5 concentration back-calculation sequence, and the modulation formula is wherein, is the PM2.5 concentration sequence after magnitude modulation, is the standard deviation of the PM2.5 concentration observation sequence in the M-year modeling period, is the standard deviation of the M-year PM2.5 concentration back-calculation sequence; the magnitude modulation results of different three-dimensional meteorological circulation factors are arithmetically averaged, and if the integrated TCC is better than the TCC of a single factor, the prediction model trained is determined; S5, complete model: real-time data is projected to the N left singular vectors obtained by training to obtain the prediction result, the prediction result is subjected to magnitude modulation, the multiple-factor modulation results are arithmetically averaged, and the final PM2.5 concentration forecast value is generated.
[0006] As a further preferred scheme of the present application, m+n=30 in step S3.1, wherein m , n is selected from any one of the following combinations: (0, 30), (5, 25), (10, 20), (15, 15), (20, 10), (25, 5), (30, 0); seven kinds of 30-day continuous forecast factor sequences are constructed.
[0007] As a further preferred scheme of the present application, the process is repeated M times in the step S4.2, specifically: taking each of the years in the modeling period as the target year and the remaining M-1 years as the modeling data in turn, after the back-calculation of all the years is completed, the PM2.5 concentration back-calculation data of the complete M years are integrated.
[0008] As a further preferred scheme of the present application, the process is repeated M times in the step S4.2, specifically: taking each of the years in the modeling period as the target year and the remaining M-1 years as the modeling data in turn, after the back-calculation of all the years is completed, the PM2.5 concentration back-calculation data of the complete M years are integrated. is the standard deviation of the PM2.5 concentration observation sequence in the M-year modeling period at the spatial grid point x2x y2 and the time dimension of 30 days; the
[0009] is the standard deviation of the PM2.5 concentration back-calculation sequence in the M-year modeling period at the same spatial grid point x2x y2 and the time dimension of 30 days.
[0010] As a further preferred scheme of the present application, the arithmetic mean of the multi-factor integration in the step S4.4 and the step S5 is to calculate the average value with the same weight for the magnitude modulation results corresponding to all the optimal three-dimensional meteorological circulation factors screened out.
[0011] As a further preferred scheme of the present application, the real-time data projection in the step S5 is realized by the following formula: wherein x1x y1 is the longitude and latitude grid point of the forecast factor space range, m is the observation data days, n is the model data days, m+n=30, is the forecast target day.
[0012] As a further preferred scheme of the present application, the prediction result calculation in the step S5 is realized by the following formula: wherein, is the forecast quantity, and x2x y2 is the longitude and latitude grid point of the PM2.5 forecast quantity.
[0013] As a further preferred scheme of the present application, the magnitude modulation in the step S5 is realized by the following formula: wherein, is the forecast result after magnitude modulation; is the observation standard deviation, is the model output standard deviation, and finally, the multi-factor modulation result is integrated to obtain the final real-time prediction result of the PM2.5 concentration.
[0014] Compared with the prior art, the present application provides a PM2.5 sub-season prediction method fusing circulation evolution and overlapping time window, which has the following beneficial effects: 1. The built prediction model can predict the concentration of PM2.5 on a sub-seasonal time scale, filling the "prediction gap" between weather scale (within 10 days) and short-term climate prediction (more than 30 days).
[0015] 2. Since the PM2.5 concentration is closely related to the coupling relationship of the near-synchronous circulation factor, the statistical relationship between the meteorological circulation factor and the PM2.5 concentration is used, the large-scale circulation factor predicted by the S2S mode, and the idea of combining dynamics and statistics is used to predict the future PM2.5 time series using the time series of the meteorological circulation factor in the early period-synchronous period, which is easy to understand the physical mechanism behind, effectively improves the prediction skill, and builds a complete prediction system. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 Figure 1 is a distribution map of the prediction skill of the PM2.5 concentration in the winter of North China 6-30 days in advance based on the fusion circulation evolution and the overlapping time window of the PM2.5 sub-seasonal prediction method of different observation-mode data combination schemes of the ECWMF mode, wherein (a) is based on the z500 factor, (b) is based on the v850 factor; Figure 2 Figure 2 is a spatial distribution map of the prediction skill of the PM2.5 concentration in the winter of North China 6-30 days in advance based on the fusion circulation evolution and the overlapping time window of the PM2.5 sub-seasonal prediction method of the ECWMF mode and the z500+z200+u200+u500+v200+v500+v850+t850+q1000 multi-factor integration; Figure 3 Figure 3 is a schematic diagram of the application effect of the PM2.5 sub-seasonal prediction method based on the fusion circulation evolution and the overlapping time window of the ECWMF mode and the z500+z200+u200+u500+v200+v500+v850+t850+q1000 multi-factor integration in the winter of North China, wherein (a) is a spatial distribution map of the PM2.5 concentration anomaly in North China in December 2013; Figure 4 Figure 4 is a schematic diagram of the application effect of the PM2.5 sub-seasonal prediction method based on the fusion circulation evolution and the overlapping time window of the ECWMF mode and the z500+z200+u200+u500+v200+v500+v850+t850+q1000 multi-factor integration in the winter of North China, wherein (b) is a spatial distribution map of the 10-day prediction result of the prediction method of the present application on the pollution process from December 15 to 25, 2013; Figure 5 Figure 5 is a schematic diagram of the method flow of the present application. DETAILED DESCRIPTION
[0017] It should be noted that the technical solutions of various embodiments can be combined with each other, but it must be based on the realization of ordinary skilled in the art, when the combination of technical solutions appear contradictory or can not be realized, it should be considered that the combination of technical solutions does not exist, also not within the scope of protection required by the invention.
[0018] Existing technical solutions: I. PM2.5 concentration prediction method: 1) Numerical mode prediction method Typical representative: WRF-Chem, WRF-CMAQ, etc. Atmospheric chemical coupling model Technical features: Mechanism modeling based on physical and chemical processes, including complete processes such as emission sources, chemical reactions, dry and wet deposition Time scale: Mainly applicable to 0-10 days of weather scale prediction 2) Statistical prediction method Typical representative: Time series analysis (ARIMA), multiple linear regression, etc. Technical features: Dependence on historical statistical relationship, high computational efficiency Time scale: Short-term climate prediction above the monthly scale 3) Machine learning prediction method Typical representative: Random forest, LSTM, Transformer algorithm, etc. Technical features: Data-driven nonlinear modeling capability Time scale: Prediction attempts of various time scales 4) Hybrid prediction method Typical representative: WRF-machine learning coupled model Technical features: Combination of numerical mode and machine learning advantages Time scale: Mainly applicable to 0-10 days of weather scale prediction II. Current subseasonal prediction method: 1) Dynamic mode method: Mainly S2S (subseasonal-to-seasonal) dynamic mode, including more than ten domestic and foreign business prediction mode subseasonal meteorological element return and real-time prediction results, and S2S multi-mode data set has been established.
[0019] 2) Traditional statistical method: For example, STPM spatial-temporal projection model.
[0020] 3) Dynamic-statistical combined method: Dynamic-statistical downscaling subseasonal prediction method based on low-frequency increment space-time coupling, etc.
[0021] 4) Machine learning method: The "Fu Xi" machine learning model established by Fudan University and the National Climate Center can make extended-range forecasts of various meteorological variables on a sub-seasonal scale.
[0022] To overcome the drawbacks of the prior art, as a specific embodiment of the present application: comprising the following steps 1) Data preparation Collect observed and model forecast data of three-dimensional meteorological circulation factors (pressure field, geopotential height field, wind field and other variables), as well as observed PM2.5 concentration data.
[0023] 2) Data preprocessing According to the modeling year (a total of M years), all the data is removed from the climatic cycle to extract the abnormal signal, 10-30 day filtering to extract the sub-seasonal signal, according to the lead-lag correlation of meteorological factors and PM2.5 concentration to determine the prediction factor area, select the mode with cumulative variance contribution rate of 95% for Empirical Orthogonal Function (Empirical Orthogonal Function, EOF) reconstruction, and extract the predictable signal.
[0024] 3) Model building Using the continuous 30-day prediction factor sequence obtained by step 2) during the modeling period (a total of M years) and the corresponding future 30-day PM2.5 sequence, singular value decomposition (Singular Value Decomposition, SVD) is performed to extract the high coupling mode between the left field (prediction factor field) and the right field (PM2.5 field), and the first N modes with cumulative variance contribution rate of 95% are selected to obtain N pairs of left and right singular vectors and their corresponding time coefficients.
[0025] Further, the specific steps are as follows: 3.1) Select the prediction factor sequence (pressure field, geopotential height field, wind field or other variables) of the observation data before the prediction starting day for continuous m days and after the prediction initial day for continuous n days (m+n=30, m takes 0, 5, 10, 15, 20, 25, 30; n takes 30, 25, 20, 15, 10, 5, 0; a total of 7 combinations) to construct the SVD left field scatter matrix, and the observation data of the PM2.5 prediction object for 30 consecutive days after the prediction initial day to construct the SVD right field scatter matrix, and perform SVD decomposition on the left field scatter matrix and the right field scatter matrix: Let the left-field forecast factor matrix be denoted by PM2.5 concentration dynamics-statistical sub-seasonal prediction method (dimension s1×t, where s1 = PM2.5 concentration dynamics-statistical sub-seasonal prediction method 1 × y1 × (m+n) is a combination of spatial and temporal dimensions, PM2.5 concentration dynamics-statistical sub-seasonal prediction method 1 and y1 are the latitude and longitude grid points of the forecast factor spatial range, representing the evolution of the forecast factor field over 30 consecutive days, and t is the sample size, i.e., the number of forecast start dates), and the right-field forecast object matrix be denoted by Y (dimension s2×t, where s2 = PM2.5 concentration dynamics-statistical sub-seasonal prediction method 2 × y2 × 30; PM2.5 concentration dynamics-statistical sub-seasonal prediction method 2, and y2 are the latitude and longitude grid points of the forecast quantity). Then the cross-covariance matrix of the two fields is... : in, It is an s1×s2 matrix, whose elements Let represent the covariance between the i-th spatial-temporal variable in the left field and the j-th spatial-temporal variable in the right field. The general formula for decomposing this cross-covariance matrix using SVD is: Where: U (dimension s1×M, M=min(s1,s2)) is a left singular vector matrix, and each column of it... This represents a spatial-temporal evolution type (i.e., left mode) of the forecast factor field. V (dimension s²×M) is a right singular vector matrix, and each column of it... This represents a spatial distribution type of the predicted object field (i.e., the right mode). ∑(dimension M×M) is a diagonal matrix whose non-negative elements on the diagonal are... These are singular values, representing the magnitude of the covariance of the i-th pair of coupled modes.
[0026] After decomposition, the original field can be represented as a linear combination of all modes:
[0027] 3.2) Taking the cumulative variance contribution rate of 95%, we obtain the evolution coupling modes and time coefficients of the top N forecast factors and PM2.5: Where t is the forecast start date, The left field of SVD consists of a sequence of forecast factors for 30 days before and after the forecast start date (m+n=30); The right field of SVD consists of a sequence of forecast objects for the next 30 days after the forecast start date; Let i be the left singular eigenvector of the forecast factor after SVD decomposition. For the corresponding i-th time coefficient; Let be the right singular eigenvector of the predicted object after SVD decomposition. y1 represents the corresponding i-th time coefficient, N represents the number of modes obtained with a 95% variance contribution rate, y2 represents the latitude and longitude grid points of the predicted factor spatial range in the dynamic-statistical sub-seasonal prediction method 1, and y2 represents the latitude and longitude grid points of the predicted PM2.5 spatial range in the dynamic-statistical sub-seasonal prediction method 2.
[0028] 4) Model Training Using k-fold cross-validation, one year is selected as the target year, and the remaining M-1 years are used as the modeling objects to reconstruct the backtesting results for the target year. This process is repeated until complete backtesting data for M years is obtained. Based on the backtesting data, the optimal forecasting factor, the number of m and n combinations for different factor modeling, and the factor combinations for multi-factor integration are determined. The specific steps are as follows: 4.1) For each forecast factor, a k-fold cross-validation is performed. When reconstructing the target year, the reanalysis data for the consecutive m days before the forecast initiation date and the model forecast data for the consecutive n days after the forecast initiation date are used to construct a 30-day forecast factor sequence. The least squares method is used to project this sequence onto the SVD left singular eigenvector used in the modeling with M-1 year data. The resulting projection coefficients are used as estimates of the SVD right modal time coefficients. The dot product is then applied to the SVD right singular eigenvector to obtain the reconstructed 30-day PM2.5 concentration anomaly sequence. in, The forecast target date for the reconstruction year. This represents the reconstructed forecast. The 30-day forecast variable sequence obtained by accumulating N sets of modes is used as the backcalculation result; other variables are handled similarly.
[0029] 4.2) Repeat the above steps to obtain backtesting data for M years for each factor with different combinations of m and n.
[0030] 4.3) For each factor, the back-calculated data obtained from different combinations of m and n over M years are compared with the observed data over M years to calculate the TCC (Time Correlation Coefficient). Based on the principle that the TCC of the regional average of the reconstruction output results and the regional average of the observed anomalous fields is larger and has a longer duration, the factors that can be used to build the prediction model and the number of m and n combinations corresponding to each factor are selected: Where i represents the start date and j represents the target date. This represents the prediction for day j from day i. Represents the forecast average for day i. This indicates an observational anomaly on day j. express The average value is M, which represents the number of time series. The TCC ranges from -1 to 1. The closer the TCC is to 1, the higher the model's predictive skill.
[0031] 4.4) The back-calculation results of m+n combinations of different selected factors are modulated by magnitude, and the modulated results are arithmetically averaged. The TCC of the multi-factor ensemble is calculated. If it is better than the prediction technique of single factors, the final prediction model can be determined. in, The result of the magnitude modulation for reconstructing the year; Let M be the standard deviation of the observed sequence for the modeling period in year M. The standard deviation of the complete M-year series output by the model; the other variables are the same as above.
[0032] 5) Complete Model 4) After the steps are completed, in actual modeling, the selected predictive factors and their corresponding m and n combinations of M-year data are used for modeling to obtain N sets of coupled modes corresponding to different factors. In real-time forecasting, the observation data of the consecutive m days before the forecast date and the model data of the following n days are combined to form a continuous 30-day time series and projected onto the corresponding N left singular eigenvectors. The resulting projection coefficients are multiplied by the N singular vectors of the right field, and the results of the multiplication of N pairs of modes are summed to obtain the output of a single factor model. The magnitude modulation is then performed, and finally the arithmetic mean of all modulation results is calculated to obtain the real-time forecast sequence for the next 30 days. in, To forecast the target date, The forecast quantity is calculated by summing N modalities to obtain the forecast variable sequence for the next 30 days, which is then used as the output of a dynamical-statistical downscaling model based on the fusion of observation and model data. Subsequently, the final forecast result is modulated by the ratio of the observed standard deviation to the model output standard deviation. in, The forecast results are after magnitude modulation; For the observed standard deviation, Output the standard deviation for the model; the other variables are the same as above.
[0033] Finally, the final real-time forecast of PM2.5 concentration is obtained by integrating the results of multi-factor modulation.
[0034] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural transformations made using the contents of the present invention's specification and drawings under the inventive concept of the present invention, or direct / indirect applications in other related technical fields, are included within the patent protection scope of the present invention.
Claims
1. A PM2.5 sub-seasonal prediction method integrating circulation evolution and overlapping time windows, characterized in that, Includes the following steps: S1. Data preparation: Collect observational data and model forecast data of three-dimensional meteorological circulation factors, as well as PM2.5 concentration observation data; S2. Data Preprocessing: Remove climatological cycles from all data in step S1 to extract anomalous signals. Use a 10-30 day bandpass filter to extract sub-seasonal signals at the 10-30 day scale. Determine the forecast factor region based on the lead-lag correlation between the three-dimensional meteorological circulation factor and PM2.5 concentration. Perform empirical orthogonal function decomposition on the sub-seasonal signals of the three-dimensional meteorological circulation factor in this region. Select the modal reconstructed forecastable signals with a cumulative variance contribution rate ≥95%. S3, Model Building: S3.1 For each forecast start date, construct multiple forecast factor evolution sequences for 30 consecutive days: each sequence consists of three-dimensional meteorological circulation factor observation data for m days before the forecast start date and three-dimensional meteorological circulation factor data for n days after the forecast start date. When constructing SVD, the three-dimensional meteorological circulation factor for n days is based on observation data, and the three-dimensional meteorological circulation factor for n days during projection is provided by S2S model forecast data, where m+n=30; S3.2, Define the left field matrix X and the right field matrix Y: The dimension of the left field matrix X is s1×t, where s1=x1×y1×(m+n) is the combination of spatial and temporal dimensions, x1 and y1 are the latitude and longitude grid points of the forecast factor spatial range determined in step S2, representing the evolution of the forecast factor field over 30 consecutive days, and t is the sample size on the forecast start date. The dimension of the right-field matrix Y is s²×t, where s² = x²×y²×30, and x² and y² are the number of latitude and longitude grid points representing the PM2.5 concentration, respectively. The cross-covariance matrix between the left-field matrix X and the right-field matrix Y is calculated. in It is an s1×s2 matrix. medium elements This represents the covariance between the i-th spatial-temporal variable of the left-field matrix X and the j-th spatial-temporal variable of the right-field matrix Y; S3.3, Regarding the cross-covariance matrix Perform SVD singular value decomposition, the decomposition formula is: Where U is the left singular vector matrix, and each column of it... A spatial distribution type of the forecast factor field; It is a right singular vector matrix, and each column of it... This represents a spatial distribution type of the predicted object field; For a diagonal singular value matrix, the non-negative elements on its diagonal are... , that is, singular values, which represent the first singular value. The magnitude of the covariance of the coupled modes; After decomposition, the original field can be represented as a linear combination of all modes: ; The cumulative variance contribution rate of 95% was used to obtain the evolution coupling modes and time coefficients of the top N forecast factors and PM2.5: Where t is the forecast start date, X is the SVD left field, which consists of the forecast factor sequence for 30 days before and after the forecast start date (m+n=30); Y is the SVD right field, which consists of the forecast object sequence for 30 days after the forecast start date. Let i be the left singular eigenvector of the forecast factor after SVD decomposition. For the corresponding i-th time coefficient; Let be the right singular eigenvector of the predicted object after SVD decomposition. y1 and y2 are the latitude and longitude grid points of the spatial range of the forecast factor, respectively, and x2 and y2 are the latitude and longitude grid points of the spatial range of the forecast PM2.
5. S4. Model Training: S4.
1. A k-fold cross-validation test is used, where the value of k equals the total number of years M in the modeling period. Each test uses one year as the target year, and the remaining M-1 years are used as modeling data. The 30-day forecast factor evolution sequence of the target year (composed of m days of observation data and n days of S2S model forecast data) is projected onto the left singular vector matrix U obtained based on the M-1 year modeling data using the least squares method to obtain the projection coefficients: in, The forecast target date for the reconstruction year. For the predicted amount of reconstruction; the projection coefficients Dot product of right singular vector matrix The PM2.5 concentration back-calculation sequence was obtained. ; S4.2 Repeat step S4.1 M times to obtain complete PM2.5 concentration back-calculation data for M years; S4.3 Calculate the time correlation coefficient between the back-calculated PM2.5 concentration data for year M and the observed PM2.5 concentration data for the same period. The calculation formula is as follows: in, This represents the calculated PM2.5 concentration from the i-th reporting date to the j-th target date. Let be the average PM2.5 concentration calculated from the first reporting day (i). Let J be the observed outlier value of PM2.5 concentration on the j-th target day. The mean of the observed outliers of PM2.5 concentration; M represents the number of time series; TCC ranges from -1 to 1. The closer the TCC is to 1, the higher the model prediction skill. The optimal three-dimensional meteorological circulation factors and the corresponding m and n combinations are selected for prediction. S4.
4. The back-calculation sequence of PM2.5 concentration after screening is modulated by magnitude. The modulation formula is as follows: in, This is a PM2.5 concentration sequence after magnitude modulation. Let M be the standard deviation of the PM2.5 concentration observation series during the modeling period of year M. The standard deviation of the back-calculated PM2.5 concentration sequence for year M is used; the arithmetic mean of the modulation results of different three-dimensional meteorological circulation factors is used for multi-factor integration to determine the trained prediction model. S5. Complete Model: Project real-time data onto N left singular vectors obtained from training to obtain prediction results, perform magnitude modulation on the prediction results, integrate the multi-factor modulation results, and generate the final PM2.5 concentration forecast value.
2. The PM2.5 subseasonal prediction method based on the fusion of circulation evolution and overlapping time windows according to claim 1, characterized in that, In step S3.1, m+n=30, where ( m , n Choose from any of the following combinations: (0,30),(5,25),(10,20),(15,15),(20,10),(25,5),(30,0); Construct 7 consecutive 30-day forecast factor sequences.
3. The PM2.5 subseasonal prediction method based on the fusion of circulation evolution and overlapping time windows according to claim 1, characterized in that, The process of repeating the process M times in step S4.2 is as follows: each year in the modeling period is taken as the target year, and the remaining M-1 years are taken as the modeling data. After the back calculation of all years is completed, the complete PM2.5 concentration back calculation data for M years is obtained.
4. The PM2.5 subseasonal prediction method based on the fusion of circulation evolution and overlapping time windows according to claim 1, characterized in that, The steps described in S4.4 Let x be the standard deviation of the PM2.5 concentration observation sequence over the M-year modeling period, at spatial grid points x2×y2 and over a 30-day time dimension; Let x be the standard deviation of the back-calculated PM2.5 concentration series for year M across the same spatial grid x2×y2 and the time dimension of 30 days.
5. The PM2.5 subseasonal prediction method based on the fusion of circulation evolution and overlapping time windows according to claim 1, characterized in that, The arithmetic mean of the multi-factor integration in steps S4.4 and S5 is the average value calculated by weighting the magnitude modulation results corresponding to all the selected optimal three-dimensional meteorological circulation factors.
6. The PM2.5 subseasonal prediction method based on the fusion of circulation evolution and overlapping time windows according to claim 1, characterized in that, The real-time data projection in step S5 is achieved through the following formula: ; Where x1×y1 represents the latitude and longitude grid points of the spatial range of the forecast factor, m is the number of days of observation data, n is the number of days of model data, and m+n=30. Forecast target date.
7. The PM2.5 subseasonal prediction method based on the fusion of circulation evolution and overlapping time windows according to claim 6, characterized in that, The prediction result calculation in step S5 is achieved through the following formula: in, x2×y2 represents the forecast amount, and x2×y2 represents the latitude and longitude grid points of the PM2.5 forecast amount.
8. The PM2.5 subseasonal prediction method based on the fusion of circulation evolution and overlapping time windows according to claim 7, characterized in that, The magnitude modulation in step S5 is achieved through the following formula: in, The forecast results are after magnitude modulation; Observational standard deviation The model outputs the standard deviation, and finally, the final real-time forecast of PM2.5 concentration is obtained by integrating the results of multi-factor modulation.