A method for constructing an ozone pollution weather condition index

CN122596402APending Publication Date: 2026-08-18HUBEI PROVINCIAL ACADEMY OF ECO-ENVIRONMENTAL SCIENCES(PROVINCIAL ECOLOGICAL ENVIRONMENT ENGINEERING ASSESSMENT CENTER)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610688735.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0004]本发明所要解决的技术问题是:提供一种臭氧污染气象条件指数的构建方法,解决现有通用型臭氧污染气象条件指数构建本地化泛化性差、权重确定主观性强的技术问题

Benefits of technology

1、权重客观可靠:利用随机森林与SHAP分析自动量化各气象因子的贡献,避免了传统方法中人为赋权的主观性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122596402A_ABST
    Figure CN122596402A_ABST
Patent Text Reader

Abstract

The present application relates to the field of meteorological monitoring, and discloses a method for constructing an ozone pollution meteorological condition index, comprising: obtaining historical ozone and meteorological data of multiple cities, and dividing the data into a modeling set and a verification set after preprocessing; constructing a random forest regression model with ozone as a label, optimizing parameters through grid search cross-validation, and calculating weight coefficients of each meteorological factor through SHAP analysis; comparing multi-strategy binning through four strategies of equal width, equal frequency, clustering and decision tree optimal binning, introducing a physical trend consistency penalty term into the objective function, and optimizing the optimal binning interval that meets the physical monotonicity law; calculating the interval division index, and obtaining the ozone pollution meteorological condition index OPMI through weighted summation; and S5, dividing the potential level and verifying the reliability. The present application solves the problems of poor pertinence, subjective weight and non-fine interval division of the existing index, and improves the precision and business applicability of ozone pollution meteorological potential evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of meteorological monitoring, and in particular to a method for constructing an ozone pollution meteorological condition index. Background Technology

[0002] Ozone is one of the primary pollutants affecting urban air quality improvement in my country in recent years, and its formation mechanism is closely related to meteorological conditions. Ozone is not a primary emission pollutant, but a secondary pollutant generated from nitrogen oxides and volatile organic compounds through photochemical reactions under sunlight. Therefore, meteorological factors—especially temperature, solar radiation, relative humidity, wind speed, and wind direction—have a significant controlling effect on the formation, accumulation, diffusion, and removal of ozone. Under the premise of relatively stable short-term emission sources, the evolution of meteorological conditions often becomes a key external factor determining the level of ozone concentration and whether pollution exceeds standards. Accurately quantifying the contribution of meteorological conditions to ozone pollution and constructing a comprehensive index that can characterize the relationship between meteorological factors and ozone concentration response is of significant scientific importance and practical application value for ozone pollution forecasting and early warning, the formulation of prevention and control measures, and regional joint prevention and control.

[0003] Currently, scholars both domestically and internationally have conducted extensive research on the correlation between meteorological conditions and air pollution, and have proposed various meteorological condition indices. For example, the retention index, stable weather index, and atmospheric self-cleaning capacity index for fine particulate matter (PM2.5) have been widely used in air quality forecasting. However, most of these indices are designed based on the pollution characteristics of PM2.5 and do not fully consider the unique meteorological conditions required for ozone formation—ozone requires "photochemical" meteorological conditions of high temperature, strong radiation, and relatively low humidity, while PM2.5 relies more on "accumulation" conditions of stable weather and high humidity. The two have fundamentally different meteorological responses. Furthermore, although the China Meteorological Administration has required all provinces and cities to conduct ozone pollution meteorological condition level forecasting, most regions generally use general indices or linear weighted methods based on local experience. For example, existing technologies include methods for calculating ozone pollution meteorological risk indices based on nine meteorological factors, including daily average temperature, daily maximum temperature, sunshine duration, wind speed, wind direction, boundary layer wind speed, weak wind duration, rainfall, and relative humidity, by artificially setting fixed weights and using linear piecewise functions. This type of method has obvious shortcomings: First, the weighting coefficients rely on expert experience or simple statistics, lacking the ability to quantify the nonlinear interactions of multiple factors; second, the factor interval division uses fixed linear segments (such as (Tavg-20) / 15), which makes it difficult to accurately characterize the complex nonlinear response relationship between meteorological factors and ozone concentration; third, the index structure is fixed and lacks the ability to adapt to different regional characteristics. When moving to areas with strong spatiotemporal heterogeneity of meteorological conditions, such as different metropolitan areas, the forecast accuracy decreases significantly. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method for constructing an ozone pollution meteorological condition index, thereby solving the technical problems of poor localization and generalization of existing general ozone pollution meteorological condition index construction and strong subjectivity in weight determination.

[0005] Specifically, this invention provides a method for constructing an ozone pollution meteorological condition index, the method comprising the following steps: S1. Obtain historical ozone observation data and historical meteorological observation data from multiple cities, perform missing data imputation, outlier removal and standardization preprocessing on the historical ozone observation data and historical meteorological observation data, and divide them into modeling dataset and independent validation dataset according to time series stratified sampling. S2. Using ozone concentration in the modeling dataset as labels and meteorological factors as candidate features, a random forest regression model is constructed. The model parameters are optimized through grid search cross-validation. SHAP interpretability analysis is introduced to calculate the absolute SHAP value of each meteorological factor. After normalization, the weight coefficients of each meteorological factor are obtained. S3. For the key meteorological factors selected in S2, four binning strategies are used for comparison: equal-width binning, equal-frequency binning, clustering binning, and decision tree optimal binning. The objective functions are the significance of the difference in ozone concentration between intervals, the monotonicity of the intervals, and the physical rationality. A physical trend consistency penalty term based on the principle of atmospheric photochemistry is introduced into the objective function to select the optimal binning interval that conforms to the physical monotonicity law for each key meteorological factor. S4. Based on the optimal binning intervals obtained in S3, calculate the interval sub-index corresponding to each interval for each meteorological factor. The interval sub-index is the ratio of the average ozone concentration in the interval to the total average ozone concentration during the study period. Then, the interval sub-index corresponding to each meteorological factor falling into the interval at the same time is weighted and summed with the weighting coefficients obtained in S2 to obtain the weighted comprehensive ozone pollution meteorological condition index OPMI. S5. Based on the independent validation dataset, the ozone concentration distribution, exceedance rate and accuracy under different OPMI intervals are statistically analyzed, the meteorological potential level of ozone pollution is classified, and the forecast reliability of OPMI is verified by correlation coefficient, average deviation and root mean square error.

[0006] A storage medium storing instructions and data for implementing a method for constructing an ozone pollution meteorological condition index.

[0007] An apparatus for constructing an ozone pollution meteorological condition index includes: a processor and a storage medium; the processor loads and executes instructions and data in the storage medium to implement a method for constructing an ozone pollution meteorological condition index.

[0008] The beneficial effects provided by this invention are: 1. Objective and reliable weights: Random forest and SHAP analysis are used to automatically quantify the contribution of each meteorological factor, avoiding the subjectivity of manual weighting in traditional methods; 2. Scientific and reasonable binning: Integrating multiple binning strategies and introducing a physical trend penalty term ensures that the interval division satisfies both statistical significance and strictly conforms to the atmospheric photochemical laws governing ozone formation; 3. Strong local adaptability: It supports independent training of localized parameter systems for different cities, which significantly improves the forecast accuracy and generalization ability of the index in different regions. 4. Excellent forecasting capability: The constructed OPMI index is highly correlated with ozone concentration, and has been independently verified to have small deviation and low error, which can accurately characterize the comprehensive impact of meteorological conditions on ozone pollution; 5. Facilitates business applications: After real-time access to meteorological data, it can quickly calculate indices and output potential levels, and can be directly embedded into existing forecasting and early warning systems, providing efficient decision support for pollution prevention and control. Attached Figure Description

[0009] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a schematic diagram of the hardware device operation according to an embodiment of the present invention. Detailed Implementation

[0010] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0011] Before formally describing the present invention, a general description of the solution of the present invention will be given first to facilitate understanding.

[0012] Example 1 Please refer to Figure 1 The present invention provides a method for constructing an ozone pollution meteorological condition index, comprising the following steps: S1. Obtain historical ozone observation data and historical meteorological observation data from multiple cities, perform missing data imputation, outlier removal and standardization preprocessing on the historical ozone observation data and historical meteorological observation data, and divide them into modeling dataset and independent validation dataset according to time series stratified sampling. It should be noted that step S1 further includes: S11. Collect hourly ozone observation data from multiple cities, as well as hourly meteorological observation data on temperature, relative humidity, wind direction, wind speed, and air pressure during the same period. S12. Remove abnormal records where the ozone concentration is less than the first preset value or greater than the second preset value. Use spatiotemporal interpolation to interpolate continuous or intermittent missing data and standardize meteorological factors to eliminate the influence of dimensions. S13. Divide the data from the previous four years into a modeling dataset and the data from the latest year into an independent validation dataset.

[0013] As one embodiment, step S1 of the present invention is implemented as follows: First, data is acquired from the ambient air quality monitoring network and meteorological observation network of the target area. The "multiple cities" can be any urban cluster or region for which an ozone pollution meteorological condition index needs to be constructed, such as the Wuhan metropolitan area in central China, which includes nine cities: Wuhan, Huangshi, Ezhou, Xiaogan, Huanggang, Xianning, Xiantao, Tianmen, and Qianjiang. The acquired data includes: Historical ozone observation data: hourly ozone concentration (unit: μg / m³), with a time span of at least 5 consecutive years, for example, from January 1, 2019 to December 31, 2023.

[0014] Historical meteorological observation data: Hourly meteorological factor data from the same period as ozone observations, including air temperature (°C), relative humidity (%), wind direction (°), wind speed (m / s), and air pressure (hPa). Data are sourced from automatic weather stations in various cities.

[0015] Data preprocessing includes the following sub-steps: Outlier Removal: The normal range for ozone concentration is set to a first preset value to a second preset value. In this embodiment, the first preset value is set to 0 μg / m³ (values ​​below this are considered invalid), and the second preset value is set to 1000 μg / m³ (values ​​exceeding this are considered instrument malfunctions or extreme events). All records with ozone concentrations less than 0 μg / m³ or greater than 1000 μg / m³ are removed. For meteorological factors, obviously erroneous values ​​are removed based on the physical probability range of each factor, such as temperatures less than -50℃ or greater than 50℃, wind speeds less than 0 or greater than 50 m / s, etc.

[0016] Missing value imputation: For continuously missing (e.g., a device malfunction lasting for several hours) or intermittently missing data points, spatiotemporal interpolation methods are used. Specifically, within the same city, time imputation is performed using the weighted average of valid data from two hours before and after the missing time; if the continuous missing data in the same city exceeds 6 hours, spatial imputation is performed using inverse distance weighting, referencing data from neighboring cities at the same time.

[0017] Standardization preprocessing: To eliminate the dimensional influence of each meteorological factor, Z-score standardization is performed independently for each meteorological factor. The calculation formula is as follows: ,in This is the mean of the factor in the modeling data. denoted as standard deviation. After standardization, the meteorological factor data follow a standard normal distribution with a mean of 0 and a standard deviation of 1.

[0018] Dataset partitioning: Following the time series sequence, data from the previous four years (e.g., January 1, 2019 to December 31, 2022) were divided into a modeling dataset for model training and parameter determination; data from the most recent year (e.g., January 1, 2023 to December 31, 2023) were divided into an independent validation dataset for final index performance evaluation. Stratified time series sampling was employed to ensure that the validation dataset included data samples from different seasons and pollution levels.

[0019] Through the above preprocessing, a complete, clean, time-aligned, standardized environment dataset is obtained.

[0020] S2. Using ozone concentration in the modeling dataset as labels and meteorological factors as candidate features, a random forest regression model is constructed. The model parameters are optimized through grid search cross-validation. SHAP interpretability analysis is introduced to calculate the absolute SHAP value of each meteorological factor. After normalization, the weight coefficients of each meteorological factor are obtained. It should be noted that step S2 further includes: S21. Using air temperature, relative humidity, wind direction, wind speed, air pressure, and precipitation as candidate feature factors, and ozone hourly concentration as the output label, a random forest regression model is constructed. S22. The grid search and cross-search validation method is used to optimize the parameters of the number of trees, maximum depth and minimum number of leaf nodes of the random forest, so as to obtain the random forest regression model with the best fitting ability and generalization ability. S23. Calculate the SHAP value of each meteorological factor based on the optimal model, normalize the absolute value of the SHAP of each meteorological factor, and obtain the weight coefficient of each meteorological factor.

[0021] As one embodiment, step S2 of the present invention is implemented as follows: On the modeling dataset, ozone hourly concentration ( Using the unit μg / m³ as the prediction target (label), and candidate meteorological factors (temperature, relative humidity, wind direction, wind speed, air pressure, precipitation) as input features (independent variables), a random forest regression model is constructed.

[0022] Model Construction: RandomForestRegressor from Python's scikit-learn library is used. Initial parameters are set as follows: number of trees (n_estimators) = 100, maximum depth (max_depth) = 10, minimum number of leaf nodes (min_samples_leaf) = 5. Input feature matrix X (n_samples × 6), output vector y (n_samples × 1).

[0023] Grid Search Cross-Validation: To obtain the optimal model parameters, grid search (GridSearchCV) combined with 5-fold cross-validation was used for optimization. The parameter space searched was: number of trees [50, 100, 200], maximum depth [5, 10, 15, 20], and minimum number of leaf nodes [2, 5, 10]. The mean squared error (MSE) was used as the evaluation metric. After grid search, the optimal parameter combination was obtained. In the Wuhan metropolitan area data of this embodiment, the optimal parameters were: number of trees = 150, maximum depth = 12, and minimum number of leaf nodes = 4.

[0024] SHAP interpretability analysis: The contribution of each meteorological factor to ozone concentration prediction is calculated using the SHAP library (Shapley Additive exPlanations). Specifically, the optimally trained random forest model is used to predict ozone concentration for all samples in the modeling dataset, and then shap.TreeExplainer is called to calculate the SHAP value for each factor for each sample. The absolute SHAP value of each factor reflects the magnitude of its influence on ozone concentration prediction.

[0025] Weight normalization: Calculate the average of the absolute SHAP values ​​of each factor across all samples, obtaining an absolute SHAP value vector. Then, normalize this vector so that the sum of the factor weights is 1. The normalization formula is: ,in The number of meteorological factors is m (m=6 in this example). The weighting coefficients for each meteorological factor are then obtained. For example, in the Wuhan metropolitan area data, the weight of temperature is approximately 0.35, relative humidity is approximately 0.28, wind speed is approximately 0.20, air pressure is approximately 0.10, wind direction is approximately 0.05, and precipitation is approximately 0.02.

[0026] S3. For the key meteorological factors selected in S2, four binning strategies are used for comparison: equal-width binning, equal-frequency binning, clustering binning, and decision tree optimal binning. The objective functions are the significance of the difference in ozone concentration between intervals, the monotonicity of the intervals, and the physical rationality. A physical trend consistency penalty term based on the principle of atmospheric photochemistry is introduced into the objective function to select the optimal binning interval that conforms to the physical monotonicity law for each key meteorological factor. The construction and introduction of the physical trend consistency penalty term in step S3 is as follows: S31. Based on the physical mechanism of ozone photochemical generation and atmospheric diffusion, predefine the physical a priori monotonic relationship between each key meteorological factor and ozone concentration: temperature and ozone concentration are positively correlated and monotonically increasing, relative humidity and ozone concentration are negatively correlated and monotonically decreasing, and wind speed and ozone concentration are negatively correlated and monotonically decreasing. In this embodiment, the following three rules are defined: Temperature: If the interval The lower limit > the interval The upper limit should be [missing information]. (The average ozone concentration increases with rising temperature.)

[0027] Relative humidity (RH): If the range The lower limit > the interval The upper limit should be [missing information]. (The average ozone concentration decreases as humidity increases).

[0028] Wind speed (WindSpeed): If the interval The lower limit > the interval The upper limit should be [missing information]. (The mean ozone concentration decreases as wind speed increases).

[0029] S32. In the binning optimization process, for each candidate binning scheme, traverse all adjacent intervals and check whether there are interval pairs that violate the physical prior monotonic relationship. Record a violation event each time a violation occurs. For example, consider a candidate binning scheme for air temperature: Interval 1: [0~10℃], mean ozone 60; Interval 2: [10~20℃], mean ozone 55. Here, the mean of interval 2 is lower than that of interval 1, violating the positively correlated monotonically increasing relationship, and a violation event is recorded.

[0030] S33. Calculate the monotonicity violation degree of this binning scheme, using the following formula: ,in Let j be the weight of the violation event. The corresponding penalty coefficient is used; the more severe the violation of physical trends, the greater the penalty coefficient. In this embodiment, the following settings are provided: (Each violation carries the same weight). Penalty coefficient. Classified according to the severity of the violation: If it is only a reversal of the mean ozone concentration between adjacent intervals but the absolute value of the difference is less than 5 μg / m³, then ; If the reversal difference is between 5 and 15 μg / m³, then ; If the reversal difference is greater than 15 μg / m³, then .

[0031] For example, in the above temperature scheme, the reversal difference is 5 μg / m³, then This violation of contribution If a plan is violated multiple times, the violations are accumulated.

[0032] S34. The monotonicity violation degree V is incorporated as a negative indicator into the comprehensive objective function, which is expressed as: Where D is the significance index of the difference in ozone concentration between intervals, M is the interval monotonicity score, and V is the degree of monotonicity violation. These are the weighting coefficients; In this embodiment, D is the variance analysis F value of the mean ozone value in each interval (normalized so that D is between 0 and 1). M is the monotonicity uniformity ratio: for air temperature, the ratio of monotonically increasing adjacent interval logarithms to the total adjacent interval logarithms; for relative humidity and wind speed, the monotonically decreasing ratio. Let's take values ​​of 0.4, 0.3, and 0.3 respectively. For example, a scheme has 4 intervals and 3 pairs of adjacent intervals, of which 2 pairs conform to a monotonic relationship, so M = 2 / 3 ≈ 0.667; if there is one violation (V = 1), then F = 0.4 × D - 0.3 × 0.667 - 0.3 × 1 = 0.4D - 0.5. The larger D is, the larger F is.

[0033] S35. Taking the maximization of the comprehensive objective function F as the optimization objective, select the optimal binning interval for each key meteorological factor that has strong distinguishing ability, good monotonicity and the least violation of physical trends.

[0034] In this embodiment, for air temperature, relative humidity, and wind speed, the F-value of each candidate binning scheme (4 strategies × K = 3~8 intervals) is calculated, and the scheme with the largest F-value is selected as the optimal binning. The final results are as described above: air temperature is binned using a decision tree with 4 intervals, relative humidity using an equal-frequency binning with 4 intervals, and wind speed using a clustering binning with 3 intervals. These schemes do not violate the physical monotonicity (V=0) and have high D-values.

[0035] S4. Based on the optimal binning intervals obtained in S3, calculate the interval sub-index corresponding to each interval for each meteorological factor. The interval sub-index is the ratio of the average ozone concentration in the interval to the total average ozone concentration during the study period. Then, the interval sub-index corresponding to each meteorological factor falling into the interval at the same time is weighted and summed with the weighting coefficients obtained in S2 to obtain the weighted comprehensive ozone pollution meteorological condition index OPMI. It should be noted that the formula for calculating the interval sub-index in step S4 is as follows:

[0036] in, Let i be the interval sub-index when meteorological factor i falls into the t-th interval. Let i be the average ozone concentration when meteorological factor i falls within the t-th interval. The total average concentration of ozone during the study period; The weighted composite ozone pollution meteorological conditions index (OPMI) is calculated using the following formula: in, is the normalized weighting coefficient of the i-th meteorological factor, and n is the number of key meteorological factors.

[0037] As one embodiment, step S4 of the present invention is implemented as follows: First, calculate the total average ozone concentration during the study period (i.e., the time span of the modeling dataset, such as 2019-2022). In this embodiment, the total average ozone concentration in the Wuhan metropolitan area modeling dataset is 85 μg / m³.

[0038] For each key meteorological factor and its first Calculate the average ozone concentration of all samples within an optimal binning interval. Then calculate the interval sub-index:

[0039] This sub-index reflects the degree to which the ozone concentration deviates from the overall average level when the meteorological factor falls within a certain range. For example, the average ozone concentration corresponding to the temperature range [>28℃] is 120 μg / m³. This indicates that the ozone concentration in this range is 41% higher than the average level.

[0040] For any given time (or any sample), obtain the actual observed values ​​of each key meteorological factor at that time, determine which optimal bin interval it falls into, and thus obtain the interval sub-index of that factor at that time. Then, using the weighting coefficients obtained in step S2... (Air temperature 0.35, relative humidity 0.28, wind speed 0.20, other factors with smaller weights can be ignored or retained) Perform a weighted summation:

[0041] in This refers to the number of key meteorological factors involved in the construction (n=3 in this example). For instance, if the temperature at a certain moment is 28℃, it falls within the range [>28℃]. With a relative humidity of 50%, the ozone concentration falls within the range of 45% to 60%, where the average ozone concentration is 78 μg / m³. The wind speed is 1.0 m / s, falling within the [≤1.2] range, where the average ozone concentration is 95 μg / m³. Then OPMI = 1.41×0.35 + 0.92×0.28 + 1.12×0.20 = 0.4935 + 0.2576 + 0.224 = 0.975. This value is close to 1, indicating that the comprehensive impact of meteorological conditions on ozone generation is close to the average level.

[0042] S5. Based on the independent verification dataset, statistically analyze the ozone concentration distribution, exceedance rate, and accuracy rate under different OPMI intervals, divide the meteorological potential levels of ozone pollution, and verify the forecasting reliability of OPMI through the correlation coefficient, mean deviation, and root mean square error.

[0043] It should be noted that the method for dividing the meteorological potential levels of ozone pollution in step S5 is as follows: Based on the independent verification dataset, statistically analyze the ozone exceedance rate and extreme value risk corresponding to the sorted OPMI values from small to large, and divide the OPMI interval into four levels: low potential, medium potential, high potential, and extremely high potential. Use the correlation coefficient R to evaluate the linear correlation degree between OPMI and ozone concentration, and use the mean deviation MB and root mean square error RMSE to evaluate the fitting accuracy and forecasting reliability of OPMI for ozone concentration.

[0044] As an embodiment, the specific implementation method of step S5 of the present invention is as follows: Use the independent verification dataset (the data of the latest year, such as 2023), calculate the OPMI value for each moment, and conduct a comparative analysis with the actual ozone concentration.

[0045] Level division: Sort all the OPMI values in the verification dataset from small to large, divide them into four equal intervals, or divide them according to the mutation points of the ozone exceedance rate. In this embodiment, statistically analyze the ozone exceedance rate (the proportion of samples with ozone concentration exceeding 160 μg / m³) and the average ozone concentration corresponding to each OPMI interval, and determine the following thresholds: Low potential: OPMI ≤ 0.75, exceedance rate < 10%; Medium potential: 0.75 < OPMI ≤ 0.95, exceedance rate 10% - 30%; High potential: 0.95 < OPMI ≤ 1.15, exceedance rate 30% - 60%; Extremely high potential: OPMI > 1.15, exceedance rate > 60%; Verification indicators: Calculate the correlation coefficient R between OPMI and ozone concentration. In this embodiment, R reaches 0.78, indicating a high positive correlation between the two.

[0046] Calculate the mean deviation MB: , in this embodiment, MB is -0.02, close to zero, indicating no systematic deviation.

[0047] Calculate the root mean square error (RMSE): In this embodiment, the RMSE is 0.18, indicating a small prediction error.

[0048] The above verification shows that the OPMI index constructed in this invention can effectively characterize the meteorological potential of ozone pollution and has high forecast reliability.

[0049] It should be noted that the method also includes: For different cities, localized parameters are trained according to steps S1 to S5 to form a localized sub-index parameter system applicable to each city.

[0050] As one embodiment, this invention can be extended to different cities. The specific implementation is as follows: For each city within the target area (e.g., nine cities including Wuhan, Huangshi, and Ezhou in the Wuhan metropolitan area), steps S1 to S5 are executed independently. Each city uses its local historical ozone and meteorological data to train a city-specific database. Random forest model and corresponding SHAP weight coefficients ( ); Optimal binning intervals and interval sub-indices for each key meteorological factor ( ); The mapping threshold between OPMI and ozone potential level.

[0051] Due to differences in geographical location, climate characteristics, and pollution source distribution among cities, localized parameters will vary. For example, Wuhan, as a large city, experiences a significant heat island effect, so temperature may have a higher weighting; while Huangshi may be affected by regional transport, making wind direction a more prominent factor. The parameter system of this invention can be stored as a configuration file, and the corresponding parameter set can be called based on the city identifier during business applications.

[0052] It should be noted that the method also includes business application steps: S61. Real-time access to meteorological observation data of the target city, and perform the same standardized preprocessing on the real-time meteorological data as in S1; Specifically, a data interface is established with meteorological observation stations to receive hourly data on temperature, relative humidity, wind speed, wind direction, air pressure, and precipitation in real time. For newly acquired real-time data, outlier checks are first performed (e.g., an alarm is triggered if the temperature exceeds -50 to 50°C), and then Z-score transformation is performed using the same standardized parameters as the modeling dataset (i.e., pre-stored historical mean and standard deviation).

[0053] S62. Substitute the preprocessed real-time meteorological data into the random forest model determined in S2 to identify the value range of each meteorological factor, retrieve the optimal binning interval determined in S3 and the sub-index calculated in S4, and calculate OPMI in real time according to the weighted summation formula in S4. Specifically, the standardized real-time feature vector is input into the pre-trained random forest model (optional; if only bin identification is required, the model can directly determine the interval to which the data falls). The model then determines which optimal bin interval each key meteorological factor currently belongs to and looks up the corresponding sub-index in a table. Combined with pre-stored weighting coefficients Calculate OPMI = Σ(K×W). The calculation process can be completed in seconds.

[0054] S63. Based on the potential levels defined in S5, output the ozone pollution meteorological potential level forecast results for the current time and future time periods.

[0055] Specifically, based on the real-time OPMI value and pre-stored level thresholds (low, medium, high, and extremely high potential), the current ozone pollution meteorological potential level is output. Simultaneously, using 24-hour meteorological forecast data provided by numerical weather prediction models, the above steps can be repeated to calculate the future OPMI value hourly and output the forecast level. The forecast results can be presented in text, color-coded (green, yellow, orange, red), or GIS map overlay format, providing decision support for ecological and environmental management departments.

[0056] Example 2 Please see Figure 2 , Figure 2 This is a schematic diagram of the hardware device in operation according to an embodiment of the present invention. The hardware device specifically includes: an ozone pollution meteorological condition index construction device 401, a processor 402, and a storage medium 403.

[0057] An ozone pollution meteorological condition index construction device 401: The ozone pollution meteorological condition index construction device 401 implements the ozone pollution meteorological condition index construction method.

[0058] Processor 402: The processor 402 loads and executes the instructions and data in the storage medium 403 to implement the method for constructing an ozone pollution meteorological condition index.

[0059] Storage medium 403: The storage medium 403 stores instructions and data; the storage medium 403 is used to implement the method for constructing an ozone pollution meteorological condition index.

[0060] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for constructing an ozone pollution meteorological condition index, characterized in that: Includes the following steps: S1. Obtain historical ozone observation data and historical meteorological observation data from multiple cities, perform missing data imputation, outlier removal and standardization preprocessing on the historical ozone observation data and historical meteorological observation data, and divide them into modeling dataset and independent validation dataset according to time series stratified sampling. S2. Using ozone concentration in the modeling dataset as labels and meteorological factors as candidate features, a random forest regression model is constructed. The model parameters are optimized through grid search cross-validation. SHAP interpretability analysis is introduced to calculate the absolute SHAP value of each meteorological factor. After normalization, the weight coefficients of each meteorological factor are obtained. S3. For the key meteorological factors selected in S2, four binning strategies are used for comparison: equal-width binning, equal-frequency binning, clustering binning, and decision tree optimal binning. The objective functions are the significance of the difference in ozone concentration between intervals, the monotonicity of the intervals, and the physical rationality. A physical trend consistency penalty term based on the principle of atmospheric photochemistry is introduced into the objective function to select the optimal binning interval that conforms to the physical monotonicity law for each key meteorological factor. S4. Based on the optimal binning intervals obtained in S3, calculate the interval sub-index corresponding to each interval for each meteorological factor. The interval sub-index is the ratio of the average ozone concentration in the interval to the total average ozone concentration during the study period. Then, the interval sub-index corresponding to each meteorological factor falling into the interval at the same time is weighted and summed with the weighting coefficients obtained in S2 to obtain the weighted comprehensive ozone pollution meteorological condition index OPMI. S5. Based on the independent validation dataset, the ozone concentration distribution, exceedance rate and accuracy under different OPMI intervals are statistically analyzed, the meteorological potential level of ozone pollution is classified, and the forecast reliability of OPMI is verified by correlation coefficient, average deviation and root mean square error.

2. The method for constructing an ozone pollution meteorological condition index as described in claim 1, characterized in that: Step S1 further includes: S11. Collect hourly ozone observation data from multiple cities, as well as hourly meteorological observation data on temperature, relative humidity, wind direction, wind speed, and air pressure during the same period. S12. Remove abnormal records where the ozone concentration is less than the first preset value or greater than the second preset value. Use spatiotemporal interpolation to interpolate continuous or intermittent missing data and standardize meteorological factors to eliminate the influence of dimensions. S13. Divide the data from the previous four years into a modeling dataset and the data from the latest year into an independent validation dataset.

3. The method for constructing an ozone pollution meteorological condition index as described in claim 1, characterized in that: Step S2 further includes: S21. Using air temperature, relative humidity, wind direction, wind speed, air pressure, and precipitation as candidate feature factors, and ozone hourly concentration as the output label, a random forest regression model is constructed. S22. The grid search and cross-search validation method is used to optimize the parameters of the number of trees, maximum depth and minimum number of leaf nodes of the random forest, so as to obtain the random forest regression model with the best fitting ability and generalization ability. S23. Calculate the SHAP value of each meteorological factor based on the optimal model, normalize the absolute value of the SHAP of each meteorological factor, and obtain the weight coefficient of each meteorological factor.

4. The method for constructing an ozone pollution meteorological condition index as described in claim 1, characterized in that: The construction and introduction of the physical trend consistency penalty term in step S3 is as follows: S31. Based on the physical mechanism of ozone photochemical generation and atmospheric diffusion, predefine the physical a priori monotonic relationship between each key meteorological factor and ozone concentration: temperature and ozone concentration are positively correlated and monotonically increasing, relative humidity and ozone concentration are negatively correlated and monotonically decreasing, and wind speed and ozone concentration are negatively correlated and monotonically decreasing. S32. In the binning optimization process, for each candidate binning scheme, traverse all adjacent intervals and check whether there are interval pairs that violate the physical prior monotonic relationship. Record a violation event each time a violation occurs. S33. Calculate the monotonicity violation degree of this binning scheme, using the following formula: ,in Let j be the weight of the violation event. The corresponding penalty coefficient is used; the more severe the violation of physical trends, the greater the penalty coefficient. S34. The monotonicity violation degree V is incorporated as a negative indicator into the comprehensive objective function, which is expressed as: Where D is the significance index of the difference in ozone concentration between intervals, M is the interval monotonicity score, and V is the degree of monotonicity violation. These are the weighting coefficients; S35. Taking the maximization of the comprehensive objective function F as the optimization objective, select the optimal binning interval for each key meteorological factor that has strong distinguishing ability, good monotonicity and the least violation of physical trends.

5. The method for constructing an ozone pollution meteorological condition index as described in claim 1, characterized in that: The formula for calculating the interval sub-index in step S4 is as follows: in, Let i be the interval sub-index when meteorological factor i falls into the t-th interval. Let i be the average ozone concentration when meteorological factor i falls within the t-th interval. The total average concentration of ozone during the study period; The weighted composite ozone pollution meteorological conditions index (OPMI) is calculated using the following formula: in, is the normalized weighting coefficient of the i-th meteorological factor, and n is the number of key meteorological factors.

6. The method for constructing an ozone pollution meteorological condition index as described in claim 1, characterized in that: The method for classifying the meteorological potential level of ozone pollution in step S5 is as follows: Based on an independent validation dataset, the ozone exceedance rate and extreme risk corresponding to the OPMI values ​​sorted from smallest to largest were statistically analyzed, and the OPMI range was divided into four levels: low potential, medium potential, high potential and very high potential. The correlation coefficient R was used to assess the linear correlation between OPMI and ozone concentration, while the mean deviation (MB) and root mean square error (RMSE) were used to assess the fitting accuracy and forecast reliability of OPMI for ozone concentration.

7. The method for constructing an ozone pollution meteorological condition index as described in claim 1, characterized in that: The method also includes: For different cities, localized parameters are trained according to steps S1 to S5 to form a localized sub-index parameter system applicable to each city.

8. The method for constructing an ozone pollution meteorological condition index as described in claim 1, characterized in that: The method also includes business application steps: S61. Real-time access to meteorological observation data of the target city, and perform the same standardized preprocessing on the real-time meteorological data as in S1; S62. Substitute the preprocessed real-time meteorological data into the random forest model determined in S2 to identify the value range of each meteorological factor, retrieve the optimal binning interval determined in S3 and the sub-index calculated in S4, and calculate OPMI in real time according to the weighted summation formula in S4. S63. Based on the potential levels defined in S5, output the ozone pollution meteorological potential level forecast results for the current time and future time periods.

9. A storage medium, characterized in that: The storage medium stores instructions and data to implement the method for constructing an ozone pollution meteorological condition index as described in any one of claims 1 to 8.

10. A device for constructing an ozone pollution meteorological condition index, characterized in that: include: A processor and a storage medium; the processor loads and executes instructions and data in the storage medium to implement the method for constructing an ozone pollution meteorological condition index as described in any one of claims 1 to 88.