Logarithmic mode temperature forecast correction method based on PM2.5 concentration

By correcting the numerical model temperature forecast based on a double regression model of PM2.5 concentration, the problems of low efficiency and insufficient accuracy of temperature forecast in existing technologies are solved, high-precision temperature forecast is achieved, forecast deviations under different pollution conditions are adapted, and the work efficiency and forecast quality of forecasters are improved.

CN120670932APending Publication Date: 2025-09-19河北省气象台
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510643241.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In the existing technology, the correction of temperature forecasts mainly relies on manual subjective judgment, which is inefficient and cannot meet the needs of rapid correction of large numbers of documents. In addition, the forecast deviation is significant under different pollution conditions, making it difficult to achieve high-precision temperature forecasts.

Method used

Through a dual regression model based on PM2.5 concentration, historical data are analyzed using univariate linear regression and divided into high-concentration and low-concentration samples. The regression equations yhigh=a1·x+b1 and ylow=a2·x+b2 are established respectively. The corresponding regression model is selected according to the PM2.5 concentration to correct the temperature forecast. Combined with the sliding training period strategy, the model parameters are dynamically updated.

Benefits of technology

It has improved the accuracy and reliability of temperature forecasts, reduced the workload of forecasters, achieved a shift from subjectivity to objectivity, improved forecast efficiency and quality, and significantly improved forecast accuracy, especially under heavy pollution conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670932A_ABST
    Figure CN120670932A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of weather forecast correction, in particular to a PM2.5 concentration-based logarithmic value mode temperature forecast correction method, which comprises the following steps of: S1, collecting hourly hourly temperature data and daily highest and lowest temperature data in the past three years, and simultaneously acquiring PM2.5 concentration data in a corresponding time period; s2, deviation analysis is carried out; and S3, selecting 155g / m as a PM2.5 concentration threshold value, and dividing the data sample into a high-concentration sample (Chigh) and a low-concentration sample (Craw). And S4, determining which regression model is used to carry out temperature forecast correction by predicting the PM2.5 concentration of the current day. According to the method, on the basis of the statistical relation between the PM2.5 concentration and the temperature forecast deviation, the forecast deviation under different pollution concentration conditions is coped with through a double regression model strategy, effective correction of traditional numerical mode forecast is achieved, and the accuracy and reliability of temperature forecast are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of weather forecast correction, and in particular to a temperature forecast correction method based on a logarithmic model of PM2.5 concentration. Background Art

[0002] According to the overall strategy of the China Meteorological Administration's "2017 Work Plan for National Intelligent Grid Weather Forecasting," a seamless intelligent grid forecast service from minutes to 10 days should be established, implementing a "two-level integration, three-tiered layout" intelligent grid forecasting strategy. This strategy will establish a national-provincial-municipal-county three-tiered service structure, with forecasting services primarily centralized at the national and provincial levels. Provincial levels should revise guiding grid forecast products according to the principle of "mandatory revisions for days 0-3, and as-needed revisions for days 4-10."

[0003] Currently, the four elements of precipitation, temperature, wind, and phase state are corrected three-hourly within a 72-hour period, requiring a total of 96 files to be corrected. The maximum and minimum temperatures are corrected 24-hourly within a 72-hour period, requiring a total of 6 files to be corrected. This results in over 100 files requiring correction. Assuming a manual correction of each file takes two minutes, this would require over 200 minutes, or more than three hours. Furthermore, as operations progress, more forecast elements will require correction, which is clearly impossible to accomplish with purely manual, subjective corrections by forecasters. This requires strengthened research into objective correction technology, promoting a shift in forecast correction methods from primarily subjective correction supplemented by objective correction to primarily objective correction supplemented by subjective correction, with the majority of corrections being completed using objective methods.

[0004] To address the above issues, the present invention uses temperature forecasting as its research subject. Through comprehensive analysis and continuous testing of historical and forecast data, it proposes a method for correcting numerical model temperature forecasts based on PM2.5 concentration. This objective correction technique can accomplish a large amount of work that cannot be done manually. Simultaneously, the experience of local forecasters is incorporated into the objective method, significantly improving the efficiency of forecast corrections. Taking temperature forecasting as a breakthrough, the present invention utilizes the interplay between meteorological conditions and air pollution levels to propose a method for correcting numerical model temperature forecasts based on PM2.5 concentration, producing high-quality temperature forecasts. Summary of the Invention

[0005] The present invention aims to provide a method for correcting temperature forecasts based on logarithmic models of PM2.5 concentration. Two regression models (high-concentration and low-concentration models) are established based on PM2.5 concentration to correct temperature forecasts. This strategy allows for flexible adjustment of forecast deviations to address varying environmental pollution conditions, improving the accuracy of the corrected temperature forecasts.

[0006] In order to achieve the above technical objectives and the above technical effects, the present invention is implemented through the following technical solutions:

[0007] A method for correcting temperature forecasts based on a logarithmic model of PM2.5 concentration includes the following steps:

[0008] S1: Collect hourly temperature data and daily maximum and minimum temperature data for the past three years, and obtain PM2.5 concentration data for the corresponding time period;

[0009] S2: Perform deviation analysis and numerically predict temperature (T model ) and actual temperature (T actual ) is defined as the deviation D, that is, D = T model -T actual ;Analyze the relationship between PM2.5 concentration and deviation D in historical data through a univariate linear regression model;

[0010] S3: Based on the results of the statistical analysis in step S2, 155 μg / m 3 As the threshold of PM2.5 concentration, the data samples are divided into high concentration samples (C high ) and low concentration samples (C low ). Construct a linear regression model for each type of sample:

[0011] For high concentration samples, establish equation y high =a1·x+b1;

[0012] For low concentration samples, establish equation y low =a2·x+b2;

[0013] Here, x represents the forecast temperature of the original numerical model, y represents the corrected temperature after regression, and a and b are regression coefficients obtained from historical data. A sliding training period is used to select samples, that is, the 30 days before the forecast date, the previous year, and the 30 days before and after the same date two years ago are used as training samples. This selection strategy ensures the similarity of the training samples with the climate conditions on the forecast day, thereby improving the adaptability of the model.

[0014] S4: Decide which regression model to use for temperature forecast correction by predicting the PM2.5 concentration of the day. If the PM2.5 concentration prediction value of the day is greater than 155μg / m 3 , select the high concentration regression model y high =a1·x+b1 for correction; otherwise, select the low concentration model y low =a2·x+b2, the temperature T predicted by the numerical model model Substitute into the corresponding regression equation and calculate the revised temperature forecast value.

[0015] The present invention is based on the discovery of an obvious phenomenon in normal forecasts. Under heavy pollution weather conditions such as haze, numerical forecasts are significantly over-predicted, that is, the forecast deviation pattern is different from that under non-polluted weather conditions. This is because numerical forecasts do not fully consider the aerosol conditions in the near-surface atmosphere when processing near-surface temperature forecasts. Therefore, the present invention proposes a two-classification model method to correct the numerical model temperature forecast. The innovation of the present invention lies in making full use of the mutual influence mechanism of meteorological conditions and environmental conditions. The study found that at 155μg / m 3 Using the PM2.5 concentration as the cutoff point, the samples are divided into two categories. Regression modeling establishes a relationship between numerical forecasts and actual conditions based on historical samples, and uses the principle of least squares to find the minimum systematic deviation. When selecting such a model, the samples used in the modeling determine its applicability. Obviously, using inappropriate samples for modeling will result in poor forecast results. Therefore, to make the model more targeted, we studied the patterns of numerical forecast deviation and, based on this, employed classification modeling to produce objective forecasts.

[0016] Based on the level of PM2.5, the present invention divides the modeling samples into two categories: one is that the PM2.5 concentration is greater than 155μg / m 3 The first category is samples with PM2.5 concentration less than or equal to 155μg / m 3 In this way, the modeling samples are divided into two categories, that is, samples with similar deviation patterns are put together. The model trained with these samples is more accurate and can better reflect the deviation patterns under such conditions. In the forecast stage, with the help of environmental meteorological forecast, if the predicted PM2.5 concentration is greater than 155μg / m 3 , the model obtained by modeling with high concentration samples is used. If the predicted PM2.5 concentration is less than or equal to 155μg / m 3 , the model obtained by modeling with low-concentration samples is used. This increases the pertinence of the prediction model.

[0017] Based on the statistical relationship between PM2.5 concentration and temperature forecast deviation, the present invention adopts a dual regression model strategy to deal with the forecast deviation under different pollution concentration conditions. At the same time, with the help of a dynamic sliding training period, the adaptability and accuracy of the model are optimized, thereby achieving effective correction of traditional numerical model forecasts and improving the accuracy and reliability of temperature forecasts.

[0018] Beneficial effects of the present invention:

[0019] This paper analyzes the statistical relationship between PM2.5 concentration and temperature forecast deviation in depth, identifying the positive deviation trend under high PM2.5 concentration conditions and the negative deviation trend under low concentration conditions. This statistical relationship lays the foundation for building a dual regression model, that is, using different regression models under different pollution conditions to correct forecast deviations. This systematic modeling based on big data analysis not only understands the specific impact of pollutants on temperature forecasts, but also establishes two sets of models (high concentration and low concentration) to determine the impact of pollutants on temperature forecasts. high =a1·x+b1,y low =a²·x+b²), enhancing the consistency of the samples used in modeling and improving the model's relevance, thereby improving the accuracy and reliability of forecasts. The model's parameters a and b are derived from rigorous historical data statistics, scientifically capturing the impact of changes in pollutant concentrations on temperature deviations, effectively reducing forecast errors. This method produces highly accurate temperature forecasts, ranking among the top of all subjective and objective forecasts produced by the Hebei Provincial Meteorological Observatory.

[0020] This invention is based on the discovery that when PM2.5 concentration is high, the temperature forecast is always too high. By analyzing the relationship between PM2.5 concentration and the temperature forecast deviation of numerical forecast, it is found that positive deviation is more common at high concentrations and negative deviation is more common at low concentrations. The modeling samples are divided into two categories: one is that the PM2.5 concentration is greater than 155μg / m 3 The first category is samples with PM2.5 concentration less than or equal to 155μg / m 3 In this way, the modeling samples are divided into two categories, that is, samples with similar deviation patterns are put together. The model trained with these samples is more accurate and can better reflect the deviation patterns under such conditions. In the forecast stage, with the help of environmental meteorological forecast, if the predicted PM2.5 concentration is greater than 155μg / m 3 , the model obtained by modeling with high concentration samples is used. If the predicted PM2.5 concentration is less than or equal to 155μg / m 3 , the model derived from low-concentration samples is used. This increases the specificity of the forecast model from a mathematical and statistical perspective. Furthermore, by fully leveraging the interaction between meteorological and environmental conditions, higher-quality forecasts are achieved, particularly improving temperature forecasts under heavy pollution conditions with high PM2.5 concentrations, thus addressing the shortcomings of existing forecasting methods.

[0021] The present invention adopts a sliding training period method to dynamically adjust and update model parameters. By selecting the 30 days before the forecast date, the 30 days before and after the same date one year and two years ago as samples, the timeliness and similarity of the sample climate conditions are guaranteed. This innovative sample selection strategy enables the regression model to automatically update over time to respond to the latest climate change trends. This automated mechanism not only greatly reduces the manual adjustment burden on forecasters, realizes the transformation of forecast correction from the previous subjective to objective, but also improves forecast efficiency, allowing forecasters to devote more energy to other important meteorological analysis and enhance overall business capabilities.

[0022] This paper systematically collects and processes temperature data, PM2.5 concentrations, and numerical model forecast data from the past three years to build a comprehensive historical database. Based on this, big data analysis techniques are used to reveal the potential relationship between pollutants and temperature deviations. This large-scale data processing not only improves information processing capabilities but also provides a solid basis for model establishment and parameter optimization. Data preprocessing, such as data cleaning, outlier detection, and missing value processing, ensures data continuity and consistency, providing high-quality data input for subsequent linear regression analysis and making the estimation of model parameters more accurate.

[0023] By combining the rich experience of local forecasters and regional characteristics, the present invention integrates objective model results with subjective experience. This combination makes the revised temperature forecast not only scientific, but also has localized adaptability, which can better capture local climate changes and weather patterns. In addition, the technical solution design has extremely strong scalability. It is not only suitable for temperature forecasts, but also lays a technical foundation for the future expansion of objective corrections to other meteorological elements (such as precipitation, wind speed, etc.). By further introducing other meteorological factors and pollutant concentrations and constructing more complex prediction models, this solution can meet more diversified weather forecast needs and promote the transformation of weather forecasts from traditional to modern intelligent; using this method, objective temperature forecasts can be produced twice a day to support intelligent grid forecasting services.

[0024] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0026] Figure 1Schematic diagram of the forecast process

[0027] Figure 2 This is a schematic diagram of the highest temperature accuracy of this method (Hebei MOS);

[0028] Figure 3 Schematic diagram of the minimum temperature accuracy of this method (Hebei MOS). DETAILED DESCRIPTION

[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0030] Example 1

[0031] The data preparation for the temperature forecast correction method based on the logarithmic model of PM2.5 concentration described in this embodiment is as follows:

[0032] This method requires the use of actual temperature data, PM2.5 concentration data, and numerical model forecast data from the past three years.

[0033] The real-time data includes hourly temperatures since January 1, 2020, the highest and lowest temperatures in the past 24 hours twice a day (08:00 and 20:00), and hourly PM2.5 concentration data.

[0034] Numerical model: The European Centre has provided array forecast results twice a day (08:00 and 20:00) since January 1, 2020. The longest forecast period each time is 240 hours. The forecast elements include hourly temperature forecast, 6-hour maximum temperature forecast, and minimum temperature forecast.

[0035] Computational forecast model:

[0036] The temperature data of the past three years and the numerical model data were used to back-calculate the forecast model using the univariate linear regression method.

[0037] The relationship between the EC fine grid forecast deviation and PM2.5 concentration is statistically analyzed. It shows that under high PM2.5 concentration conditions, the majority of deviations are positive, while under low PM2.5 concentration conditions, the majority are negative, with a weak linear relationship. The sample is averaged at 9 points and it is found that when the PM2.5 concentration is 155μg / m 3All of the above errors are positive, and further screening of the few positive deviation samples reveals a unique pattern. For the minimum temperature forecast, there is no clear relationship between the forecast deviation and PM2.5 concentration. When the samples are sorted from high to low PM2.5 concentration and then averaged over nine points, the average value is generally below 0°C, fluctuating around the average. Minimum temperature forecast deviation does not vary significantly with PM2.5 concentration.

[0038] Using the univariate linear regression method, the univariate linear regression equations for the daily maximum and minimum temperatures at each forecast time (24 hours, 48 ​​hours, 72 hours, 96 hours, 120 hours, 144 hours, 168 hours, 192 hours, 216 hours, and 240 hours) at 142 national-level stations in Hebei Province were established:

[0039] y=a·x+b (1)

[0040] In formula (1), y is the regression correction value, x is the forecast value of the time-sensitive model, a is the regression coefficient, and b is the constant term.

[0041] In the formula, a and b are the quasi-symmetrical sliding training period method, and samples from the past three years are selected from the 30 days before the forecast date, the 30 days before and after the same date in the previous year, and the 30 days before and after the same date in the year before last. 3 As the threshold, all samples were divided into two categories, with PM2.5 concentration higher than 155μg / m 3 The first category is the PM2.5 concentration below 155 μg / m 3 As one type, separate the models and get the corresponding a, b values ​​a1, b1, a2, b2. Get the corresponding prediction equation

[0042] y=a1·x+b1 (2)

[0043] y=a²·x+b² (3)

[0044] Make forecast: In the forecast stage, according to the forecast model made previously and the PM2.5 concentration predicted by the environmental meteorological forecast, if the PM2.5 concentration is greater than 155μg / m 3 Using formula (2), for the predicted PM2.5 concentration below 155μg / m 3 The numerical model prediction results are substituted into the corresponding forecast model using formula (3) to produce hourly temperature, daily maximum and daily minimum temperature forecasts.

[0045] Example 2

[0046] The method for correcting a temperature forecast based on a logarithmic model of PM2.5 concentration described in this embodiment includes the following steps:

[0047] During the data preparation phase, we first collected hourly temperature data and daily maximum and minimum temperature data for the past three years, along with PM2.5 concentration data for the corresponding time periods. The underlying logic of this step is to accumulate sufficient historical data to support subsequent statistical analysis and model building. Collecting numerical model forecast data from the European Center is also a key task during this phase, aiming to provide a basic temperature forecast for subsequent error analysis and bias correction.

[0048] Entering the data analysis and regression model construction stage, the first step is to conduct deviation analysis. Specifically, the deviation D is defined as the numerical forecast temperature (T model ) and actual temperature (T actual ), that is, D=T model -T actual Using a linear regression model to analyze the relationship between PM2.5 concentration and the deviation D in historical data, we found that when PM2.5 concentrations were high, the deviation was mostly positive, while at lower concentrations, the deviation tended to be negative. This reflects the impact of air pollution levels on temperature forecast accuracy and guides the construction of subsequent models.

[0049] In data classification and linear regression, based on the results of statistical analysis, 155 μg / m 3 As the threshold of PM2.5 concentration, the data samples are divided into high concentration samples (C high ) and low concentration samples (C low ). Build a linear regression model for each type of sample: For high concentration samples, establish equation y high =a1·x+b1; For low concentration samples, establish equation y low = a²·x+b². Here, x represents the original numerical model's forecast temperature, y represents the corrected temperature after regression, and a and b are regression coefficients derived from historical data. A sliding training period is used to select samples, using the 30 days before the forecast date, the previous year, and the 30 days before and after the same date two years ago as training samples. This selection strategy ensures similarity between the training samples and the forecast day's climate conditions, thereby improving the model's adaptability.

[0050] During the forecasting phase, the PM2.5 concentration of the day is predicted to determine the corresponding regression model for temperature forecast correction. If the PM2.5 concentration forecast value for the day is greater than 155μg / m 3 , select the high concentration regression model y high =a1·x+b1 for correction; otherwise, select the low concentration model y low =a2·x+b2. The temperature T predicted by the numerical model model Substitute into the corresponding regression equation and calculate the revised temperature forecast value.

[0051] Based on the statistical relationship between PM2.5 concentration and temperature forecast deviation, a dual regression model strategy is adopted to deal with the forecast deviation under different pollution concentration conditions. At the same time, with the help of a dynamic sliding training period, the adaptability and accuracy of the model are optimized, thereby achieving effective correction of traditional numerical model forecasts and improving the accuracy and reliability of temperature forecasts.

[0052] In this embodiment, step S1 includes the following sub-steps:

[0053] S1.1: Collect real-time temperature data: Utilize automatic weather stations and manual observation data to ensure the reliability and accuracy of data sources. Record hourly temperatures, as well as the maximum and minimum temperatures at 08:00 and 20:00 daily.

[0054] S1.2: Collect PM2.5 concentration data: Obtain hourly PM2.5 concentration values ​​from environmental protection departments or meteorological monitoring stations. The data must strictly correspond to the temperature data in time.

[0055] S1.3: Obtain numerical model forecast data: Obtain numerical model temperature forecasts twice daily from the European Centre, ensuring that the forecast data cover hourly temperatures, maximum temperatures within 6 hours, and minimum temperatures.

[0056] S1.4: Data cleaning and correction: All collected data are tested for outliers, data points with obvious errors are removed, missing values ​​are filled or eliminated, and the continuity and consistency of the data are ensured.

[0057] S1.5: Data calibration and alignment: Synchronize the temperature data with the PM2.5 concentration data time axis to ensure the temperature at each time point is T actual (t), numerical model predicted temperature T model (t), and PM2.5 concentration C PM2.5 (t) Strict correspondence is required to facilitate subsequent analysis and modeling.

[0058] Example 3

[0059] The data preparation phase of the PM2.5 concentration-based logarithmic model temperature forecast correction method described in this embodiment specifically includes the following steps:

[0060] 1. Collection of real-time temperature data:

[0061] Hourly temperature data: Starting January 1, 2020, actual hourly temperature data will be collected. This means 24 hourly temperature records will be obtained each day. To ensure data continuity and accuracy, the collected data will undergo quality control to remove outliers and missing values.

[0062] Daily maximum and minimum temperature data: Record the daily maximum and minimum temperatures at 8:00 AM and 8:00 PM. This data also undergoes quality control to ensure consistency and integrity with hourly temperature data.

[0063] 2. Collection of PM2.5 concentration data:

[0064] Hourly PM2.5 concentration data: Starting January 1, 2020, hourly PM2.5 concentration data will also be recorded. The time points of these data must be synchronized with the hourly temperature data to ensure that the temperature and PM2.5 concentration data are aligned and matched on the time axis in subsequent analysis.

[0065] 3. Collection of numerical model forecast data:

[0066] Numerical model forecast data from the European Centre: The data covers forecasts issued twice daily (at 08:00 and 20:00) since January 1, 2020. The forecast time span is up to 240 hours and includes:

[0067] Hourly temperature forecast: temperature forecast value for each hour.

[0068] 6-hour maximum and minimum temperature forecast: Provides the maximum and minimum temperatures forecast within a 6-hour interval.

[0069] Quality-controlled data: Forecast data must be quality-checked to ensure it is free of missing data, has reasonable margins of error, and is time-aligned with actual temperature and PM2.5 data.

[0070] It should be understood that by collecting hourly temperature data and PM2.5 concentration data over the past three years, the adequacy of the sample size can be ensured, thus providing a solid data foundation for subsequent statistical analysis.

[0071] Collecting daily maximum and minimum temperature data can help identify extremes in daily temperature fluctuations, thereby enhancing the understanding of temperature forecast biases.

[0072] Collecting numerical model forecast data can provide basic forecast data for subsequent model construction and facilitate deviation analysis and correction.

[0073] In this embodiment, the data analysis and regression model building phase specifically includes:

[0074] 1. Definition and calculation of deviation:

[0075] The temperature forecast deviation D is defined as the numerical forecast temperature (T model ) minus the actual temperature (T actual ). Formula:

[0076] D=T model -T actual

[0077] The deviation D reflects the error between the predicted temperature and the actual temperature and is used as the dependent variable in the model analysis.

[0078] 2. Analysis of the relationship between PM2.5 concentration and deviation:

[0079] Collect and organize historical data, including actual temperature, numerical forecast temperature and PM2.5 concentration.

[0080] A scatter plot between PM2.5 concentration and deviation D was drawn using a visualization tool to preliminarily observe the relationship between them.

[0081] Perform a univariate linear regression analysis and set the PM2.5 concentration (C PM2.5 ) is the independent variable, deviation D is the dependent variable, and a linear regression equation is constructed:

[0082] D=beta0+beta1·C PM2.5

[0083] Calculate the coefficients beta0 and beta1 in the regression analysis and verify their significance in the model through statistical tests (such as t-test).

[0084] 3. Data classification:

[0085] According to the PM2.5 concentration, the setting is 155 micrograms per cubic meter (μg / m 3 ) is the threshold value, and the data samples are divided into two categories:

[0086] High concentration samples (C high ):PM2.5 concentration is greater than 155μg / m 3

[0087] Low concentration samples (C low ): PM2.5 concentration is less than or equal to 155μg / m 3

[0088] 4. Linear regression model construction:

[0089] Construct linear regression models for two types of samples respectively:

[0090] For high concentration samples (C high ), using the formula:

[0091] y high =a1·x+b1

[0092] For low concentration samples (C low ), using the formula:

[0093] y low =a2·x+b2

[0094] In these formulas, x represents the temperature predicted by the original numerical model, y represents the temperature after correction by the model, and a1, b1, a2, and b2 are the regression coefficients obtained from data analysis.

[0095] 5. Determination of regression coefficient:

[0096] Use a sliding training period to select representative samples for training:

[0097] A sample of 30 days prior to the forecast day is included to ensure that recent climate conditions are taken into account.

[0098] Data samples from 30 days before and after the same date in the previous year and 30 days before and after the same date two years ago were selected to capture changes in annual climate patterns.

[0099] The least squares method (OLS) is used to calculate the regression coefficients a1, b1, a2, and b2, so that the model fits the historical samples optimally.

[0100] In this embodiment, the data classification and linear regression stage specifically includes:

[0101] 1. Implement a sliding training period strategy

[0102] Time window selection

[0103] Dynamic time frame: Use three time window data to train and update the model to increase the generalization ability of the model:

[0104] Data from the last 30 days is used to reflect the latest climate and pollution trends.

[0105] Capture seasonal climate characteristics 30 days before and after the same day last year.

[0106] The 30 days before and after the same day two years ago are used to account for year-to-year variations.

[0107] Continuous update strategy

[0108] Parameter update: As new data is ingested, model parameters are updated daily to ensure that the model responds to the latest changes in meteorological and pollutant concentrations.

[0109] 2. Model verification and update

[0110] Cross-validation: Split the data into training and test sets, and use cross-validation to evaluate the model to ensure its robustness on independent data.

[0111] Calculate the evaluation indicators: root mean square error (RMSE) and mean absolute error (MAE) to evaluate the accuracy of the model:

[0112]

[0113]

[0114] Based on these indicators, the prediction ability of the model under different PM2.5 concentration conditions was determined.

[0115] Based on the validation results, the classification thresholds and model parameters are adjusted as necessary. Real-time data is introduced for online updates to ensure that the model adapts to the changing environment.

[0116] In this embodiment, the forecast making stage specifically includes:

[0117] Obtain the PM2.5 concentration forecast value for the day from the environmental meteorological forecast system, select the high concentration or low concentration model according to the PM2.5 concentration forecast value classification, and output the corrected temperature forecast

[0118] 1. Hourly temperature forecast

[0119] For each hourly temperature forecast value T modelhourly , and use the corresponding PM2.5 concentration model for correction:

[0120] If PM2.5 predict >155, then T correctedhourly =a1·T modelhourly +b1

[0121] If PM2.5 predict ≤155, then T correctedhourly =a2·T modelhourly +b2

[0122] The daily maximum temperature forecast is the same as the daily minimum temperature forecast.

[0123] 2. PM2.5 concentration will significantly affect the deviation of numerical model temperature forecast.

[0124] Under high-concentration pollution conditions, the temperature forecast deviation is usually positive, while under low-concentration conditions, the deviation is mostly negative.

[0125] Therefore, different regression models are selected to correct the temperature forecast to make it more accurate.

[0126] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. A method for correcting temperature forecasts based on a logarithmic model of PM2.5 concentration, characterized in that: The following steps are involved: S1: Collect hourly temperature data and daily maximum and minimum temperature data for the past three years, and obtain PM2.5 concentration data for the corresponding time period; S2: Perform deviation analysis and numerically predict temperature T model With the actual temperature T actual The difference between them is defined as the deviation D, that is, D = T model - T actual ;Analyze the relationship between PM2.5 concentration and deviation D in historical data through a univariate linear regression model; S3: Based on the results of the statistical analysis in step S2, 155µg / m³ is selected as the threshold of PM2.5 concentration, and the data samples are divided into high-concentration samples C high and low concentration sample C low ; Construct a linear regression model for each type of sample: S4: Determine which regression model to use for temperature forecast correction by predicting the PM2.5 concentration for the day.

2. The method for correcting temperature forecasts based on a logarithmic model of PM2.5 concentration according to claim 1, wherein: The step S1 includes the following sub-steps: S1.1: Collect real-time temperature data: Use automatic weather stations and manual observation data to ensure the reliability and accuracy of data sources; record hourly temperatures, as well as the maximum and minimum temperatures at 08:00 and 20:00 daily; S1.2: Collect PM2.5 concentration data: Obtain hourly PM2.5 concentration values ​​from environmental protection departments or meteorological monitoring stations. The data must strictly correspond to the temperature data in time. S1.3: Obtain numerical model forecast data: Obtain numerical model temperature forecasts twice daily from the European Centre, ensuring that the forecast data covers hourly temperatures, maximum and minimum temperatures within 6 hours; S1.4: Data cleaning and correction: All collected data are tested for outliers, data points with obvious errors are removed, missing values ​​are filled or eliminated, and the continuity and consistency of the data are ensured; S1.5: Data calibration and alignment: Synchronize the temperature data with the PM2.5 concentration data time axis to ensure the temperature at each time point is T actual (t), numerical model predicted temperature T model (t), and PM2.5 concentration C PM2.5 (t) Strict correspondence is required to facilitate subsequent analysis and modeling.

3. The method for correcting temperature forecasts based on a logarithmic model of PM2.5 concentration according to claim 1, wherein: The step S2 includes the following sub-steps: Define the temperature forecast deviation D as the numerical forecast temperature T model Subtract the actual temperature T actual the differences between; formula: D = T model - T actual The deviation D reflects the error between the predicted temperature and the actual temperature and is used as the dependent variable in the model analysis; Collect and organize historical data, including actual temperature, numerical temperature forecast, and PM2.5 concentration; Use visualization tools to draw a scatter plot between PM2.5 concentration and deviation D to preliminarily observe the relationship between them; Perform a univariate linear regression analysis and set the PM2.5 concentration C PM2.5 As the independent variable and the deviation D as the dependent variable, a linear regression equation is constructed: D = beta0 + beta1 ·C PM2.5 Calculate the coefficients beta0 and beta1 in the regression analysis and verify their significance in the model through statistical tests.

4. The method for correcting temperature forecasts based on a logarithmic model of PM2.5 concentration according to claim 1, wherein: The step S3 includes the following sub-steps: For high concentration samples, establish equation y high = a1·x + b1; For low concentration samples, establish equation y low = a2·x + b2; Here, x represents the forecast temperature of the original numerical model, y represents the corrected temperature after regression, and a and b are regression coefficients obtained from historical data. A sliding training period is used to select samples, that is, the 30 days before the forecast date, the previous year, and the 30 days before and after the same date two years ago are used as training samples. This selection strategy ensures the similarity of the training samples with the climate conditions on the forecast day, thereby improving the adaptability of the model.

5. The method for correcting temperature forecasts based on a logarithmic model of PM2.5 concentration according to claim 1, wherein: The step S4 specifically includes: if the PM2.5 concentration forecast value on that day is greater than 155µg / m³, select the high concentration regression model y high = a1·x + b1 for correction; otherwise, select the low concentration model y low = a2·x + b2, the temperature T model Substitute into the corresponding regression equation and calculate the revised temperature forecast value.