Spatial downscaling method and system for obtaining national 1-kilometer daily-scale meteorological grid live analysis data of more than 30 years
By acquiring and matching meteorological data and topographic parameters, the LightGBM model was used to train the daily meteorological grid real-time analysis data of more than 30 years with small error and high accuracy, which solved the problems of large error and low accuracy in the existing technology, and achieved high-quality meteorological data set generation.
Patent Information
- Application Number
- CN202510512983.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-05
AI Technical Summary
The existing technology is difficult to provide real-time analysis data of 1 kilometer daily meteorological grids for more than 30 years with small error and high accuracy.
By obtaining meteorological data from 1990 to 2022, using ERA5-Land data and ground site observation data for training samples, combining terrain parameter data to generate training sample data sets, using LightGBM model for machine learning model training, and generating live analysis data of 1 kilometer daily scale meteorological grids nationwide for more than 30 years.
A 1-kilometer resolution meteorological grid live analysis data set in China's land area from 1990 to 2022 with small error and high accuracy has been achieved to meet practical application needs.
Smart Images

Figure CN120429640A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of meteorological data product development, and more particularly to a spatial downscaling method and system for obtaining over 30 years of national 1 km daily grid real-time analysis data. Background Art
[0002] Climate zoning is a process of categorizing global or regional climates using relevant indicators based on research objectives and industry requirements. The purpose of climate zoning is to survey and analyze the availability of climate resources across regions and promote their rational development and utilization. To meet the operational needs of agricultural climate resource surveys and zoning, the development of technologies for producing longer-term, higher-resolution grid-based real-time data is a pressing priority. Summary of the Invention
[0003] To this end, the technical problem to be solved by the present invention is to provide a spatial downscaling method and system for obtaining more than 30 years of national 1 km daily scale meteorological grid real-time analysis data, so as to obtain a 1 km resolution meteorological grid real-time analysis dataset for the Chinese land area from 1990 to 2022 with small error and high accuracy.
[0004] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0005] A spatial downscaling method for obtaining 1 km daily scale meteorological grid real-time analysis data for more than 30 years nationwide comprises the following steps: step (1), meteorological data acquisition: meteorological data from 1990 to 2022 are acquired, including ground station observation data, HRCLDAS grid real-time data, ERA5-Land data and terrain parameter data; step (2), ERA5-Land data and ground station observation data for the whole year of 2021 are selected as training samples; daily data matching is performed on the ERA5-Land data and ground station observation data of that year; the completed daily data matching samples are combined with terrain parameter data to generate a downscaling method. The training sample data set required by the model; Step (3), model training and verification: Use the training sample data set to train the constructed downscaling model for machine learning model training to form a machine learning model that meets spatial downscaling; During the model training process, the 2020 station observation data is selected as the true value to compare and evaluate the error characteristics of the grid products before and after downscaling, so as to test the model training effect; Step (4), data set generation and evaluation: Use the machine learning model that meets spatial downscaling to generate more than 30 years of national 1 km daily scale meteorological grid real-time analysis data products, and use ground station observation data and HRCLDAS grid real-time data as true values to evaluate the generated data set products.
[0006] In the above spatial downscaling method for obtaining more than 30 years of national 1 km daily scale meteorological grid real-time analysis data, in step (1), the ground station observation data are data after quality control by the China Meteorological Administration, including manual observation data from 1990 to 2010 and automatic observation data from 2010 to 2020; the HRCLDAS grid real-time data include hourly data of 2m temperature, 2m humidity, 10m wind speed and other elements from 2010 to 2022; the ERA5-Land data include hourly 2m temperature, surface temperature and ground incident solar shortwave radiation from 1990 to 2022. The terrain parameter data include digital elevation data, slope and aspect data, land cover classification data and normalized difference vegetation index data; the digital elevation data is a 1 km resolution nationwide DEM data generated by resampling the original 90-meter spatial resolution data jointly released by NASA and the Ministry of Economy, Trade and Industry of Japan; the slope and aspect data are calculated based on the digital elevation data; the land cover classification data is the Global Land Cover Classification Dataset (GLCC) provided by the United States Geological Survey; the normalized vegetation index data is a 1 km resolution vegetation index product provided by the MODIS sensor of NASA.
[0007] In the above-mentioned spatial downscaling method for obtaining the actual analysis data of the national 1 km daily scale meteorological grid for more than 30 years, in step (2), the specific method of matching the training samples is as follows: extracting the meteorological element data required for the downscaling model training from the ground station observation data, and performing statistical calculations of the daily maximum, daily minimum and daily average values of the extracted meteorological elements; extracting the meteorological element data required for the downscaling model training from the ERA5-Land data, and performing statistical calculations of the daily maximum, daily minimum and daily average values of the extracted meteorological elements; matching the preprocessed ground station observation data and the ERA5-Land data according to the daily data of each meteorological element; merging the matched daily data and connecting them with the terrain parameter data in space and time to generate the training sample data set required for the downscaling model.
[0008] The above-mentioned spatial downscaling method obtains the actual analysis data of the national 1 km daily scale meteorological grid for more than 30 years. The meteorological elements include temperature, relative humidity and wind speed.
[0009] In the above-mentioned spatial downscaling method for obtaining the actual analysis data of the national 1 km daily scale meteorological grid for more than 30 years, in step (3), before training the downscaling model, the training sample data set is quality controlled to filter out the sample data with observation and model errors within ±3 standard deviations.
[0010] In the above spatial downscaling method for obtaining the national 1 km daily scale meteorological grid real-time analysis data for more than 30 years, in step (3), the machine learning model is the LightGBM model.
[0011] The above-mentioned spatial downscaling method obtains the actual analysis data of the national 1 km daily scale meteorological grid for more than 30 years. When training the downscaling model, the sample characteristics include terrain parameter data, longitude and latitude, time, and the daily maximum, daily minimum and daily average of each meteorological element. The sample label includes the error between the daily maximum, daily minimum and daily average of each meteorological element model data and the observation data. The error indicators include correlation coefficient, mean deviation and root mean square error.
[0012] In the above-mentioned spatial downscaling method for obtaining the actual analysis data of the national 1 km daily scale meteorological grid for more than 30 years, in step (3), the method for optimizing the parameters of the number of leaves and node depth of the downscaling model is as follows: 5 copies are automatically randomly selected from the training set as training and 1 copy is used as an evaluation test through the grid search parameter adjustment method; when the node depth remains unchanged, the correction ability of the downscaling model with different numbers of leaves is compared by sliding, thereby determining the leaf number parameter; then, the correction ability of the downscaling model with different node depths is tested, thereby determining the node depth parameter.
[0013] The above-mentioned spatial downscaling method for obtaining the national 1 km daily scale meteorological grid actual analysis data for more than 30 years, in step (4), uses a machine learning model that meets the spatial downscaling requirements, inputs the ERA5-Land data from 1990 to 2022 obtained in step (1), and generates the national 1 km daily scale meteorological grid actual analysis data products for more than 30 years; the evaluation indicators include error indicators and grid image quality indicators; the error indicators include correlation coefficient, mean deviation and root mean square error, and the grid image quality indicators include peak signal-to-noise ratio, structural similarity index and learning perception image block similarity index; the ground station observation data is used as the true value to evaluate the error indicators of the generated dataset product, and the HRCLDAS grid actual data is used as the true value to evaluate the grid image quality indicators of the generated dataset product.
[0014] A spatial downscaling system for obtaining 30 years of national 1 km daily scale meteorological grid real-time analysis data is used to implement the above-mentioned spatial downscaling method for obtaining 30 years of national 1 km daily scale meteorological grid real-time analysis data, including the following modules: a data download module for completing the download of original ERA5-Land data, ground station observation data, and terrain parameter data; a data source module for completing the statistical calculation and storage of the daily maximum, daily minimum and daily average values of each meteorological element data in the HRCLDAS grid real-time data and ERA5-Land data; a model training module for completing the statistical calculation and storage of the daily maximum, daily minimum and daily average values of each meteorological element data in the ground station observation data. The statistical calculation of the daily average value is performed and matched with the ERA5-Land data for daily data; the matched daily data are then merged and spatially and temporally connected with the terrain parameter data to generate a training sample data set based on national ground stations, and finally the matched training sample data set is used for model training; the historical backcalculation module generates more than 30 years of national 1 km daily scale meteorological grid real-time analysis data products by calling the trained downscaling machine learning model and ERA5-Land data; the ground station observation data is used as the true value to evaluate various error indicators, and the HRCLDAS grid real-time data is used as the true value to evaluate various image quality indicators.
[0015] The technical solution of the present invention achieves the following beneficial technical effects:
[0016] The present invention is based on ERA5-Land reanalysis data. By extracting the characteristics of each meteorological element in the ERA5-Land reanalysis data, the daily maximum, daily minimum and daily average of each meteorological element are statistically calculated, and the data are matched with the observation data of ground stations across the country. The matched daily data are combined with terrain parameter data to generate a training sample data set required for the downscaling model; using this data set, the LightGBM model is adopted as the machine learning training model, and the parameters of the number of leaves and node depth are optimized to train a downscaling model; using this model, based on more than 30 years of ERA5-Land reanalysis data, more than 30 years of national 1 km daily scale meteorological grid real-time analysis data are developed, so that a 1 km resolution meteorological grid real-time analysis data set of the land area of China from 1990 to 2022 can be obtained. Taking the ground station observation data and HRCLDAS grid actual data as the true values, the error indicators and image quality indicators of the generated long-term downscaled meteorological grid actual analysis data were evaluated respectively. The results showed that the 1-km resolution meteorological grid actual analysis dataset for the Chinese land area from 1990 to 2022 developed by the method of the present invention has small error and high accuracy, which can meet the needs of practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Schematic diagram of the spatial downscaling process of daily-scale meteorological grid real-time analysis data in an embodiment of the present invention;
[0018] Figure 2 and Figure 3 They are respectively the data distribution diagrams of the daily average temperature samples before and after quality control in the embodiment of the present invention;
[0019] Figure 4 、 Figure 5 and Figure 6 They are respectively graphs showing the importance analysis results of the sample data of daily maximum temperature, daily average temperature and daily minimum temperature in an embodiment of the present invention;
[0020] Figure 7 、 Figure 8 and Figure 9 They are respectively graphs showing the importance analysis results of the sample data of daily maximum relative humidity, daily average relative humidity and daily minimum relative humidity in an embodiment of the present invention;
[0021] Figure 10 and Figure 11 They are respectively graphs showing the importance analysis results of the sample data of daily maximum wind speed and daily average wind speed in an embodiment of the present invention;
[0022] Figure 12 and Figure 13 These are the comparative analysis results of parameter adjustment for different leaf numbers and different node depths in the embodiments of the present invention;
[0023] Figure 14a 、 Figure 14b and Figure 14c These are the comparison results of the three error indicators of correlation coefficient, mean deviation and root mean square error of the daily maximum temperature in the embodiment of the present invention;
[0024] Figure 15a 、 Figure 15b and Figure 15c These are the comparison results of the three network image quality indicators, namely the peak signal-to-noise ratio, structural similarity index, and image similarity index of the daily maximum temperature in the embodiment of the present invention;
[0025] Figure 16a 、 Figure 16b and Figure 16c These are the comparison results of the three error indicators of correlation coefficient, mean deviation and root mean square error of the daily minimum temperature in the embodiment of the present invention;
[0026] Figure 17a 、 Figure 17b and Figure 17c These are the comparison results of the three network image quality indicators, namely, the peak signal-to-noise ratio, structural similarity index, and image similarity index of the daily minimum temperature in the embodiment of the present invention;
[0027] Figure 18a 、 Figure 18b and Figure 18c These are the comparison results of the three error indicators of correlation coefficient, mean deviation and root mean square error of daily average temperature in the embodiment of the present invention;
[0028] Figure 19a 、 Figure 19b and Figure 19c These are the comparison results of the three network image quality indicators, namely the peak signal-to-noise ratio, structural similarity index, and image similarity index of the daily average temperature in the embodiment of the present invention;
[0029] Figure 20a 、 Figure 20b and Figure 20c These are the comparison results of the three network image quality indicators, namely the peak signal-to-noise ratio, structural similarity index, and image similarity index, of the maximum daily relative humidity in an embodiment of the present invention;
[0030] Figure 21a 、 Figure 21b and Figure 21c These are the comparison results of the three error indicators of correlation coefficient, mean deviation and root mean square error of daily minimum relative humidity in the embodiment of the present invention;
[0031] Figure 22a 、 Figure 22b and Figure 22c These are the comparison results of the three network image quality indicators, namely the peak signal-to-noise ratio, structural similarity index, and image similarity index, for the daily minimum relative humidity according to an embodiment of the present invention;
[0032] Figure 23a 、 Figure 23b and Figure 23c These are the comparison results of the three error indicators of correlation coefficient, mean deviation and root mean square error of the daily average relative humidity in the embodiment of the present invention;
[0033] Figure 24a 、 Figure 24b and Figure 24c These are the comparison results of the three network image quality indicators, namely the peak signal-to-noise ratio, structural similarity index, and image similarity index of the daily average relative humidity in the embodiment of the present invention;
[0034] Figure 25a 、 Figure 25b and Figure 25c These are the comparison results of the three error indicators of correlation coefficient, mean deviation and root mean square error of the daily maximum wind speed in the embodiment of the present invention;
[0035] Figure 26a 、 Figure 26b and Figure 26cThese are the comparison results of the three network image quality indicators, namely the peak signal-to-noise ratio, structural similarity index, and image similarity index of the maximum daily wind speed in the embodiment of the present invention;
[0036] Figure 27a 、 Figure 27b and Figure 27c These are the comparison results of the three error indicators of correlation coefficient, average deviation and root mean square error of daily average wind speed in the embodiment of the present invention;
[0037] Figure 28a 、 Figure 28b and Figure 28c These are the comparison results of the three network image quality indicators evaluation: peak signal-to-noise ratio, structural similarity index, and image similarity index of the daily average wind speed in the embodiment of the present invention. DETAILED DESCRIPTION
[0038] 1. Spatial downscaling method for meteorological grid live analysis data
[0039] 1.1 Technical Route
[0040] This example uses a machine learning downscaling method based on ERA5-Land reanalysis data to complete machine learning model training and realize the development of a daily scale real-time analysis dataset. The specific method flow is as follows Figure 1 As shown, it mainly includes the following steps:
[0041] (1) Meteorological data acquisition: Develop a data download program to obtain the daily maximum, minimum, and average temperatures, daily maximum, minimum, and average relative humidity, and daily maximum and average wind speeds required for the task through the ERA5 official website data access interface. The element data required for downscaling include air temperature, dew point temperature, humidity, wind speed, surface air temperature, soil temperature, and solar radiation.
[0042] (2) Training sample matching: The ERA5-Land daily maximum, minimum, and average temperature, daily maximum, minimum, and average relative humidity, daily maximum and average wind speed data and the station observation daily maximum, minimum, and average temperature, daily maximum, minimum, and average relative humidity, daily maximum and average wind speed data were processed respectively. After data matching, the sample data required for the downscaling model, such as daily maximum, minimum, and average temperature, daily maximum, minimum, and average relative humidity, daily maximum and average wind speed, were generated in combination with other auxiliary data such as terrain elevation.
[0043] (3) Model training and testing: The machine learning model is trained using the matched sample data. The model effectiveness is tested based on the sample data groups to form a machine learning model that meets the requirements of spatial downscaling. The station observation data are used as the true value to compare and evaluate the error characteristics of the grid products before and after downscaling.
[0044] (4) Dataset development and evaluation: Based on the trained model, the downscaling calculation of each element is performed to develop a 1 km resolution daily maximum, minimum, and average temperature, daily maximum, minimum, and average relative humidity, and daily maximum and average wind speed real-time analysis dataset for the land area of China from 1990 to 2022. The overall quality of the historical dataset is evaluated according to the evaluation indicators. The grid product comparison evaluation is conducted using the ground station observation data and the 1 km resolution multi-source fusion real-time analysis product of the National Meteorological Information Center as the true value. The evaluation indicators include error indicators and grid image quality indicators. Error indicators include correlation coefficient, mean deviation, root mean square error, etc. Grid image quality indicators include peak signal-to-noise ratio, structural similarity index, image similarity, etc.
[0045] 1.2 Data source information
[0046] (1) Site observation data
[0047] Station data acquisition methods are divided into manual observation and automatic observation, depending on the time period. From 1990 to 2010, station observation data was primarily acquired through manual observation, with approximately 2,400 observation stations nationwide. From 2010 to 2020, station observation data was primarily acquired through automatic observation, with approximately 50,000 observation stations nationwide. The China Meteorological Administration's ground-based automatic station observation data undergoes a strict quality control process, and the meteorological station data used in this example are all quality-controlled data.
[0048] (2) HRCLDAS grid live data
[0049] The High Resolution China Meteorological Administration Land Data Assimilation System (HRCLDAS) uses 1km-resolution terrain data and employs physical inversion, terrain correction, and multigrid variational analysis techniques to produce 1km-resolution meteorological driving data. To address the high-resolution and large-data-volume characteristics of land surface simulations, a model-parallel computing solution has been designed, along with an efficient soil moisture simulation system. This system generates grid-based fusion products in real time, providing a series of near-real-time product data. Its meteorological driving field data products include 2-meter temperature, 2-meter humidity, 10-meter wind speed, surface pressure, precipitation, and shortwave radiation. Within China, these products are superior to comparable international counterparts and offer higher spatiotemporal resolution.
[0050] (3) ERA5-Land data (ERA5-Land reanalysis data)
[0051] European fifth-generation land surface reanalysis data (ERA5-Land). ERA5-Land is a high-resolution global land reanalysis dataset provided by the European Centre for Medium-Range Weather Forecasts (ECMWF). It is the land portion of the ERA5 climate reanalysis, with a higher spatial resolution, a grid spacing of approximately 9 kilometers, and provides hourly data. The ERA5-Land data coverage is from 1950 to the present, and will be continuously updated to support land monitoring applications. The dataset contains high-resolution information on a variety of surface variables, such as air temperature, evaporation, snow cover, soil moisture, etc. These data are very valuable for analyzing climate trends and anomalies, conducting hydrological research, initializing numerical weather forecasts and climate models, and supporting various applications in water resources, land, and environmental management. This example downloads three data closely related to the project elements (temperature, humidity, and wind) from ERA5-Land hourly at 2m air temperature, surface temperature, and ground incident solar shortwave radiation from 1990 to 2022. Through statistical analysis, the highest and lowest average values of each element per day are calculated. The model was trained using data from 2021 as input, and the downscaling results were evaluated using data from 2020. The trained model was then used to downscale the 1-km resolution daily maximum, minimum, and average temperatures, daily maximum, minimum, and average relative humidity, and daily maximum and average wind speed data from 1990 to 2022.
[0052] (4) Topographic parameter data
[0053] The terrain parameter information used in this embodiment includes digital elevation data, slope and aspect data, land cover classification data, normalized difference vegetation index, etc.
[0054] A. Digital Elevation Model (DEM) data is a model that represents the undulations of the Earth's surface terrain. It records the ground elevation value of each point through a regularly spaced grid of points. DEM can be regarded as a snapshot of the Earth's surface elevation, usually in digital form, so it can be quickly processed and analyzed by computers. The DEM data used in this embodiment is jointly released by the National Aeronautics and Space Administration (NASA) and the Ministry of Economy, Trade and Industry (METI) of Japan. The original data has a spatial resolution of 90 meters. After resampling, it generates nationwide DEM data with a resolution of 1 km.
[0055] B. Slope and aspect data are two important terrain characteristics in terrain analysis. Slope refers to the angle between a tangent plane at a point on the ground surface and the horizontal plane, usually expressed in degrees. A larger value indicates steeper terrain. Aspect, on the other hand, refers to the downhill direction with the greatest rate of change from each pixel to its adjacent pixel. It is a measure of the change in slope direction at a location on the ground surface. It is usually expressed as an angle ranging from 0 to 360°, measured clockwise from due north. The slope and aspect data used in this embodiment are calculated based on digital elevation data.
[0056] C. Land cover classification data. Land cover classification datasets are important datasets that describe the natural and man-made features of the Earth's surface. They have important application value for environmental change research, national geographic conditions monitoring, and sustainable development planning. The Global Land Cover Characterization (GLCC) dataset used in this example is a global land cover classification dataset provided by the United States Geological Survey (USGS). This dataset is mainly based on 1 km resolution Advanced Very High Resolution Radiometer (AVHRR) 10-day Normalized Difference Vegetation Index (NDVI) synthetic data, obtained through unsupervised classification methods.
[0057] D. Normalized Difference Vegetation Index (NDVI) data is an important tool for monitoring vegetation coverage and growth using remote sensing technology. NDVI quantifies vegetation density and health by comparing the reflectance of near-infrared and red light bands. The NDVI dataset used in this example is the MODIS NDVI dataset, a series of vegetation index products provided by the MODIS (Moderate Resolution Imaging Spectroradiometer) sensor of the National Aeronautics and Space Administration (NASA). It includes three different spatial resolutions: 250 meters, 500 meters, and 1000 meters. This example uses a 1000-meter resolution product.
[0058] 1.3 Model Method
[0059] (1) LightGBM model
[0060] Light Gradient Boosting Machine (LightGBM) is used. LightGBM is an efficient gradient boosting decision tree (GBDT) model, characterized by its speed and high memory efficiency. GBDT is an ensemble learning algorithm commonly used for classification, regression, and ranking tasks. It sequentially constructs decision trees to minimize a loss function. The training data for each round of decision trees is derived from the residuals of the previous round, and the model is continuously iterated to optimize prediction accuracy. GBDT performs well in many machine learning tasks, but as data size and feature dimensionality increase, its computational overhead and storage requirements increase, becoming a bottleneck. To accelerate training, LightGBM implements two key optimizations: gradient-based one-side sampling (GOSS) and exclusive feature bundling (EFB). Especially when working with high-dimensional, sparse features or massive amounts of data, LightGBM can significantly improve model training speed without significantly compromising accuracy compared to other GBDT implementations (such as XGBoost). In addition, LightGBM also supports histogram-based feature segmentation and leaf node priority growth strategy, making it particularly superior in processing big data scenarios. Therefore, LightGBM becomes a highly scalable machine learning model suitable for large-scale data training.
[0061] (2) Sample data generation
[0062] The data for the whole year of 2021 are selected as sample data for model training. Data quality control is performed before model training, and sample data other than those with observation and model errors within ±3 standard deviations are screened out. For detailed analysis, see Figure 2 and Figure 3 .
[0063] The same method is used to screen the quality of other factor data, and the sample size is shown in the following table:
[0064] Table 1 Statistics of the number of samples of each factor model
[0065]
[0066]
[0067] (3) Feature importance analysis
[0068] Feature importance analysis is a property used to assess the importance of features in machine learning models. It analyzes the impact of each feature in the model to determine how important each feature is to the model's predictions. This helps us better understand how the model makes predictions and can be used for feature selection, improving model performance and interpretability.
[0069] Temperature data downscaling calculation uses longitude, latitude, month (sine-cosine transformation) data, maximum, minimum, and average data of ERA5 meteorological elements (including surface temperature, 2m temperature, and radiation), and surface parameters (including vegetation index, surface type, elevation, slope, and aspect) as basic features for machine learning model training to organize sample data and perform feature importance analysis. The results of daily maximum, minimum, and average temperature analysis are shown in Figures 4 to 6 .
[0070] Humidity data downscaling calculation uses longitude, latitude, month (sine-cosine transformation) data, daily maximum, minimum, and average relative humidity data after ERA5 meteorological element conversion, and surface parameters (including vegetation index, surface type, elevation, slope, and aspect) as basic features for machine learning model training to organize sample data and perform feature importance analysis. The analysis results of daily maximum, minimum, and average relative humidity are shown in Figures 7 to 9 .
[0071] Wind speed data downscaling calculation uses longitude, latitude, month (sine-cosine transformation) data, daily maximum and average wind speed data of ERA5 meteorological elements, and surface parameters (including vegetation index, surface type, elevation, slope, and aspect) as basic features for machine learning model training to organize sample data and perform feature importance analysis. The results of daily maximum and average wind speed analysis are shown in Figures 10 and 11 .
[0072] (4) Model training and parameter optimization
[0073] In this embodiment, the model parameters are selected using the "experience + enumeration within a certain range" method (currently a common method for hyperparameter adjustment). For the LightGBM model, the most critical parameters are the number of model leaves and node depth. During the model training process, these two key parameters are mainly tuned, and the other parameters are given based on experience.
[0074] The 2021 annual average daily temperature data is used as the model sample, and 5 of them are automatically randomly selected for training and 1 for evaluation through the grid search parameter adjustment method. First, under the condition that the node depth remains unchanged, the correction ability of the model under different leaf numbers is compared. Figure 12As shown in the figure, as the number of leaves changes from 100 to 1400, the error between the corrected and observed temperatures decreases. As the number of leaves increases from 100 to 800, the root mean square error decreases from 0.044°C to 0.001°C. As the number of leaves increases from 800, the root mean square error becomes more stable. Therefore, the number of leaves is set to 800 for the maximum temperature model parameter selection.
[0075] At the same time, we further tested the correction capabilities of different node depth models, such as Figure 13 As shown in the figure, five leaf numbers of 100, 200, 300, 400, and 500 are selected to compare the correction capabilities of nodes with node depths of 10, 20, 30, 40, 50, 60, and 70. It can be seen that under different leaf numbers, when the node depth increases from 10 to 20, the root mean square error decreases significantly. When the node depth continues to increase from 20, the error is relatively stable under different leaf numbers, and the correction error is no longer sensitive to changes in node depth.
[0076] The parameters of the downscaling model for daily mean temperature are as follows:
[0077]
[0078] Through the above parameter adjustment process, the analysis process of each factor parameter is basically consistent with the daily average temperature, and the same model parameters are selected for model training.
[0079] 1.4 Product Evaluation
[0080] The product evaluation process of this embodiment uses ground station observation data and the 1-kilometer resolution multi-source fusion real-time analysis product of the National Meteorological Information Center as the true values, and compares and evaluates them with the original ERA5-Land and downscaled grid products. The evaluation indicators include error indicators and grid image quality indicators. Error indicators include correlation coefficient, mean deviation, root mean square error, etc. Grid image quality indicators include peak signal-to-noise ratio, structural similarity index, image similarity, etc.
[0081] (1) The error index is calculated as follows:
[0082] Mean error (ME):
[0083] Root Mean Square Error (RMSE):
[0084] Mean Absolute Error (MAE):
[0085] Correlation coefficient (COR):
[0086] In the above formula, O is the observation data, and G is the grid data that matches the observation time and space.
[0087] (2) The grid image quality indicators are as follows:
[0088] A. Peak Signal-to-Noise Ratio (PSNR)
[0089]
[0090] Where X1 and X2 are the data values for each grid of the two types of grid products, respectively, and MSE is the mean square error. PSNR is a metric for evaluating image quality; a larger value indicates less image distortion.
[0091] B. Structural Similarity Index Measure (SSIM)
[0092]
[0093] Where X1 and X2 are the data values of each grid of the two types of grid products, μ X1 , μ X2 are the mean values of two types of grid products, σ X1 , σ X2 are the variances of two types of grid products, σ X1X2 The covariance of the two grids, C1 and C2, are constants that default to the maximum value of the grid multiplied by 0.01 and 0.03, respectively. The SSIM value range is [0, 1], with larger values indicating more similar images.
[0094] C. The Learned Perceptual Image PatchSimilarity (LPIPS) metric
[0095]
[0096] Among them, M is the number of local blocks into which the image is divided, w j is the weight of the j-th layer neural network feature, Φ j is the activation function output of the j-th layer neural network, x1 is the downscaled grid data image, x2 is the image generated by HRCLDAS data, and p is the power of the norm.
[0097] LPIPS is used to measure the difference between two images. The lower the value, the more similar the two images are, and vice versa.
[0098] 2. Spatial downscaling of meteorological grid live analysis data
[0099] 2.1 Server Program Deployment
[0100] The program of this embodiment is deployed based on four servers to implement corresponding program functions, namely, a data download server, a data source server, a model training server, and a historical backcalculation server.
[0101] (1) Data download server
[0102] The data download server is deployed in the DMZ and primarily downloads ERA5-Land raw data and imports it to the intranet server in a one-way manner. See Table 2 for a list of related programs.
[0103] Server IP address: 10.0.66.8; Program path: / home / lwg / era5 / code / download / .
[0104] Table 2
[0105]
[0106] (2) Data source server
[0107] The data source server primarily stores raw data. Daily statistical calculations for various elements are performed on this server, and the results are copied to other computing servers. See Table 3 for a list of related programs.
[0108] Server IP address: 10.40.26.115; data mount directory: / CMADAAS_mnt (the original CMADAAS path has been mounted to another account); program path: / share_testbed / hs / .
[0109] Table 3
[0110]
[0111] (3) Model training server
[0112] The model training server is a new HPC server used for machine learning model development, training, model validation, and subsequent model inference (dataset computation) and related evaluation and mapping. This server is a computer machine that primarily provides CPU computing resources and can submit jobs to multiple compute nodes for computation. Server address: sc-2.hpc.cma; program path: / g11 / hanshuai / max_min_data; temperature, humidity, and wind speed each have their own processing program, with a consistent program structure. See Table 4 for a list of processing programs for each element.
[0113] Table 4
[0114]
[0115] (4) Historical Recalculation Server
[0116] The historical backcalculation service is used to backcalculate historical data. ERA5-Land raw data and trained models are copied to this server, and the model is invoked to backcalculate data for various meteorological elements. Historical observation data is also downloaded from this server for data evaluation and verification. Server address: pird.nmic.cn; program path: / g11 / wangql / code / ; temperature, humidity, and wind speed each have their own processing program, with a consistent structure. See Table 5 for a list of processing programs for each element.
[0117] Table 5
[0118]
[0119] 3. Production and evaluation of spatial downscaling datasets for meteorological grid live analysis data
[0120] Based on the model validation results, the trained model was used to generate a 33-year dataset from 1990 to 2022. This dataset covers a total of 12,053 days. Eight meteorological element files were generated for each day, including the daily maximum, minimum, and average temperature; the daily maximum, minimum, and average relative humidity; and the daily maximum and average wind speed. A total of 96,424 files should have been generated. The actual number of generated dataset files was consistent: 12,053 files for each element, for a total of 96,424 files, as shown in the table below.
[0121] Table 6
[0122]
[0123]
[0124] The historical data set is stored in the " / g11 / wangql / data" directory of the historical back-calculation server. Data files are stored according to three factors: temperature, relative humidity, and wind speed. The disk occupies approximately 3.71T. The occupancy of each type of data is shown in the following table:
[0125] Table 7
[0126] elements Temperature relative humidity wind speed Space occupied (T) 1.2 1.5 1.01
[0127] 3.1 Maximum temperature
[0128] The daily maximum temperature dataset is stored in the " / g11 / wangql / data / tair / eval" directory on the historical backcalculation server. The spatial resolution is 0.01° × 0.01°, the temporal resolution is 1 day, and it is stored in NetCDF format. The file name is "Pred_max_{yyyymmdd}.nc". The temperature product storage structure includes four variables: LON, LAT, Pred, and era5. LON and LAT are coordinate variables in degrees. Pred and era5 are the daily maximum temperature element variables of the downscaled product and era5, respectively, in K. The specific file structure is shown in Table 8.
[0129] Table 8
[0130]
[0131] Data from January and August 2020 (representing winter and summer, respectively) were selected for comparison, and the spatial distribution of the daily maximum temperature data was compared. The daily maximum temperature data observed at the station from 1990 to 2022 were used as the reference test source, and the error index comparison test was performed with the original ERA5-Land reanalysis data and the daily maximum temperature dataset generated by this embodiment. The correlation coefficient index ERA5-Land is 0.89, and the dataset generated by this embodiment is 0.91; the average deviation difference index ERA5-Land is -1.63℃, and the dataset generated by this embodiment is -0.49℃; the root mean square error index ERA5-Land is 3.3℃, and the dataset generated by this embodiment is 2.7℃. The specific image results are as follows: Figures 14a to 14c shown.
[0132] The HRCLDAS data from January and early August 2020 were selected as reference data, and the grid image quality was compared with the original ERA5-Land reanalysis data and the dataset generated in this example. The indicators included peak signal-to-noise ratio, structural similarity index, image similarity index, etc. The specific results are as follows: Figures 15a to 15c shown.
[0133] 3.2 Lowest temperature on the day
[0134] The daily minimum temperature dataset is stored in the directory " / g11 / wangql / data / tair / eval" on the historical backcalculation server. The spatial resolution is 0.01°×0.01° and the temporal resolution is 1 day. It is stored in NetCDF format and the file name is "Pred_min_{yyyymmdd}.nc". The daily minimum temperature data storage structure includes four variables: LON, LAT, Pred, and era5. LON and LAT are coordinate variables in degrees. Pred and era5 are the daily minimum temperature element variables of the downscaled product and era5, respectively, in K. The specific file structure is shown in Table 8.
[0135] Data from January and August 2020 (representing winter and summer, respectively) were selected for comparison, and the spatial distribution of daily minimum temperature data was compared.
[0136] The daily minimum temperature data observed at the station from 1990 to 2022 were used as a reference test source. Error index comparison tests were performed with the original ERA5-Land reanalysis data and the daily minimum temperature dataset generated by this embodiment. The correlation coefficient index of ERA5-Land was 0.9, and that of the dataset generated by this embodiment was 0.92; the mean deviation index of ERA5-Land was -0.5℃, and that of the dataset generated by this embodiment was -0.35℃; the root mean square error index of ERA5-Land was 3.21℃, and that of the dataset generated by this embodiment was 2.83℃. The specific image results are shown in Figure 2. Figures 16a to 16c shown.
[0137] The HRCLDAS data from January and early August 2020 were selected as reference data, and the grid image quality was compared with the original ERA5-Land reanalysis data and the dataset generated in this example. The indicators included peak signal-to-noise ratio, structural similarity index, image similarity index, etc. The specific results are as follows: Figures 17a to 17c shown.
[0138] 3.3 Daily average temperature
[0139] The daily mean temperature dataset is stored in the directory " / g11 / wangql / data / tair / eval" on the historical backcalculation server. The spatial resolution is 0.01° × 0.01°, the temporal resolution is 1 day, and it is stored in NetCDF format. The file name is "Pred_mean_{yyyymmdd}.nc". The storage structure of the daily mean temperature product includes four variables: LON, LAT, Pred, and era5. LON and LAT are coordinate variables in degrees. Pred and era5 are the daily mean temperature element variables of the downscaled product and era5, respectively, in K. The specific file structure is shown in Table 8.
[0140] Data from January and August 2020 (representing winter and summer, respectively) were selected for comparison, and the spatial distribution of daily average temperature data was compared.
[0141] The daily average temperature data observed at the station from 1990 to 2022 was used as a reference test source. Error index comparison tests were performed with the original ERA5-Land reanalysis data and the daily average temperature dataset generated by this embodiment. The correlation coefficient index of ERA5-Land was 0.93, and that of the dataset generated by this embodiment was 0.94; the mean deviation index of ERA5-Land was -0.5℃, and that of the dataset generated by this embodiment was -0.1℃; the root mean square error index of ERA5-Land was 2.33℃, and that of the dataset generated by this embodiment was 2.08℃. The specific image results are shown in Figure 2. Figures 18a to 18c shown.
[0142] The HRCLDAS data from January and early August 2020 were selected as reference data, and the grid image quality was compared with the original ERA5-Land reanalysis data and the dataset generated in this example. The indicators included peak signal-to-noise ratio, structural similarity index, image similarity index, etc. The specific results are as follows: Figures 19a to 19c shown.
[0143] 3.4 Daily maximum relative humidity
[0144] The daily maximum relative humidity dataset is stored in the " / g11 / wangql / data / rhus / eval" directory on the historical backcalculation server. The spatial resolution is 0.01° × 0.01° and the temporal resolution is 1 day. It is stored in NetCDF format and the file name is "MAX_RHUS_{yyyymmdd}.nc". The temperature product storage structure includes four variables: LON, LAT, Pred, and era5. LON and LAT are coordinate variables in degrees. Pred and era5 are the daily maximum relative humidity element variables of the downscaled product and era5, respectively, in %. The specific file structure is shown in Table 8.
[0145] Data from January and August 2020 (representing winter and summer, respectively) were selected for comparison, and the spatial distribution of daily maximum relative humidity data was compared.
[0146] There is no daily maximum humidity observation in the station data, so no error index evaluation and comparison is performed.
[0147] The HRCLDAS data from January and early August 2020 were selected as reference data, and the grid image quality was compared with the original ERA5-Land reanalysis data and the dataset generated in this example. The indicators included peak signal-to-noise ratio, structural similarity index, image similarity index, etc. The specific results are as follows: Figures 20a to 20c shown.
[0148] 3.5-day minimum relative humidity
[0149] The daily minimum relative humidity dataset is stored in the " / g11 / wangql / data / rhus / eval" directory on the historical backcalculation server. The spatial resolution is 0.01° × 0.01°, the temporal resolution is 1 day, and it is stored in NetCDF format. The file name is "MIN_RHUS_{yyyymmdd}.nc". The daily minimum relative humidity data storage structure includes four variables: LON, LAT, Pred, and era5. LON and LAT are coordinate variables in degrees. Pred and era5 are the daily minimum relative humidity element variables of the downscaled product and era5, respectively, in %. The specific file structure is shown in Table 8.
[0150] Data from January and August 2020 (representing winter and summer, respectively) were selected for comparison, and the spatial distribution of daily minimum humidity data was compared.
[0151] The daily minimum relative humidity data observed at the daily station from 1990 to 2022 were used as the reference test source. The error index comparison test was performed with the original ERA5-Land reanalysis data and the daily minimum relative humidity dataset generated by this embodiment. The correlation coefficient index of ERA5-Land was 0.79, and that of the dataset generated by this embodiment was 0.8; the mean deviation index of ERA5-Land was 5.1%, and that of the dataset generated by this embodiment was 4.69%; the root mean square error index of ERA5-Land was 13.19%, and that of the dataset generated by this embodiment was 12.94%. The specific image results are as follows: Figures 21a to 21c As shown.
[0152] The HRCLDAS data from January and early August 2020 were selected as reference data, and the grid image quality was compared with the original ERA5-Land reanalysis data and the dataset generated in this example. The indicators included peak signal-to-noise ratio, structural similarity index, image similarity index, etc. The specific results are as follows: Figures 22a to 22c As shown.
[0153] 3.6 Daily average relative humidity
[0154] The daily average relative humidity dataset is stored in the " / g11 / wangql / data / rhus / eval" directory on the historical backcalculation server. The spatial resolution is 0.01° × 0.01°, the temporal resolution is 1 day, and it is stored in NetCDF format. The file name is "MIN_RHUS_{yyyymmdd}.nc". The data storage structure of the daily minimum relative humidity includes four variables: LON, LAT, Pred, and era5. LON and LAT are coordinate variables in degrees. Pred and era5 are the daily average relative humidity element variables of the downscaled product and era5, respectively, in %. The specific file structure is shown in Table 8.
[0155] Data from January and August 2020 (representing winter and summer respectively) were selected for comparison, and the spatial distribution of daily minimum relative humidity data was compared. The daily minimum relative humidity data observed at daily stations from 1990 to 2022 were used as reference test sources, and error index comparison tests were performed with the original ERA5-Land reanalysis data and the daily minimum relative humidity dataset generated by this embodiment. The correlation coefficient index ERA5-Land is 0.79, and the dataset generated by this embodiment is 0.79; the average deviation difference index ERA5-Land is -0.22%, and the dataset generated by this embodiment is -0.62%; the root mean square error index ERA5-Land is 11.1%, and the dataset generated by this embodiment is 10.69%. The specific image results are as follows: Figures 23a to 23c As shown. The HRCLDAS data of January and early August 2020 were selected as reference materials, and the grid image quality was compared with the original ERA5-Land reanalysis data and the dataset generated in this embodiment. The indicators include peak signal-to-noise ratio, structural similarity index, image similarity index, etc. The specific results are shown as follows. Figures 24a to 24c shown.
[0156] 3.7 Maximum wind speed on the day
[0157] The daily maximum wind speed dataset is stored in the directory " / g11 / wangql / data / wind / eval" on the historical backcalculation server. The spatial resolution is 0.01°×0.01° and the temporal resolution is 1 day. It is stored in NetCDF format and the file name is "MAX_WIND_{yyyymmdd}.nc". The storage structure of the daily maximum wind speed data includes four variables: LON, LAT, Pred, and era5. LON and LAT are coordinate variables in degrees. Pred and era5 are the daily maximum wind speed element variables of the downscaled product and era5, respectively, in °C. The specific file structure is shown in Table 8.
[0158] Data from January and August 2020 (representing winter and summer, respectively) were selected for comparison, and the spatial distribution of the daily maximum wind speed data was compared. The daily maximum wind speed data observed at the stations from 1990 to 2022 were used as the reference test source, and the error index comparison test was performed with the original ERA5-Land reanalysis data and the daily maximum wind speed dataset generated by this embodiment. The correlation coefficient index ERA5-Land is 0.48, and the dataset generated by this embodiment is 0.58; the average deviation difference index ERA5-Land is -1.33m / s, and the dataset generated by this embodiment is -1.07m / s; the root mean square error index ERA5-Land is 2.51m / s, and the dataset generated by this embodiment is 2.21m / s. The specific image results are as follows: Figures 25a to 25c shown.
[0159] The HRCLDAS data from January and early August 2020 were selected as reference data, and the grid image quality was compared with the original ERA5-Land reanalysis data and the dataset generated in this example. The indicators included peak signal-to-noise ratio, structural similarity index, image similarity index, etc. The specific results are as follows: Figures 26a to 26c shown.
[0160] 3.8 Daily average wind speed
[0161] The daily average wind speed dataset is stored in the directory “ / g11 / wangql / data / wind / eval” on the historical backcalculation server. The spatial resolution is 0.01°×0.01° and the temporal resolution is 1 day. It is stored in NetCDF format and the file name is “MEAN_WIND_{yyyymmdd}.nc”. The daily average wind speed data storage structure includes four variables: LON, LAT, Pred, and era5. LON and LAT are coordinate variables in degrees. Pred and era5 are the daily average wind speed element variables of the downscaled product and era5, respectively, in m / s. The specific file structure is shown in Table 8.
[0162] Data from January and August 2020 (representing winter and summer, respectively) were selected for comparison, and the spatial distribution of daily average wind speed data was compared. The daily average wind speed data observed at daily stations from 1990 to 2022 were used as reference test sources, and error index comparison tests were performed with the original ERA5-Land reanalysis data and the daily average wind speed dataset generated by this embodiment. The correlation coefficient index ERA5-Land is 0.53, and the dataset generated by this embodiment is 0.59; the average deviation difference index ERA5-Land is 0.29m / s, and the dataset generated by this embodiment is 0.0m / s; the root mean square error index ERA5-Land is 1.18m / s, and the dataset generated by this embodiment is 1.03m / s. The specific image results are as follows: Figures 27a to 27cshown.
[0163] The HRCLDAS data from January and early August 2020 were selected as reference data, and the grid image quality was compared with the original ERA5-Land reanalysis data and the dataset generated in this example. The indicators included peak signal-to-noise ratio, structural similarity index, image similarity index, etc. The specific results are as follows: Figures 28a to 28c shown.
Claims
1. A spatial downscaling method for obtaining over 30 years of national 1 km daily scale meteorological grid real-time analysis data, characterized in that: The steps include: Step (1) Acquisition of meteorological data: Acquisition of meteorological data from 1990 to 2022, including ground station observation data, HRCLDAS grid real-time data, ERA5-Land data, and terrain parameter data; Step (2), matching of training samples: select the ERA5-Land data and ground station observation data for the whole year of 2021 as training samples; perform daily data matching on the ERA5-Land data and ground station observation data of that year; combine the completed daily data matching samples with the terrain parameter data to generate the training sample dataset required for the downscaling model; Step (3), model training and testing: using the training sample data set to train the constructed downscaling model, forming a machine learning model that meets the requirements of spatial downscaling; During the model training process, data from 2020 were selected to evaluate the downscaling results, in order to compare the error characteristics of the grid products before and after downscaling, thereby verifying the effectiveness of the model training; Step (4), dataset generation and evaluation: A machine learning model that satisfies spatial downscaling is used to generate a national 1 km daily scale meteorological grid real-time analysis data product for more than 30 years, and the generated dataset products are evaluated using ground station observation data and HRCLDAS grid real-time data as the true value.
2. The spatial downscaling method for obtaining 30 years or more of national 1 km daily scale meteorological grid real-time analysis data according to claim 1 is characterized in that: In step (1), the ground station observation data are quality-controlled data from the China Meteorological Administration, including manual observation data from 1990 to 2010 and automatic observation data from 2010 to 2020; HRCLDAS grid live data includes hourly data on 2m temperature, 2m humidity, and 10m wind speed from 2010 to 2022; The ERA5-Land data include hourly data on 2m air temperature, surface temperature, and incident solar shortwave radiation from 1990 to 2022. Topographic parameter data include digital elevation data, slope and aspect data, land cover classification data, and normalized difference vegetation index data. The digital elevation data is a 1-kilometer resolution nationwide DEM data generated by resampling the original 90-meter spatial resolution data jointly released by NASA and the Ministry of Economy, Trade and Industry of Japan. The slope and aspect data are calculated based on digital elevation data; the land cover classification data are from the Global Land Cover Classification Dataset (GLCC) provided by the United States Geological Survey; and the normalized difference vegetation index data are from the 1 km resolution vegetation index product provided by the MODIS sensor of NASA.
3. The spatial downscaling method for obtaining 30 years or more of national 1 km daily scale meteorological grid real-time analysis data according to claim 2, characterized in that: In step (2), the specific method of training sample matching is as follows: extracting the meteorological element data required for downscaling model training from the ground station observation data, and performing statistical calculations of the daily maximum, daily minimum and daily average values of the extracted meteorological elements; extracting the meteorological element data required for downscaling model training from the ERA5-Land data, and performing statistical calculations of the daily maximum, daily minimum and daily average values of the extracted meteorological elements; matching the preprocessed ground station observation data and the ERA5-Land data according to the daily data of each meteorological element; merging the matched daily data and connecting them with the terrain parameter data in space and time to generate the training sample data set required for the downscaling model.
4. The spatial downscaling method for obtaining 30 years or more of national 1 km daily scale meteorological grid real-time analysis data according to claim 3, characterized in that: Meteorological elements include air temperature, relative humidity and wind speed.
5. The spatial downscaling method for obtaining 30 years or more of national 1 km daily scale meteorological grid real-time analysis data according to claim 1, characterized in that: In step (3), before training the downscaling model, the training sample data set is quality controlled to filter out sample data with observation and model errors within ±3 standard deviations.
6. The spatial downscaling method for obtaining 30 years or more of national 1 km daily scale meteorological grid real-time analysis data according to claim 1, characterized in that: In step (3), the machine learning model is the LightGBM model.
7. The spatial downscaling method for obtaining 30 years or more of national 1 km daily scale meteorological grid real-time analysis data according to claim 6, characterized in that: During downscaling model training: sample features include terrain parameter data, longitude and latitude, time, and the daily maximum, daily minimum, and daily average values of each meteorological element. Sample labels include the errors between the daily maximum, daily minimum, and daily average values of each meteorological element model data and the observed data. Error indicators include correlation coefficient, mean deviation, and root mean square error.
8. The spatial downscaling method for obtaining 30 years or more of national 1 km daily scale meteorological grid real-time analysis data according to claim 6, characterized in that: In step (3), the method for optimizing the number of leaves and node depth of the downscaling model is as follows: 5 copies are automatically and randomly selected from the training set as training and 1 copy is used as an evaluation test through the grid search parameter adjustment method; when the node depth remains unchanged, the correction ability of the downscaling model with different numbers of leaves is compared to determine the leaf number parameter; then the correction ability of the downscaling model with different node depths is tested to determine the node depth parameter.
9. The spatial downscaling method for obtaining 30 years or more of national 1 km daily scale meteorological grid real-time analysis data according to claim 1, characterized in that: In step (4), a machine learning model that satisfies spatial downscaling is used as input for the ERA5-Land data from 1990 to 2022 obtained in step (1) to generate a national 1 km daily scale meteorological grid real-time analysis product for more than 30 years; the evaluation indicators include error indicators and grid image quality indicators; the error indicators include correlation coefficient, mean deviation and root mean square error, and the grid image quality indicators include peak signal-to-noise ratio, structural similarity index and learning perception image block similarity index; the ground station observation data is used as the true value to evaluate the error indicators of the generated dataset product, and the HRCLDAS grid real-time data is used as the true value to evaluate the grid image quality indicators of the generated dataset product.
10. A spatial downscaling system that obtains over 30 years of national 1 km daily scale meteorological grid real-time analysis data, characterized by: The spatial downscaling method for obtaining 30 years or more of national 1 km daily scale meteorological grid real-time analysis data as described in any one of claims 1 to 9 comprises the following modules: Data download module, used to download original ERA5-Land data, ground station observation data, and terrain parameter data; The data source module is used to complete the statistical calculation and storage of the daily maximum, minimum and average values of each meteorological element data in the HRCLDAS grid real-time data and ERA5-Land data; The model training module is used to complete the statistical calculation of the daily maximum, minimum, and average values of various meteorological element data in the ground station observation data, and match them with the ERA5-Land data for daily data. The matched daily data are then merged and spatially and temporally connected with the terrain parameter data to generate a training sample dataset based on ground stations across the country. Finally, the matched training sample dataset is used for model training. The historical backcalculation module generates more than 30 years of national 1 km daily scale meteorological grid real-time analysis data products by calling the trained downscaling machine learning model and ERA5-Land data. It evaluates various error indicators using ground station observation data as the true value, and evaluates various image quality indicators using HRCLDAS grid real-time data as the true value.
Citation Information
Patent Citations
Regional wind energy resource refined evaluation method combining regional climate, local daily change and multi-scale interaction characteristics
CN112541654A
Deep learning-based ERA5 precipitation product downscaling method
CN116484189A
Local hundred-meter minute-level humiture rainfall real-time analysis system and construction method thereof
CN119066355A
Ozone concentration forecasting method based on peak perception and influence factor analysis
CN119623754A
The method for managing and displaying heat distribution diagram by received images from satellites
KR102694631B1