Ground relative humidity rasterization method and system and computer equipment
The optimal model was selected by using the signal degrees of freedom and generalized cross-validation minimum likelihood estimation method, and combined with the local thin disk smoothing spline method, the problem of uncertain relative humidity rasterization accuracy was solved, and high-precision meteorological data rasterization was achieved.
Patent Information
- Application Number
- CN202410286802.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-13
- Publication Date
- 2025-10-03
AI Technical Summary
In the existing technology, the optimal parameters for relative humidity rasterization are unclear, resulting in uncertain rasterization accuracy of meteorological data at the regional scale and possible errors.
The signal degrees of freedom are used as the evaluation index of the relative humidity model. The optimal model is selected by combining the generalized cross-validation minimum and maximum likelihood estimation methods. The relative humidity is rasterized using the local thin disk smoothing spline method to ensure an optimal balance between the smoothness of the rasterized surface and the fidelity of the data.
The accuracy of relative humidity rasterization is improved, the smooth continuity of the rasterization results and the authenticity of the data are ensured, the error is reduced, and high-precision rasterization of meteorological data at the regional scale is achieved.
Smart Images

Figure CN120744338A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of ground monitoring technology, and in particular to a ground relative humidity rasterization method, system and computer equipment. Background Art
[0002] Meteorological elements in environmental factors are the data basis for inventions in fields such as resources, environment, and global change. Rasterized meteorological data are important parameters of various geological models, hydrological models, ecological models, and climate models. The observation of meteorological data mainly depends on meteorological stations, but their spatial distribution is uneven. Meteorological data in areas outside the stations can only be estimated from observations at neighboring stations, that is, meteorological data rasterization. Therefore, it is of great significance to improve the rasterization accuracy of meteorological data at the regional scale.
[0003] Currently, common methods for rasterizing meteorological data include the nearest neighbor method, moving average method, thin plate spline method, and kriging. The main disadvantage of the nearest neighbor method is that it does not consider the distribution of surrounding points, resulting in discontinuous and jagged edges in the rasterized results. The moving average method assumes that data changes are linear, which may not be true in reality, and it also smoothes out high-frequency information in the data. The thin plate spline method is not only computationally complex but also oversmoothes the data, resulting in a loss of detail in the rasterized results. Kriging requires prior knowledge of the statistical characteristics of the data and has relatively high computational complexity. Compared with traditional interpolation methods, the thin plate spline method, based on geostatistical interpolation technology, can achieve accurate rasterization without considering prior knowledge or physical processes. Furthermore, this method, while being computationally simple and having a simple data structure, ensures an optimal balance between rasterized surface smoothness and data fidelity, ensuring smooth and continuous rasterized surfaces.
[0004] ANUSPLIN, a specialized software for rasterizing meteorological data, is a multidimensional, multifactorial spatial rasterization method based on thin disk splines. By introducing covariates, this method achieves higher accuracy and smoother, more continuous rasterization results. More importantly, it can rasterize multiple surfaces simultaneously, making it particularly suitable for rasterizing time series meteorological data. Currently, ANUSPLIN has limited application in my country and is primarily used to rasterize meteorological elements such as temperature, precipitation, and visibility. The optimal parameters for rasterizing relative humidity remain unclear, resulting in uncertainty about the accuracy of rasterizing meteorological data at the regional scale. Its applicability has not been verified, and it may be susceptible to errors. Summary of the Invention
[0005] In view of the fact that the optimal parameters in the existing technology are unclear, resulting in uncertain accuracy and possible errors in the rasterization of meteorological data at the regional scale, the present invention proposes a ground relative humidity rasterization method, system and computer equipment. By adopting the signal degree of freedom as the evaluation index of the relative humidity model and taking the minimum of the generalized cross-validation minimum and the maximum likelihood estimate as the selection criteria for the optimal model of all successful relative humidity models, the best relative humidity gridding model is selected, thereby solving the problem that the existing technology has uncertain accuracy and possible errors in the rasterization of meteorological data at the regional scale.
[0006] A method for rasterizing ground relative humidity comprises the following steps:
[0007] Collect meteorological data from regional weather stations and meteorological site monitoring data;
[0008] The meteorological data of the regional weather stations and the site monitoring data were combined in various ways to construct multiple relative humidity models with daily time resolution.
[0009] The signal degrees of freedom and residual signal degrees of freedom are used as evaluation indicators for relative humidity models to determine the successful model among all relative humidity models. The successful model is determined by whether the degrees of freedom and residual signal degrees of freedom are less than half of the number of stations.
[0010] Generalized cross validation and maximum likelihood estimation were used as the selection criteria for the optimal model of all successful relative humidity models, and the best evaluation index was determined by combining the square root error.
[0011] The best evaluation index is used as a basis to calculate the daily relative humidity rasterization ratio of different successful models, and the model with the largest relative humidity rasterization ratio is used as the basis to select the best relative humidity rasterization model;
[0012] The ground monitoring data are input into the relative humidity rasterization optimal model, and the relative humidity rasterization results are output.
[0013] Furthermore, the ground monitoring data includes latitude and longitude, air pressure, temperature, DEM and spline number.
[0014] Furthermore, the local thin disk smoothing spline method is embedded in the relative humidity model, which is expressed as:
[0015] z i =f(x i )+b T y i +e i (i=1,…,N)
[0016] Where i is a point in space; z i is the dependent variable at point i in space; xi is the d-dimensional spline independent variable; f is about x i An unknown smooth function of y i is a p-dimensional independent covariate; b is y i The p-dimensional coefficient; e i has a expected value of 0 and a variance of w i σ 2 The random error of the independent variable, where w i is the known local relative coefficient of variation as a weight, σ 2 is the error variance; N is the number of samples involved in the calculation, and T is the matrix transpose.
[0017] Furthermore, the least square method is used to estimate and determine the function f and the coefficient b, and the calculation formula is:
[0018]
[0019] Where, J m (f) is f(x i ), m is called the spline degree in ANUSPLIN; ρ is the smoothness parameter, which is determined by minimizing the generalized cross validation GCV.
[0020] Furthermore, the signal degree of freedom is expressed as:
[0021] SIGNAL = trace(A)
[0022] Where A is the M×N dimensional influence line matrix, and trace is the trace of the matrix;
[0023] The remaining degrees of freedom are expressed as:
[0024] ERROR=trace(IA)=N-trace(A)
[0025] Where N is the number of sites; A is the M×N dimensional influence line matrix, whose value is greater than half of the number of sites.
[0026] Furthermore, the accuracy of the relative humidity grid model is evaluated by using a signal-to-noise ratio, and the calculation formula is expressed as:
[0027] SNR=SIGNAL / ERROR
[0028] Where SIGNAL is the signal degree of freedom; ERROR is the residual degree of freedom.
[0029] Furthermore, the generalized cross validation root mean square error is expressed as:
[0030] RTGCV=√GCV
[0031] The root mean square error is expressed as:
[0032]
[0033] Where W is a diagonal matrix.
[0034] Furthermore, the determination coefficient is used to classify the goodness of fit of the relative humidity rasterization model, and the classification results include excellent, good, fair, and poor. The calculation formula of the determination coefficient is:
[0035]
[0036] Among them, i is the monitoring site data; y i The i-th actual monitoring data; f i The relative humidity corresponding to the i-th station is rasterized; is the mean of actual monitoring data.
[0037] Furthermore, a ground relative humidity gridding system includes:
[0038] The acquisition module is used to collect meteorological data from regional weather stations and meteorological site monitoring data;
[0039] The model building module is used to combine the meteorological data of the regional weather stations and the site monitoring data in various ways to build relative humidity models with multiple time resolutions;
[0040] a successful model determination module, configured to use the signal degrees of freedom or the residual signal degrees of freedom as an evaluation index for the relative humidity model to determine a successful model among all relative humidity models; wherein the successful model is determined based on whether the degrees of freedom or the residual signal degrees of freedom are less than half of the number of stations;
[0041] An evaluation index determination module is used to use the generalized cross validation method and the maximum likelihood estimation method as the selection criteria for the optimal model of all successful relative humidity models, combined with the square root error, to determine the best evaluation index;
[0042] A selection module is used to calculate the daily relative humidity rasterization ratios of different successful models based on the best evaluation index, and select the best relative humidity rasterization model based on the largest relative humidity rasterization ratio;
[0043] The calculation module is used to input ground monitoring data into the relative humidity rasterization optimal model and output the relative humidity rasterization results.
[0044] Furthermore, a ground relative humidity rasterization computer device includes: a memory, a processor, and a computer program stored in the memory, and when the processor executes the computer program, the steps of the ground relative humidity rasterization method according to any one of claims 1 to 8 are implemented.
[0045] The present invention provides a method, system, and computer device for rasterizing ground relative humidity, which have the following beneficial effects:
[0046] The present invention adopts the signal degree of freedom as the evaluation index of the relative humidity model, and determines the successful model according to whether the degree of freedom or the residual signal degree of freedom is less than half of the number of sites. The generalized cross-validation method and the maximum likelihood estimation method are used as the selection criteria for the optimal model of all successful relative humidity models, and the best evaluation index is determined, thereby selecting the best relative humidity gridding model, and outputting the relative humidity rasterization result under the model, ensuring the optimal balance between the smoothness of the rasterized surface and the fidelity of the data, ensuring the smoothness and continuity of the rasterized surface, and making the rasterization result more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 The daily relative humidity rasterization results and monitoring data R from 2015 to 2019 in the embodiment of the present invention are 2 Statistical charts;
[0048] Figure 2 The relative humidity gridding results and monitoring values R for each site from 2015 to 2019 in the embodiment of the present invention are shown in Table 1. 2 Statistical charts;
[0049] Figure 3 The relative humidity rasterization results and monitoring values R of various cities from 2015 to 2019 in the embodiment of the present invention are shown in Table 1. 2 Statistical charts;
[0050] Figure 4 The monthly relative humidity rasterization results and monitoring data R from 2015 to 2019 in the embodiment of the present invention are 2 Statistical charts;
[0051] Figure 5 The rasterized results and monitoring data of daily relative humidity in 2020 in the embodiment of the present invention are R 2 Statistical charts;
[0052] Figure 6 The relative humidity grid results and monitoring data R for each site in 2020 in the embodiment of the present invention are as follows: 2 Statistical charts;
[0053] Figure 7 The relative humidity rasterization and monitoring data R of various cities in 2020 in the embodiment of the present invention2 Statistical charts;
[0054] Figure 8 The monthly relative humidity rasterization results and monitoring data R in 2020 in the embodiment of the present invention are 2 Statistical charts;
[0055] Figure 9 This is a flowchart for completing the optimal model selection in an embodiment of the present invention using the ground monitoring data of the China Meteorological Data Network from 2015 to 2020 as experimental data. DETAILED DESCRIPTION
[0056] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0057] The present invention proposes a method for rasterizing ground relative humidity, which specifically includes the following steps:
[0058] Collect meteorological data from regional weather stations and site monitoring data.
[0059] The meteorological data of the meteorological stations in the area and the site monitoring data were combined in various ways to construct multiple relative humidity models with a daily time resolution.
[0060] The signal degrees of freedom or residual signal degrees of freedom are used as evaluation indicators of the relative humidity model to determine the successful models of all relative humidity models; the successful models are determined based on whether the degrees of freedom or residual signal degrees of freedom are less than half of the number of stations.
[0061] The minimum of generalized cross validation and the minimum of maximum likelihood estimation were used as the selection criteria for the optimal model of all successful relative humidity models, and the best evaluation index was determined by combining the square root error.
[0062] The best evaluation index is used as the basis to count the daily relative humidity gridding proportions of different successful models, and the model with the largest relative humidity gridding proportion is used as the basis to select the best relative humidity gridding model.
[0063] Input the ground monitoring data into the relative humidity gridding optimal model and output the relative humidity gridding result.
[0064] To improve the accuracy of relative humidity rasterization in Shaanxi Province, this study used ground monitoring data from the China Meteorological Data Network from 2015 to 2020, combined with latitude and longitude, air pressure, temperature, and elevation, using the ANUSPLIN software rasterization method to explore the effects of different variable combinations and spline orders on the accuracy of the relative humidity model. Using ground monitoring data from the China Meteorological Data Network from 2015 to 2019 as experimental data, the optimal model was selected using the following steps:
[0065] In the first step, the ground monitoring data from 2015 to 2019 were preprocessed into standard data, and the preprocessed standard data were used as model input to construct a relative humidity model from 2015 to 2019 with a time resolution of daily.
[0066] In the second step, signal degrees of freedom or residual signal degrees of freedom were used as the model success evaluation indicators, and all successful models from 2015 to 2019 were selected.
[0067] In the third step, the minimum of generalized cross validation and the minimum of maximum likelihood estimation were used as the optimal model selection criteria respectively. Combined with the square root error, statistical analysis was performed to determine which of generalized cross validation and maximum likelihood estimation were the best evaluation indicators for optimal model selection.
[0068] The fourth step is to select the best model every day based on the best evaluation index, calculate the success probability of different successful models, and select the best relative humidity rasterization model based on the maximum percentage.
[0069] The fifth step is to complete the daily rasterization of relative humidity from 2015 to 2019 based on the optimal model, calculate the determination coefficient of the relative humidity rasterization results and the actual monitoring values in time and space under the model, and conduct statistical analysis to ensure the correctness of the model and clarify the covariates (geographical and meteorological factors) and model parameters required for the optimal relative humidity rasterization.
[0070] The 2020 ground monitoring data from the China Meteorological Data Network was used as validation data to verify that the optimal model is universal in time and space:
[0071] The first step is to preprocess the 2020 ground monitoring data of the China Meteorological Data Network into standard data and complete the construction of the 2020 relative humidity model with a daily time resolution based on the optimal model.
[0072] The second step is to complete the rasterization of the relative humidity in 2020 based on the constructed model.
[0073] The third step is to select the coefficient of determination between the monitoring data and the spatial rasterization results as the model validation evaluation index, and to verify that the statistical accuracy in both time and space reaches the required percentage, so as to verify that the selected optimal relative humidity rasterization model is universal in time and space.
[0074] Based on the same inventive concept, the present invention provides a ground relative humidity gridding system, comprising:
[0075] The acquisition module is used to collect meteorological data from regional weather stations and site monitoring data.
[0076] The model building module is used to combine the meteorological data of the meteorological stations in the area and the site monitoring data in various ways to build multiple relative humidity models with a time resolution of daily.
[0077] The successful model determination module is used to use the signal degrees of freedom or the residual signal degrees of freedom as the evaluation index of the relative humidity model to determine the successful model among all relative humidity models; wherein the successful model is determined based on whether the degrees of freedom or the residual signal degrees of freedom are less than half of the number of stations.
[0078] The evaluation index determination module is used to use the minimum of generalized cross validation and the minimum of maximum likelihood estimation as the selection criteria for the optimal model of all successful relative humidity models, and combine the square root error to determine the best evaluation index.
[0079] The selection module is used to use the best evaluation index as a basis to count the daily relative humidity rasterization proportions of different successful models, and use the model with the largest relative humidity rasterization proportion as a basis to select the best relative humidity rasterization model.
[0080] The calculation module is used to input ground monitoring data into the relative humidity rasterization optimal model and output the relative humidity rasterization results.
[0081] All experiments in this paper are implemented using the Python 3.8 development language with the help of the Pycharm 2022.1 compiler. The required library packages are mainly pandas (1.3.3), numpy (1.21.2), geopandas (0.9.0), rasterio (1.2.8), sklearn (0.24.2), rasterstats (0.17.0), matplotlib (3.4.3), seaborn (0.11.2), affine (2.3.0) and pyproj (3.2.1). The specific operation process is as follows:
[0082] First, the 2015-2020 meteorological data and elevation data (DEM) were preprocessed to prepare the data into the standard format for ANUSPLIN software input. The 2015-2020 meteorological data were used to prepare the experimental data and verification data. The specific preprocessing process and content include the following:
[0083] (1) Meteorological data processing: The cleaning process of meteorological station monitoring data is a key step to ensure data quality and availability, so as to ensure the accuracy of subsequent analysis and modeling. The following is the cleaning process of meteorological station monitoring data: data import, data quality check, missing value processing, outlier processing, data format consistency, deduplication processing, data formatting, and data verification.
[0084] (2) Elevation data (DEM) processing: The DEM data comes from the geospatial data cloud, with a spatial resolution of 30m*30m, and is a tiled data. The tiled data covering Shaanxi Province was downloaded, and the DEM tiled data was spliced using ARCGIS software. The spliced results were clipped using the Shaanxi Province vector mask to obtain the DEM data for the study area. Whether the DEM data needs to be resampled depends on the research needs. In this study, the spatial resolution of the average relative humidity rasterization result is 1000m, so the Shaanxi Province DEM needs to be resampled and stored in raster form. The DEM data required by ANUSPLIN is required to be written in ASCII format. The DEM data originally stored in raster form needs to be converted into ASCII format, and the correctness of the data during the conversion process must be checked and guaranteed.
[0085] ANUSPLIN software has a total of 6 modules, namely SPLINE, SELNOT, ADDNOT, LAPGRD, LAPPNT, and GCVGML. Their specific meanings are: SPLINE is applicable to thin disk spline functions of any independent variables or multiple covariates, and the data smoothness is determined by GCV or GML; SELNOT selects the initial node for SPLINE; AODNOT adds nodes; GCVGML calculates GCV or GML errors for the fitted surface for data inspection or positioning; LAPPNT calculates the point file of predicted values or Bayesian standard error estimates; LAPGRD generates a fitted surface or a Bayesian standard error surface.
[0086] The general usage process of ANUSPLIN software is as follows: (1) Execute the SPLINE command to generate six files: log and list, error covariance, surface coefficient, optimization parameters and residuals.
[0087] (1) Execute the SPLINE command to generate six files including log and list, error covariance, surface coefficient, optimization parameters and residuals.
[0088] (2) Execute the LAPGRD command to obtain the fitted surface and standard error surface from the surface coefficient file and the error covariance file. Whether to input the error covariance file or output the standard error surface depends on whether the SPLINE command outputs the error covariance file.
[0089] The SPLINE and LAPGRD modules can be implemented using either software applications or CMD commands. Using software applications for preprocessing data and setting parameters is more tedious. By comparison, implementing data preprocessing with a development language and constructing CMD commands for surface fitting with Anusplin software can reduce this tedious work and facilitate subsequent streamlined processing. The CMD command requirements for the SPLINE module are shown in Table 1, and those for the LAPGRD module are shown in Table 2.
[0090] Secondly, according to the rasterization requirements of the ANUSPLIN software, 45 models were constructed (SPLINE module) and run for the daily relative humidity meteorological elements from 2015 to 2019, considering different variable combinations of latitude and longitude, air pressure, temperature, DEM, and spline order (2, 3, and 4). The specific construction requirements are shown in Table 1.
[0091] Table 1 SPLINE command construction requirements
[0092]
[0093]
[0094]
[0095]
[0096]
[0097] This module will generate a log file for each model on a daily basis, which records the statistical parameters for judging the source of errors and evaluating the quality of rasterization. According to whether the signal degree of freedom (SIGNAL) of the rasterization quality indicator is less than half of the number of sites, all successful models are counted and a log summary table is formed (the GCV and GML verification comparison of the average relative humidity ANUSPLIN output results from 2015 to 2019 is shown in Table 2). The successful model summary table is statistically analyzed to determine which of the generalized cross-validation index (GCV) and maximum likelihood (GML) is used as the best model evaluation method of the present invention. Based on the selected evaluation method, the proportion of different successful models in the daily relative humidity rasterization from 2015 to 2019 is counted, and the model combination with the largest proportion is selected as the best model of the present invention, specifically:
[0098] ANUSPLIN provides a series of statistical parameters in the log file (Log file and List file) for identifying error sources and rasterization quality. These include observation data statistics (mean, variance, standard deviation, etc.), effective number of fitted surface parameters estimated by SIGNAL (signal degrees of freedom), residual degrees of freedom ERROE, smoothing parameters RHO, GCV, expected true mean square error (MSE), maximum likelihood error (GML), mean square residual (MSR), variance estimate VAR and its square root. These can be used to select the optimal model. The statistical results also provide the sequence of data points with the maximum root mean square residual (RMSRR), which can be used for data quality control to detect and eliminate errors in the position and value of the original data. The log file can also record the execution process of the program to identify problems during execution.
[0099] SIGNAL indicates the complexity of the fitted surface, while RHO balances the accuracy and smoothness of the fitted surface. A small RHO, a SIGNAL greater than half the number of observation sites, or an excessively large RHO indicates that the fitting process cannot find the optimal smoothing parameters. This suggests that the data points may be too sparse, there are short correlations in the data, or the fitting function is too complex, making the selected model unsuitable for gridding. These situations are marked with an * in ANUSPLIN. When fitting a daily surface, the SIGNAL value should have a smooth transition between days. A clear deviation from the transition trend indicates that the gridded surface for that day may have systematic errors. It can be used for preliminary data verification and generally indicates missing data.
[0100] The best model judgment criteria: GCV or GML is the smallest, the signal-to-noise ratio (SNR) (ratio of signal degrees of freedom to remaining degrees of freedom) is the smallest, the signal degrees of freedom are less than half of the site, and there is no * indicator in the model success rate judgment.
[0101] The optimal model was used to construct the relative humidity rasterization command (LAPGRD module) for the period 2015–2019. Details of the rasterization are shown in Table 2. The LAPGRD command construction and execution output of this module is the relative humidity rasterization result. The relative humidity monitoring values from 2015–2019 were compared with the relative humidity rasterization results. A statistical analysis was performed to determine whether the optimal model met the required temporal and spatial accuracy to ensure its accuracy.
[0102] Table 2 LAPGRD command construction and execution
[0103] Site longitude latitude elevation Air pressure temperature Average relative humidity 53543 109.98 39.83 1461.9 860.6 -8.5 26 53564 111.22 39.37 1036 909.8 -10.7 36 53646 109.78 38.27 1157 895.8 -10.3 31 53651 110.47 38.82 1098 903.5 -11.3 34 53663 111.82 38.92 1401 868.3 -12.5 30 53664 111.13 38.47 1012.6 913.3 -9.6 34 53723 107.38 37.8 1349.3 874 -9.6 42 53725 107.58 37.58 1360.3 872.1 -8.1 37 53735 108.8 37.62 1336.7 876.2 -6.8 31 53738 108.17 36.92 1331.4 877.6 -9.7 50
[0104] Next, the optimal model was used to construct a daily relative humidity model and rasterization command for 2020, obtaining spatial rasterization results of the relative humidity in Shaanxi Province in 2020. The 2020 relative humidity monitoring data and the ANUSPLIN rasterization results were compared, and the rasterization performance of the 2020 optimal model was statistically analyzed in time and space, verifying the universality of the optimal model in time and space. Finally, based on the above process, the optimal model for the relative humidity in Shaanxi Province using the ANUSPLIN software rasterization method was determined. This optimal model was used to generate a set of mature tools to facilitate the subsequent rasterization of relative humidity meteorological elements in Shaanxi Province.
[0105] The meteorological data used in this invention comes from the China Meteorological Data Network (https: / / data.cma.cn / data / detail / dataCode / A.0012.0001.html), mainly including three meteorological elements of daily relative humidity, air pressure and temperature for 6 years (2016-2020, 2022) from 60 meteorological stations in Shaanxi Province (32) and its surrounding areas (28). The station monitoring data provided includes longitude and latitude data.
[0106] The elevation data used in the present invention is the 90m resolution digital elevation data DEM of the geospatial data cloud SRTMDEMUTM.
[0107] The basic algorithm involved in the model in Anusplin software: Local Thin Disk Smoothing Spline Method is an extension of Thin Disk Smoothing Spline. It allows the introduction of not only spline independent variables but also linear covariate submodels, such as the relationship between temperature and altitude, precipitation and coastline, etc. By introducing covariates, the rasterization results are more accurate and the surface is more continuous and smooth.
[0108] z i =f(x i )+b T y i +e i (i=1,…,N) (1)
[0109] Where i is a point in space; z i is the dependent variable at point i in space; x i is the d-dimensional spline independent variable; f is the variable about x that needs to be estimated i An unknown smooth function of y i is a p-dimensional independent covariate; b is y i The p-dimensional coefficient; e i has a expected value of 0 and a variance of w i σ 2 The random error of the independent variable, where w i is the known local relative coefficient of variation as a weight, σ 2is the error variance, which is constant across all data points but is usually unknown; N is the number of samples involved in the calculation.
[0110] When the second term is missing, that is, the model has no covariates, the model becomes a common thin disk smoothing spline model. When the first independent variable is missing, the model is a multiple linear regression model, which is not allowed in the ANUSPLIN software.
[0111] The function f and the coefficient b can be determined by minimizing the following equation, i.e., by least squares estimation:
[0112]
[0113] Where, J m (f) is f(x i ), m is called the spline degree in ANUSPLIN, and the roughness degree in the package; ρ is a smoothness parameter, which is usually greater than 0 and plays a balancing role between data fidelity and surface roughness. It is usually determined by minimizing generalized cross validation (GCV).
[0114] Evaluation indicators:
[0115] (1) Signal freedom (SIGNAL): The signal freedom directly reflects whether the number of sample points or nodes is sufficient. If its value is too large or too small, the program will not be able to obtain the optimal value of the smoothing parameter. It is marked with * in the log file and generally cannot exceed half of the number of data points.
[0116] SIGNAL=trace(A) (3)
[0117] Where A is the M×N dimensional influence line matrix.
[0118] (2) Residual degrees of freedom (ERROR):
[0119] ERROR=trace(IA)=N-trace(A) (4)
[0120] Where N is the number of sites, and A is the M×N dimensional influence line matrix. Its value must be greater than half the number of sites.
[0121] The sum of the residual degrees of freedom and the signal degrees of freedom is N, which is the number of sites.
[0122] (3) Signal-to-noise ratio (SNR): The signal-to-noise ratio is the ratio of the signal degrees of freedom to the residual degrees of freedom. The smaller the value, the better the model accuracy.
[0123] SNR=SIGNAL / ERROR (5)
[0124] Where SIGNAL is the signal degree of freedom; ERROR is the residual degree of freedom. The SIGNAL and ERROR indicators are recorded in the log file generated by the model building run.
[0125] (4) Generalized Cross-Validation (GCV): GCV is a cross-validation-based evaluation metric that evaluates the model's fit by fitting the model on a training dataset and calculating the sum of squared prediction errors on a validation dataset. The goal of GCV is to select a smoothing parameter that minimizes the sum of squared prediction errors.
[0126] The basic principle of generalized cross-validation is to randomly select some observation points, use the remaining points to grid the randomly selected points, and compare the errors between the actual observation values and the predicted values of the selected points.
[0127] W=diag(w1,...,w N ) (6)
[0128]
[0129] Where N is the number of sites; A is the M×N dimensional influence matrix; w i (i=1,…,N) is the relative error variance. The smaller the generalized cross validation GCV value, the better the model accuracy.
[0130] (5) Root mean square error of generalized cross validation (RTGCV):
[0131] RTGCV=√GCV (8)
[0132] The root mean square error is used to measure the deviation between the predicted value and the true value. The smaller the root mean square error of generalized cross validation, the better the model accuracy.
[0133] (6) Generalized Maximum Likelihood (GML): GML is an evaluation metric based on maximum likelihood estimation. It selects the optimal smoothing parameter by maximizing the likelihood function of the model. The goal of GML is to select a smoothing parameter that maximizes the probability that the model fits the data.
[0134] (7) Root mean square residual (RTMSR): Root mean square residual is a statistical indicator used to evaluate the goodness of fit of a model. It measures the difference between the model predicted value and the observed value. The smaller the value of the root mean square residual, the better the goodness of fit of the model and the smaller the difference between the predicted value and the observed value.
[0135]
[0136] Where N is the number of sites; A is the M×N dimensional influence matrix; and W is a diagonal matrix.
[0137] (8) Root Mean Square Error (RTMSE): Root Mean Square Error is a commonly used metric to evaluate the accuracy of a prediction model or rasterization method. It can be used to assess the accuracy of the model and the reliability of the prediction. A smaller RMSE value indicates that the model has better prediction ability and the difference from the actual observation is smaller; while a larger RMSE value indicates that the model has poor prediction ability and the difference from the actual observation is larger.
[0138]
[0139] Detailed explanation of the parameters in the formula is given in (7).
[0140] (9) Determination coefficient: The present invention uses the determination coefficient to evaluate the accuracy of the rasterization results of the anusplin software and the actual site monitoring data. The calculation formula is as follows:
[0141]
[0142] Where i is the monitoring site data; y i The i-th actual monitoring data; f i is the ANUSPLIN gridded result corresponding to the i-th site; is the mean of the actual monitoring data. The closer the coefficient of determination is to 1, the better the model accuracy.
[0143] According to the range of the coefficient of determination, the goodness of fit of the model can be classified, and the results are shown in Table 3.
[0144] Table 3 Classification diagram of coefficient of determination
[0145]
[0146] Statistical Methods: The statistical methods used in this paper are frequency and repetition. Frequency refers to the number of times a specific value appears in a dataset, measuring the repetition of values within the data. Frequency refers to the relative number of occurrences of a specific value, typically expressed as a percentage or decimal. It is the ratio of frequency to the total number of data points and indicates the relative importance of a value within the dataset. Calculating frequency and repetition can effectively understand the distribution and trends of a dataset and is very useful for data aggregation, visualization, and further analysis.
[0147] Result analysis:
[0148] Model statistical analysis: For relative humidity, the 45 initially determined models were used to construct a model for the average daily data of the invention area from 2015 to 2019 (a total of 1,825 days) and to collect statistical error information. The statistical results were analyzed to determine the final model evaluation index of the present invention. Based on the determined model evaluation index, the frequency of occurrence of all successful models was counted, and the optimal model applicable to relative humidity in the invention area was determined according to the principle of maximum frequency.
[0149] (1) Determining the evaluation metric: The GCV and GML metrics are commonly used to select appropriate smoothing parameters for model construction. However, determining which metric is better depends on the specific problem and dataset, and there is no fixed answer. Therefore, the selection of evaluation metrics requires analysis and judgment based on the specific situation.
[0150] In order to determine the appropriate evaluation index, the present invention adopts a two-step method. First, according to the principle of minimum GCV and minimum GML, the models that perform well under the GCV and GML evaluation indicators are screened out day by day, and the results are summarized (see Table 4). Secondly, based on the above selection, the final evaluation index is determined by the frequency and frequency of the evaluation criteria with the smallest square root error (RT VAR, RTMSR and RTMSE). If the frequency of GML occurrence is greater than GCV, GML is used as the evaluation criterion for counting and accumulation; if the frequency of GML occurrence is less than GCV, GCV is used as the evaluation criterion for counting and accumulation; if the frequencies of the two are equal, any one can be selected as the evaluation criterion for counting and accumulation. Finally, the best evaluation index is determined by statistically analyzing the frequency of occurrence of the three.
[0151] As shown in Table 4, 45 preliminary models were used to construct a model for the daily relative humidity in the invention area from 2015 to 2019 (1825 days). The successfully constructed models (hereinafter referred to as successful models) were statistically analyzed using GCV and GML as evaluation indicators. The results showed that the number of days with good performance capabilities selected based on the minimum GCV principle was 542 days, accounting for 29.7% of the invention time; the number of days with good performance capabilities selected based on the minimum GML principle was 785 days, accounting for 43.01% of the invention time; there were 498 days (27.29%) that could be evaluated using both GCV and GML as evaluation indicators. This result indicates that GML may be a better evaluation indicator.
[0152] Table 4 Frequency statistics of models that perform well under GCV and GML evaluation indicators
[0153] GCV frequency GML frequency GCV=GML frequency total Number of days 542 785 498 1825 Proportion (%) 29.70 43.01 27.29 100
[0154] Table 5 shows that the statistical results based on the minimum selection criteria of the root mean square error (RTVAR, RTMSR, and RTMSE) show that the secondary selection model was successful based on GCV and GML as evaluation indicators. GCV was used as the evaluation indicator on 1074 days, accounting for 58.85%; GML was used as the evaluation indicator on 243 days, accounting for 13.32%; and both GCV and GML could be used as evaluation indicators on 508 days, accounting for 27.83%. This result shows that GCV is more effective as an evaluation indicator.
[0155] Table 5: Frequency of GCV and GML as evaluation criteria based on square root error statistics
[0156]
[0157] Combining Tables 4 and 5, we can see that when only GCV and GML are used as model selection evaluation indicators, the proportion of models with good performance using GML as the evaluation indicator is only slightly higher than the proportion of models with good performance using GCV as the evaluation indicator. However, when considering square root error as the model selection evaluation indicator, the proportion of models with good performance using GCV as the evaluation indicator is significantly higher than the proportion of models with good performance using GML as the evaluation indicator. Therefore, in this invention, GCV is selected as the model evaluation indicator.
[0158] (2) Determination of the optimal model: The GCV minimum principle was used to select the model for the relative humidity on a daily basis from 2015 to 2019. The models were statistically analyzed according to the variables, covariates, and abbreviations of the models. The results are shown in Table 6.
[0159] Table 6 Daily frequency statistics of the best model using GCV as the evaluation indicator from 2018 to 2019
[0160]
[0161]
[0162] As shown in Table 6, the frequency of models with latitude and longitude as variables, air pressure and temperature as covariates, and a power of 2 is 376, accounting for 20.62%; the frequency of models with latitude and longitude as variables, air pressure and temperature as covariates, and a power of 3 is 311, accounting for 17.04%; the frequency of models with latitude and longitude as variables, elevation, air pressure and temperature as covariates, and a power of 2 is 246, accounting for 13.48%; the frequency of models with latitude and longitude as variables, elevation, air pressure and temperature as covariates, and a power of 3 is 206, accounting for 11.29%; the proportion of the remaining selected models is less than 10%. Because the model with longitude and latitude as variables, air pressure and temperature as covariates, and a power of 2 accounts for the highest proportion, the surface file constructed by this model is selected for rasterization in the present invention to achieve daily rasterization of relative humidity in the invention area from 2015 to 2019, and statistical analysis of the rasterization results and site monitoring data in space and time to ensure the accuracy of the model in the daily rasterization of relative humidity from 2015 to 2019.
[0163] Statistical analysis of the accuracy of the results:
[0164] A model with latitude and longitude as variables, pressure and temperature as covariates, and a power of 2 was constructed. The output surface was used as a gridded input to fit the relative humidity surface for the invention area from 2015 to 2019. Daily gridded results of the relative humidity data for the invention area from 2015 to 2019 were obtained. The accuracy of the gridded model for the relative humidity data for the invention area from 2015 to 2019 was evaluated using daily and site-by-site calculations to determine the optimal gridded model.
[0165] The daily calculation is to calculate the relative humidity of the station in the invention area and the rasterized results of the corresponding date. 2 , summarizing the R 2 Values, to understand the time trend of the optimal model. 2 It can be seen that Figure 1 ):R 2 The number of days below 0.60 is 89 days, which is only 4% of the invention time; 2 The number of days above 0.8 is 1353 days, accounting for 74% of the invention time, indicating that the optimal model rasterized relative humidity field is correct in time series.
[0166] The calculation of each station is to calculate the relative humidity of each relative humidity station within the invention area and the R of the grid result. 2 , understand the spatial variation and performance of the model. 2 It can be seen that Figure 2 ): There are only two sites R in Hanzhong City 2Less than 0.83, the two sites are 57211 and 57127, but their R 2 The rasterization results are 0.73 and 0.83 respectively, which shows good performance. The other sites R 2 It is above 0.83, meeting the accuracy requirements and the rasterization results are good or excellent, indicating that the rasterized relative humidity of the optimal model is spatially correct.
[0167] (1) Statistical analysis by city: In the spatial dimension, the performance of the optimal model rasterized relative humidity in different cities is statistically analyzed to determine whether the optimal model rasterized relative humidity is correct in different regions.
[0168] Statistics of R in each city within the invention zone 2 ( Figure 3 ): Tongchuan, Yulin, Xixian New Area, Xianyang, Shangluo, Xi'an, Yan'an and Baoji 2 Above 0.90, it shows that the rasterization results in these cities have excellent surface capabilities; Weinan, Hanzhong, and Ankang have R 2 The rasterization results are between 0.80 and 0.90, and the overall R 2 The rasterization result is 0.94, which shows good performance. By comparing the data from different cities in the invention area, we can determine the accuracy of the optimal model rasterization relative humidity in different regions.
[0169] (2) Statistical analysis by time: In the time dimension, the performance of the optimal model rasterized relative humidity in different months is statistically analyzed by month to determine whether the optimal model rasterized relative humidity is correct in different time periods.
[0170] Statistics of R in the invention area every month 2 : Optimal model rasterization result R 2 The lowest value was in August, but it also reached 0.88, and the rasterization results showed good performance. 2 All are above 0.90, indicating excellent performance of the rasterized results. By comparing data from different months, we can confirm that the optimal model rasterized relative humidity is accurate in different months.
[0171] Verification results:
[0172] A relative humidity model for 2020 was constructed using the optimal model parameters (latitude and longitude as variables, pressure and temperature as covariates, and a power of 2). This surface was used as a rasterized input to fit the 2020 daily average relative humidity surface for the invention area, obtaining a rasterized result of the 2020 daily relative humidity for the invention area. The accuracy of the rasterized model for relative humidity data for the invention area during 2020 was evaluated using daily and site-by-site calculations. The universality of the optimal model was verified by statistically analyzing its accuracy across time and space.
[0173] The daily calculation is to calculate the relative humidity of the station in the invention area and the rasterized results of the corresponding date. 2 , summarizing R in the 2020 time range 2 Values, by time series statistical analysis R 2 It can be seen that Figure 5 ):R 2 The number of days below 0.60 is 25 days, which is only 7% of the invention time; 2 The number of days above 0.8 is 341 days, accounting for 93% of the invention time, indicating that the optimal model rasterized relative humidity field is universal in time series.
[0174] The calculation of each station is to calculate the relative humidity of each relative humidity station within the invention area and the R of the grid result. 2 , from the spatial statistical analysis R 2 It can be seen that Figure 6 ):Only the station numbered 57211 in Hanzhong City and the station numbered 57245 in Ankang City R 2 Less than or equal to 0.83, respectively 0.83 and 0.80, the rasterization results are good; other sites R 2 Above 0.83, the accuracy requirements are met and the rasterization results are good or excellent, which verifies that the optimal model relative humidity rasterization is spatially universal.
[0175] (1) Statistical analysis by prefecture-level city: Statistics of R in each prefecture-level city within the invention zone 2 , Tongchuan City, Yulin City, Xixian New Area, Xianyang City, Shangluo City, Xi'an City, Yan'an City and Baoji City R 2 Above 0.90, it shows that the rasterization results in these cities have excellent surface capabilities; Weinan, Hanzhong, and Ankang have R 2 The rasterization results are between 0.80 and 0.90, and the overall R 2 The rasterization performance is 0.94, which is good. By comparing the data from different cities in the invention area, we can verify the universality of the optimal model rasterization relative humidity in different regions.
[0176] (2) Statistical analysis by time: Statistics of R in the invention area every month2 : Optimal model rasterization result R 2 In January, August-September and February, a total of four months of R 2 The rasterization results are between 0.80 and 0.90, and the performance is good. The R 2 Above 0.90, the gridded results are excellent. By comparing data from different months, we can verify that the optimal model for gridded relative humidity is universal across different months.
[0177] Using 45 pre-selected models, we constructed a daily model of the relative humidity in the invention area from 2015 to 2019. We statistically compared the error information to determine the optimal evaluation index (GCV) for the invention. Based on this optimal evaluation index, we calculated the optimal model for each day and conducted a frequency analysis of this result. We determined that the optimal model for relative humidity rasterization is: longitude and latitude as variables, air pressure and temperature as covariates, and a power of 2. We then generated a grid based on the optimal model, and statistically analyzed the monitoring values and the accuracy of the rasterized results to demonstrate the correctness of the model selection. The daily rasterized results and the accuracy of the monitoring values for the 2020 relative humidity data verified the universality of the optimal model.
[0178] Based on the same inventive concept, the present invention also provides a ground relative humidity rasterization computer device, comprising: a memory, a processor, and a computer program stored in the memory, wherein the processor implements the steps of the ground relative humidity rasterization method when executing the computer program.
[0179] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A method for rasterizing ground relative humidity, characterized in that: The following steps are involved: Collect meteorological data from regional weather stations and meteorological site monitoring data; The meteorological data of the regional weather stations and the site monitoring data were combined in various ways to construct multiple relative humidity models with daily time resolution. The signal degrees of freedom and residual signal degrees of freedom are used as evaluation indicators for relative humidity models to determine the successful model among all relative humidity models. The successful model is determined by whether the degrees of freedom and residual signal degrees of freedom are less than half of the number of stations. Generalized cross validation and maximum likelihood estimation were used as the selection criteria for the optimal model of all successful relative humidity models, and the best evaluation index was determined by combining the square root error. The best evaluation index is used as a basis to calculate the daily relative humidity rasterization ratio of different successful models, and the model with the largest relative humidity rasterization ratio is used as the basis to select the best relative humidity rasterization model; The ground monitoring data are input into the relative humidity rasterization optimal model, and the relative humidity rasterization results are output.
2. A method for rasterizing ground relative humidity according to claim 1, characterized in that: The ground monitoring data includes latitude and longitude, air pressure, temperature, DEM and spline times.
3. A method for rasterizing ground relative humidity according to claim 1, characterized in that: The relative humidity model is embedded in the local thin disk smoothing spline method, which is expressed as: z i =f(x i )+b T y i +e i (i=1,…,N) Where i is a point in space; z i is the dependent variable at point i in space; x i is the d-dimensional spline independent variable; f is about x i An unknown smooth function of y i is a p-dimensional independent covariate; b is y i The p-dimensional coefficient; e i has a expected value of 0 and a variance of w i σ 2 The random error of the independent variable, where w i is the known local relative coefficient of variation as a weight, σ 2 is the error variance; N is the number of samples involved in the calculation, and T is the matrix transpose.
4. A method for rasterizing ground relative humidity according to claim 3, characterized in that: It also includes the use of the least squares method to estimate the function f and the coefficient b, and the calculation formula is: Where, J m (f) is f(x i ), m is called the spline degree in ANUSPLIN; ρ is the smoothness parameter, which is determined by minimizing the generalized cross validation GCV.
5. The method for rasterizing ground relative humidity according to claim 1, wherein: The signal degrees of freedom are expressed as: SIGNAL=trace(A) Where A is the M×N dimensional influence line matrix, and trace is the trace of the matrix; The remaining degrees of freedom are expressed as: ERROR=trace(IA)=N-trace(A) Where N is the number of sites; A is the M×N dimensional influence line matrix, whose value is greater than half of the number of sites.
6. A method for rasterizing ground relative humidity according to claim 5, characterized in that: The accuracy of the relative humidity gridding model is evaluated using the signal-to-noise ratio, which is expressed as follows: SNR=SIGNAL / ERROR Where SIGNAL is the signal degree of freedom; ERROR is the residual degree of freedom.
7. The method for rasterizing ground relative humidity according to claim 1, characterized in that: The generalized cross validation root mean square error is expressed as: The root mean square error is expressed as: Where W is a diagonal matrix.
8. The method for rasterizing ground relative humidity according to claim 1, characterized in that: The method also includes classifying the goodness of fit of the relative humidity rasterized model using a determination coefficient, wherein the classification results include excellent, good, fair, and poor. The calculation formula of the determination coefficient is: Among them, i is the monitoring site data; y i The i-th actual monitoring data; f i The relative humidity corresponding to the i-th station is rasterized; is the mean of actual monitoring data.
9. A ground relative humidity grid system, characterized in that: include: The acquisition module is used to collect meteorological data from regional weather stations and meteorological site monitoring data; The model building module is used to combine the meteorological data of the regional weather stations and the site monitoring data in various ways to build multiple relative humidity models with a daily time resolution; a successful model determination module, configured to use the signal degrees of freedom and the residual signal degrees of freedom as evaluation indicators for the relative humidity model to determine a successful model among all relative humidity models; wherein the successful model is determined based on whether the degrees of freedom and the residual signal degrees of freedom are less than half of the number of stations; An evaluation index determination module is used to use the generalized cross validation method and the maximum likelihood estimation method as the selection criteria for the optimal model of all successful relative humidity models, combined with the square root error, to determine the best evaluation index; A selection module is used to calculate the daily relative humidity rasterization ratios of different successful models based on the best evaluation index, and select the best relative humidity rasterization model based on the largest relative humidity rasterization ratio; The calculation module is used to input ground monitoring data into the relative humidity rasterization optimal model and output the relative humidity rasterization results.
10. A ground relative humidity rasterization computer device, characterized in that: include: A memory, a processor, and a computer program stored in the memory, wherein when the processor executes the computer program, the steps of the ground relative humidity rasterization method according to any one of claims 1 to 8 are implemented.