Multi-source satellite precipitation product fusion method and device combining sliding quadruple permutation and random forest spatial interpolation
By combining sliding quadruple arrangement and random forest space interpolation, the multi-source satellite precipitation product fusion method is solved in the prior art that the actual data cannot be effectively utilized and time-varying characteristics of errors is considered, and a higher precision precipitation estimation is achieved.
Patent Information
- Application Number
- CN202411722030.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-11-28
AI Technical Summary
The prior art cannot effectively utilize the measured data information in the integration of satellite precipitation products, and the commonly used triple and quadruple arrangement technologies cannot consider the time-varying characteristics of the error of the precipitation product collection member, resulting in insufficient precipitation estimation accuracy.
Combining the multi-source satellite precipitation product fusion method that combines sliding quadruple arrangement and random forest spatial interpolation, the error time-varying characteristics of precipitation products are estimated through sliding quadruple arrangement analysis technology, and actual measurement information is introduced through the random forest spatial interpolation model, weighted average and spatial interpolation are performed to improve the precipitation estimation accuracy.
Through the fusion method of sliding quadruple arrangement and random forest spatial interpolation, the error characteristics and spatial distribution of precipitation products can be more accurately grasped, the accuracy and reliability of precipitation estimation can be improved, and the problem that the fusion effect is affected by product quality in the prior art is overcome.
Smart Images

Figure CN119669376B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for fusing multi-source satellite precipitation products, belonging to the field of multi-source data fusion and precipitation estimation in hydrometeorology. Background Art
[0002] Precipitation is the basic driving factor of the global water cycle. Obtaining accurate precipitation information in a basin is crucial for basin water resources management and hydrological modeling. Ground monitoring data provides the most reliable precipitation data for different applications, and the accuracy of such precipitation data is the highest. However, the sparse and uneven spatial distribution of rain gauges limits its accurate grasp of precipitation spatial heterogeneity, which is particularly prominent in remote mountainous areas and underdeveloped regions. In recent decades, satellite precipitation products and reanalysis products (numerical models based on physics and dynamics) with quasi-global coverage have emerged one after another. Their higher spatio-temporal resolution and global coverage characteristics have gradually made them a new way to replace traditional observed precipitation and have been widely used in hydrology.
[0003] Satellite precipitation products inevitably have random and systematic errors, mainly caused by measurement conditions, retrieval algorithms, and satellite sensors. The performance of precipitation products is affected by background conditions, seasons, latitudes, and meteorological conditions. Direct application of the original products may bring greater uncertainties to precipitation analysis and corresponding hydrological applications. Fusing multi-source precipitation products and combining them with site observation data has become a popular method to improve the quality of satellite products. However, on the one hand, the current commonly used precipitation fusion method based on triple and quadruple permutation techniques cannot consider the time-varying characteristics of the ensemble member errors of precipitation products; on the other hand, such methods perform prior estimation of the true value when the measured data is unknown, so they cannot effectively utilize the measured data information. Machine learning algorithms usually establish a non-linear relationship between the original precipitation products and the measured precipitation on the basis of incorporating feature covariates, and then predict the true value precipitation through the original precipitation products. Such methods cannot screen precipitation products in advance and perform preliminary processing on precipitation products; in addition, the accuracy and quality of the original precipitation products will affect the machine learning fusion effect. For example, precipitation products with very poor quality may lead to overfitting problems in machine learning. Summary of the Invention
[0004] Object of the Invention: Aiming at the deficiencies of the prior art, the present invention proposes a method and device for fusing multi-source satellite precipitation products by combining sliding quadruple permutation and random forest spatial interpolation, which can make full use of the quadruple permutation analysis technology to estimate the time-varying characteristics of the errors of the original precipitation products and perform preliminary weighted average estimation, and then introduce measured information through the random forest spatial interpolation model to further improve the precipitation estimation accuracy.
[0005] Technical Solution: A method for fusing multi-source satellite precipitation products by combining sliding quadruple permutation and random forest spatial interpolation includes the following steps:
[0006] (1) Obtain multi-source satellite precipitation products, multi-source environmental covariate data, and measured station daily precipitation data. Resample the multi-source precipitation products and environmental covariate data to unify them to the same spatio-temporal resolution;
[0007] (2) Based on the unified spatio-temporal resolution, obtain the time series of multi-source satellite precipitation products at each grid point on the basin surface. Combine any four precipitation product sequences to construct a sample combination set. Estimate the error variance and error covariance of the four precipitation sequences in each sample combination within each sliding window period based on the sliding quadruple permutation analysis method. Arithmetically average the error variances and error covariances representing the same precipitation product in different sample combinations to obtain the error variance of the multi-source precipitation products and the inter-product error covariance. Construct an error matrix based on the product error variance and inter-product error covariance, calculate the weight coefficients of each product, and then calculate the weighted average satellite precipitation sequence;
[0008] (3) On the grid with measured stations, use the measured station precipitation as the dependent variable, and use the weighted average satellite precipitation sequence, multi-source environmental covariate sequence, and the precipitation sequences and station distances of surrounding measured stations as independent variables to construct a sample set, divide the training samples and validation samples, and train the random forest spatial interpolation model;
[0009] (4) According to the trained random forest spatial interpolation model, on the grid without measured stations, use the weighted average satellite precipitation sequence of the grid, the precipitation sequences and station distances of surrounding measured stations, and the environmental covariate sequence as the input of the random forest spatial interpolation model to predict the precipitation sequence of the grid point, and integrate the predicted precipitation sequences of each grid point to finally form a new set of precipitation product data.
[0010] The present invention also provides a multi-source satellite precipitation product fusion device combining sliding quadruple permutation and random forest spatial interpolation, including:
[0011] A data preprocessing module, configured to obtain multi-source satellite precipitation products, multi-source environmental covariate data, and measured station daily precipitation data, and resample the multi-source precipitation products and environmental covariate data to unify them to the same spatio-temporal resolution;
[0012] The multi-source precipitation product weighted average module is used to obtain the time series of multi-source satellite precipitation products at each grid point on the basin surface based on a unified spatio-temporal resolution, construct a sample combination set by combining any four precipitation product sequences, estimate the error variance and error covariance of the four precipitation sequences in each sample combination within each sliding window period based on the sliding quadruple permutation analysis method, perform arithmetic averaging on the error variances and error covariances representing the same precipitation product in different sample combinations to obtain the error variance of the multi-source precipitation products and the inter-product error covariance, construct an error matrix based on the product error variance and the inter-product error covariance, calculate the weight coefficients of each product, and then calculate the weighted average satellite precipitation sequence;
[0013] The random forest spatial interpolation model training module is used to construct a sample set on the grid with measured stations, taking the precipitation at the measured stations as the dependent variable and the weighted average satellite precipitation sequence, the multi-source environmental covariate sequence, the precipitation sequence and the station distance of the surrounding measured stations as the independent variables, divide the training samples and the validation samples, and train the random forest spatial interpolation model;
[0014] The precipitation prediction module is used to, based on the trained random forest spatial interpolation model, on the grid without measured stations, take the weighted average satellite precipitation sequence of the grid, the precipitation sequence and the station distance of the surrounding measured stations, and the environmental covariate sequence as the input of the random forest spatial interpolation model to predict the precipitation sequence of the grid point, and integrate the predicted precipitation sequences of each grid point to finally form a new set of precipitation product data.
[0015] The present invention also provides a computer device, including: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and when the program is executed by the processor, the steps of the multi-source satellite precipitation product fusion method combining sliding quadruple permutation and random forest spatial interpolation as described above are implemented.
[0016] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the multi-source satellite precipitation product fusion method combining sliding quadruple permutation and random forest spatial interpolation as described above are implemented.
[0017] Compared with the prior art, the present invention has the following beneficial effects:
[0018] The present invention estimates the error variance and error covariance of precipitation sequences within each sample combination set during the sliding window period based on the sliding quadruple permutation analysis technique by obtaining multi-source precipitation product sequences at each grid point on the basin surface, constructs a sample combination set, constructs an error feature matrix and calculates the weight coefficients of each precipitation product based on this, and then calculates the weighted average satellite precipitation, which can grasp the time-varying characteristics of the errors of the original precipitation products, and introduces measured precipitation information through the random forest spatial interpolation model to achieve more accurate prediction. The present invention overcomes the deficiencies in the current satellite precipitation product fusion field where measured information cannot be incorporated, as well as the problem that the fusion effect is affected by the quality of the original products when using machine learning to fuse satellite precipitation products with measured precipitation, and is of great significance for improving the fusion accuracy of satellite precipitation products. Description of the Drawings
[0019] Figure 1 It is a flowchart of a multi-source satellite precipitation product fusion method combining sliding quadruple permutation and random forest spatial interpolation according to the present invention. Detailed Embodiments
[0020] The technical solution of the present invention will be further described below with reference to the drawings.
[0021] As Figure 1 shown, a multi-source satellite precipitation product fusion method combining sliding quadruple permutation and random forest spatial interpolation includes the following steps:
[0022] Step (1), obtain multi-source satellite precipitation products, multi-source environmental covariate data, and measured station daily precipitation data, and resample the multi-source precipitation products and environmental covariate data to unify them to the same spatio-temporal resolution;
[0023] In the embodiments of the present invention, taking the Yili River Basin in Xinjiang as the research area, data of six satellite precipitation products including ERA5-Land, SM2RAIN, CHIPRS, CMORPH, IMERG, and PERSIANN-CDR from 2010 to 2020 were collected. At the same time, elevation data with a resolution of 90 m provided by SRTM data was collected, and slope, aspect, and Tpi data were calculated based on this elevation data. Surface temperature data provided by MODIS satellite, NDVI data, and 10 m wind speed and wind direction data provided by ERA5-Land were collected. The above data have different temporal and spatial resolutions. For the problem of different spatial scales, the bilinear interpolation method was used to unify different spatial resolutions to 0.1 degrees. For the problem of different time scales, the next-day scale was summed up, and the 8-day average data provided by MODIS was considered that the daily average within 8 days was equal to the 8-day average value. Finally, the multi-source satellite precipitation product data and environmental covariate data collected were unified into a spatio-temporal scale of 0.1 degrees and daily scale. There are 22 measured stations in the research area, and the measured daily precipitation data of these 22 stations from 2011 to 2019 were collected from the hydrological yearbook.
[0024] Step (2), obtain the multi-source precipitation product sequences at each grid point on the basin surface, construct a sample combination set, estimate the variation characteristics of the error variance and error covariance of the precipitation sequences within each sample combination over time based on the sliding quadruple permutation analysis technique, construct an error characteristic matrix and calculate the product weight matrix based on this, and perform weighted averaging according to the product weights to obtain the weighted average satellite precipitation data;
[0025] In the embodiments of the present invention, based on the six satellite precipitation products P1, P2, …, P6 (representing ERA5-Land, SM2RAIN, CHIPRS, CMORPH, IMERG, and PERSIANN-CDR respectively) after spatio-temporal resampling in step (1), at each grid point on the basin surface, 6 precipitation time series (P 1,1 , P 1,2 , …, P 1,t ), (P 2,1 , P 2,2 , …, P 2,t ), …, (P 6,1 , P 6,2 , …, P 6,t ) can be obtained. Where t represents the total time length, P 1,1 represents the precipitation value of the first precipitation product at the first moment, and P 1,t represents the precipitation value of the first precipitation product at the t-th moment. For these 6 time series, any 4 are combined, and a total of 15 sample combinations can be obtained. The 4 sequences within each sample combination are denoted as (W1, W2, …, W t)、(X1, X2, …, X t )、(Y1, Y2, …, Y t )、(Z1, Z2, …, Z t )。
[0026] Let each sliding window period be 101 days. If the weighted average precipitation during the period from day t to day t + k is required, then multi-source precipitation products during the period from day t - 50 to day t + k + 50 need to be collected. Given that the weighted average precipitation during the period from 20110101 to 20191231 is required in the embodiment, the first sliding window period should be from 20101112 to 20110220.
[0027] Taking the period from 20101112 to 20110220 as the first sliding window period, the four precipitation sequences in each sample combination within the sliding window period are (W1, W2, …, W 101 )、(X1, X2, …, X 101 )、(Y1, Y2, …, Y 101 )、(Z1, Z2, …, Z 101 ), and a column vector is constructed
[0028] Taking the grid point with the central coordinate of (82°E, 47°N) as an example, the W, X, Y, Z vectors of the sample combination including ERA5-Land, SM2RAIN, CHIRPS, and CMORPH within the window period from 20101112 to 20110220 are shown in Table 1.
[0029] Table 1 Results of W, X, Y, Z vectors
[0030]
[0031] Calculate the covariance between each vector
[0032] Construct the covariance matrix B and the coefficient matrix A through the covariance calculated above:
[0033]
[0034] According to the least square solution method of the quadruple permutation analysis method, by solving the equation A*U = B, the error eigenmatrix U (the error eigenmatrix is used to calculate the error variance of each vector and the error covariance between vectors) can be obtained, and the solution method is as follows:
[0035]
[0036] where β W 、βX and β Y and β Z represent the magnitude conversion coefficients of the W, X, Y, and Z sequences relative to the unknown true precipitation sequence, respectively. C pp is the variance of the unknown true precipitation sequence, and C εWεW and C εXεX and C εYεY and C εZεZ represent the error variances of the W, X, Y, and Z sequences without magnitude conversion, respectively. C εYεZ represents the error covariance between Y and Z without magnitude conversion.
[0037] Based on the error feature matrix U, the error variances of W, X, Y, and Z in this sliding window period can be calculated The error covariance between Y and Z is
[0038] In the example of the present invention, taking the grid point with the central coordinates of (82°E, 47°N) as an example, the calculation results of the error variances and error covariances of the W, X, Y, and Z vectors in the sample combinations including ERA5-Land, SM2RAIN, CHIRPS, and CMORPH in the window period from 20101112 to 20110220 are shown in Table 2.
[0039] Table 2 Calculation results of error variances and error covariances
[0040]
[0041] For the error variances and error covariances representing the same precipitation product calculated within 15 sample combinations, arithmetic means are taken (for example, if the first sample combination includes ERA5-Land, SM2RAIN, CHIRPS, and CMORPH, and the second sample combination includes ERA5-Land, SM2RAIN, CHIRPS, and IMERG, then by taking the arithmetic mean of the error variances of ERA5-Land obtained within the two sample combinations, it is considered that this arithmetic mean represents the error variance of the ERA5-Land product). Finally, the error variances of each precipitation product and the error covariances between products Construct an error matrix The weight coefficients λ of each product i (i = 1 to m) are calculated by the following formula: E ij represents the element in the i-th row and j-th column of the error matrix. The weighted average precipitation data on the 51st day (i.e., 20110101) is
[0042] In the example of the present invention, for the grid point with the central coordinate of (82°E, 47°N), the weights of six precipitation products during the window period from 20101112 to 20110220 are shown in Table 3.
[0043] Table 3 Weight Coefficients of Each Precipitation Product
[0044]
[0045] Then, add one day to the initial date of the sliding window period, that is, take 20101113 - 20110221 as the new window period, repeat the above steps to calculate the error variance of precipitation products and the error covariance between products within the window period, construct an error matrix to calculate the weight coefficients of each product, and the weighted average precipitation data on the 51st day (i.e., 20110102) within the window period is Keep moving the window period forward until the weighted average precipitation at each time point is obtained, and finally obtain the weighted average precipitation sequence.
[0046] Step (3), on the grid with measured stations, take the precipitation at the measured stations as the dependent variable, and take the weighted average satellite precipitation data, multi-source environmental covariate data, and the precipitation and station distances of surrounding measured stations as independent variables to construct a sample set, divide the training samples and validation samples, and use the random forest spatial interpolation model for training;
[0047] The random forest spatial interpolation model is based on the random forest model. By regarding n observed values near the prediction point and the distances from these observation positions to the prediction position as additional covariates in the random forest, the spatial autocorrelation between observed values is considered. Its input is the estimated precipitation of the precipitation product at the prediction point, environmental covariate data, the observed precipitation near the prediction point, and the distance from the observation point to the prediction point, and the output is the true precipitation at the prediction point. Its expression is as follows:
[0048] z(L p ) = f(P(L p ), En(L p ), z(L1), d1, z(L2), d2,..., z(L k ), d k )
[0049] Where: z(L p ) is the true precipitation at the prediction point L p , P(L p ) is the estimated precipitation of the precipitation product at the prediction point L o , En(L p ) is the environmental covariate at the prediction point L pThe environmental covariate sequence at [location] includes longitude, latitude, elevation, slope, aspect, terrain position index, land surface temperature, normalized difference vegetation index NDVI, 10m wind speed, and 10m wind direction. z(L1), z(L2), …, z(L k ) is the measured precipitation at the k nearest observed locations near the prediction point, d1, d2, …, d k is the distance between the k nearest observed locations near the prediction point and the prediction point.
[0050] Further, step (3) includes:
[0051] Based on the collected measured precipitation data, for each measured site j (j = 1~n, where n is the total number of measured sites), extract the measured precipitation P_obs at each moment of this site j,t , and construct the dependent variable vector;
[0052] Extract the precipitation P_merge after weighted average in step (2) at the same moment corresponding to the grid where this site is located j,t , environmental variables (longitude lon j , latitude lat j , elevation Elevation j , slope Slope j , aspect Aspect j , terrain position index Tpi j , land surface temperature LST j,t , normalized difference vegetation index NDVI j,t , 10m wind speed Windspeed j,t , 10m wind direction Winddirection j,t ), and the measured precipitation P_obs_near1 at the same moment of the k nearest measured points near this grid point t , P_obs_near2 t , …, P_obs_neark t , and construct the independent variable matrix with the corresponding site distances D_obs_near1, D_obs_near2, …, D_obs_neark;
[0053] In the example of the present invention, k is taken as 10, that is, considering the influence of the precipitation of the 10 nearest measured points near the target grid point on the target grid point, partial results of the independent variable and dependent variable matrices are shown in Table 4.
[0054] Table 4 Independent variable and dependent variable matrices
[0055]
[0056] Table 4 (continued) Independent variable and dependent variable matrices
[0057]
[0058]
[0059] Furthermore, the independent variable and dependent variable data samples are divided into training samples and validation samples by ten-fold cross-validation. Random sampling is performed on the model parameters of the random forest spatial interpolation model, including the number of node splitting variables mtry, the minimum node size min.node.size, and the observation-to-sample ratio sample.fraction in the decision tree. The random forest spatial interpolation model is trained under each parameter combination condition, and the optimal parameter combination is selected based on the minimum sum of squared errors of the validation data set to obtain the trained random forest spatial interpolation model.
[0060] In the example of the present invention, the parameter results of the trained random forest spatial interpolation model are shown in Table 5.
[0061] Table 5 Parameters of the Random Forest Spatial Interpolation Model
[0062]
[0063] Step (4), based on the constructed random forest spatial interpolation model, on the grid without measured stations, using the weighted average satellite precipitation data of the grid, the precipitation and station distances of the measured stations around the grid, and the environmental covariate data as model inputs, the predicted precipitation of the grid point is output, and the predicted precipitation sequences of each grid point are integrated to finally form a new set of precipitation product data;
[0064] Furthermore, the specific steps of step (4) are as follows: For any grid without measured stations, extract the precipitation P_merge weighted average in step (2) on the grid t , environmental variables (longitude lon, latitude lat, elevation Elevation, slope Slope, aspect Aspect, terrain position index Tpi, land surface temperature LST t , normalized difference vegetation index NDVI t , 10m wind speed Windspeed t , 10m wind direction Winddirection t ), and the measured precipitation P_obs_near1 t , P_obs_near2 t , …, P_obs_neark t, and the corresponding site distances \(D_{obs\_near1}\), \(D_{obs\_near2}\), …, \(D_{obs\_neark}\), to construct an independent variable matrix; use this independent variable matrix as the input of the trained random forest spatial interpolation model in step (3), output the predicted precipitation sequence at this grid point, integrate the predicted precipitation sequences \(P_{predict}\) of each grid point, and finally form a new set of precipitation product data.
[0065] In the example of the present invention, the independent variable matrix of the grid without measured sites and partial results of the predicted precipitation output by the random forest spatial interpolation model are shown in Table 6.
[0066] Table 6 Independent variable matrix of the grid without measured sites and partial results of the predicted precipitation
[0067]
[0068] Table 6 (continued) Independent variable matrix of the grid without measured sites and partial results of the predicted precipitation
[0069]
[0070]
[0071] The present invention uses common evaluation indicators of precipitation products, namely Pearson correlation coefficient, root mean square error RMSE, probability of detection POD, false alarm rate FAR, and Heidke skill score HSS, to evaluate the fusion effect of a multi-source satellite daily-scale precipitation product fusion method that combines the sliding quadruple permutation analysis technique and the random forest spatial interpolation model.
[0072] In the example of the present invention, the calculation results of the evaluation indicators are shown in Table 7, and the results are presented in the form of mean ± standard deviation. The mean is the arithmetic mean of the precipitation product evaluation indicators at different measured sites, and the standard deviation is the standard deviation of the precipitation product evaluation indicators at different measured sites.
[0073] Table 7 Evaluation indicators of the original product and the fused product
[0074]
[0075]
[0076] According to the results in Table 7, it can be seen that after fusing precipitation products from different sources using the fusion method proposed by the present invention, the precipitation simulation accuracy has been significantly improved: the correlation coefficient has been increased to above 0.70, RMSE has been reduced to below 2.5, POD has been increased to above 0.95, FAR has been reduced to below 0.47, and HSS has been increased to above 0.55.
[0077] Based on the same technical concept as the method embodiment, another embodiment of the present invention further provides a multi-source satellite precipitation product fusion device combining sliding quadruple permutation and random forest spatial interpolation, including:
[0078] A data preprocessing module, configured to obtain multi-source satellite precipitation products, multi-source environmental covariate data, and measured station daily precipitation data, resample the multi-source precipitation products and environmental covariate data, and unify them to the same spatio-temporal resolution;
[0079] A multi-source precipitation product weighted average module, configured to obtain the time series of multi-source satellite precipitation products at each grid point on the basin surface based on the unified spatio-temporal resolution, combine any four precipitation product sequences to construct a sample combination set, estimate the error variance and error covariance of the four precipitation sequences in each sample combination within each sliding window period based on the sliding quadruple permutation analysis method, perform arithmetic averaging on the error variances and error covariances representing the same precipitation product in different sample combinations to obtain the error variance of the multi-source precipitation products and the error covariance between products, construct an error matrix based on the product error variance and the error covariance between products, calculate the weight coefficients of each product, and further calculate the weighted average satellite precipitation sequence;
[0080] A random forest spatial interpolation model training module, configured to, on the grid with measured stations, use the measured station precipitation as the dependent variable, and use the weighted average satellite precipitation sequence, multi-source environmental covariate sequence, and the precipitation sequence and station distance of surrounding measured stations as independent variables to construct a sample set, divide the training samples and validation samples, and train the random forest spatial interpolation model;
[0081] A precipitation prediction module, configured to, based on the trained random forest spatial interpolation model, on the grid without measured stations, use the weighted average satellite precipitation sequence of the grid, the precipitation sequence and station distance of surrounding measured stations, and the environmental covariate sequence as the input of the random forest spatial interpolation model to predict the precipitation sequence of the grid point, integrate the predicted precipitation sequences of each grid point, and finally form a new set of precipitation product data.
[0082] It should be understood that the multi-source satellite precipitation product fusion device combining sliding quadruple permutation and random forest spatial interpolation in the embodiment of the present invention can implement all the technical solutions in the above method embodiment. The functions of its respective functional modules can be specifically implemented according to the method in the above method embodiment, and the specific implementation process can refer to the relevant descriptions in the above embodiment, which will not be elaborated here.
[0083] The present invention also provides a computer device, comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and when the program is executed by the processor, the steps of the multi-source satellite precipitation product fusion method combining sliding quadruple permutation and random forest spatial interpolation as described above are implemented.
[0084] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the multi-source satellite precipitation product fusion method combining sliding quadruple permutation and random forest spatial interpolation as described above are implemented.
[0085] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device (system), a computer device, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0086] The present invention is described with reference to the flowchart of the method according to the embodiments of the present invention. It should be understood that each process in the flowchart and the combination of the processes in the flowchart can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 process or multiple processes.
[0087] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one Figure 1 process or multiple processes.
[0088] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 process or multiple processes.
Claims
1. A multi-source satellite precipitation product fusion method combining sliding quartet permutation and random forest spatial interpolation, characterized in that: The following steps are involved: (1) Obtain multi-source satellite precipitation products, multi-source environmental covariate data, and daily precipitation data at measured sites, and resample the multi-source precipitation products and environmental covariate data to the same temporal and spatial resolution; (2) Based on a unified spatiotemporal resolution, the time series of multi-source satellite precipitation products at each grid point on the watershed surface is obtained, and any four precipitation product sequences are combined to construct a sample combination set. The error variance and error covariance of the four precipitation sequences in each sample combination within each sliding window period are estimated based on the sliding quartet permutation analysis method. The error variance and error covariance of the same precipitation products in different sample combinations are arithmetic averaged to obtain the error variance of the multi-source precipitation products and the error covariance between products. An error matrix is constructed based on the product error variance and the error covariance between products, and the weight coefficient of each product is calculated, thereby obtaining the satellite precipitation sequence after weighted average. The sliding quartet permutation analysis method constructs a column vector according to the four precipitation sequences within the sliding window period, calculates the covariance between the vectors, constructs a covariance matrix, obtains the error characteristic matrix using the least squares solution, and calculates the error variance and error covariance of the four precipitation sequences within the sliding window period according to the error characteristic matrix. (3) On the grid with measured stations, the measured station precipitation is used as the dependent variable, and the weighted average satellite precipitation series, multi-source environmental covariate series, and the surrounding measured station precipitation series and station distance are used as independent variables. A sample set is constructed, divided into training samples and validation samples, and the random forest spatial interpolation model is trained; (4) Based on the trained random forest spatial interpolation model, on the grid without measured stations, the weighted average satellite precipitation sequence of the grid, the precipitation sequence of the measured stations around the grid, the station distance, and the environmental covariate sequence are used as the input of the random forest spatial interpolation model to predict the precipitation sequence of the grid point. The predicted precipitation sequences of each grid point are integrated to finally form a new set of precipitation products.
2. The method according to claim 1, characterized in that In the step (1), the multi-source precipitation products and environmental covariate data are resampled to the same spatiotemporal resolution, including: The bilinear interpolation method is used to unify the data of different spatial resolutions to the same resolution, and the data of the next-day scale are accumulated and summed to unify to the daily scale.
3. The method according to claim 1, characterized in that In step (2), the method for constructing the precipitation sample combination set is as follows: Based on the m satellite precipitation products P1, P2, ..., P after spatiotemporal resampling in step (1), m , at each grid point on the basin surface, we get m precipitation time series (P 1,1 , P 1,2 ,…,P 1,t )、(P 2,1 , P 2,2 ,…,P 2,t ),…,(P m,1 , P m,2 ,…,P m,t ), t is the total length of time, and any four of these m time series can be combined to obtain C m 4 There are four sample combinations, and the four sequences in each sample combination are represented as (W1, W2, ..., W t )、(X1,X2,…,X t )、(Y1,Y2,…,Y t )、(Z1,Z2,…,Z t ).
4. The method according to claim 1, characterized in that In step (2), the error variance and error covariance of the four precipitation sequences in each sample combination within each sliding window period are estimated based on the sliding four-fold permutation analysis method. The specific estimation method is as follows: Assume that the duration of each sliding window is S, and the four precipitation sequences within the sliding window are (W1, W2, …, W S )、(X1,X2,…,X S )、(Y1,Y2,…,Y S )、(Z1,Z2,…,Z S ), construct a column vector Calculate the covariance between vectors The covariance matrix B and the coefficient matrix A are constructed using the covariance calculated above: According to the least squares solution of the four-fold permutation analysis method, the error characteristic matrix U is obtained by solving the equation A*U=B. The solution method is as follows: Where: β W , β X , β Y , β Z Respectively represent the magnitude conversion coefficients of W, X, Y, and Z series relative to the unknown true precipitation series, C pp is the variance of the unknown precipitation true value sequence, C εWεW , C εXεX , C εYεY , C εZεZ Respectively represent the error variance of the W, X, Y, and Z series without magnitude reduction, C εYεZ Represents the error covariance between Y and Z without magnitude reduction; According to the calculated error characteristic matrix U, the error variances of the four sequences W, X, Y, and Z during the sliding window period are calculated as follows: and the error covariance between Y and Z is 5. The method according to claim 1, characterized in that In step (2), the error matrix is expressed as: in represents the error variance of each precipitation product, Represents the error covariance between products, and the weight coefficient λ of each product i Calculated by the following formula: E ij represents the element in the i-th row and j-th column of the error matrix. The weighted average precipitation data of the k-th day is 6. The method according to claim 1, characterized in that In the step (3), the environmental covariate data include longitude, latitude, elevation, slope, aspect, terrain position index, surface temperature, normalized difference vegetation index NDVI, 10m wind speed, and 10m wind direction data.
7. The method according to claim 1, characterized in that In step (3), the random forest spatial interpolation model expression is as follows: z(L p )=f(P(L p ),En(L p ),z(L1),d1,z(L2),d2,...,z(L k ),d k ) Where: z(L p ) is the predicted point L p The actual precipitation at p ) is the precipitation product at the prediction point L o Estimated precipitation at p ) is the predicted point L p The sequence of environmental covariates at , z(L1), z(L2), …, z(L k ) is the measured precipitation at the k nearest observation locations near the prediction point, d1, d2, …, d k is the distance between the predicted point and the k nearest observation positions near the predicted point.
8. A multi-source satellite precipitation product fusion device combining sliding quartet permutation and random forest spatial interpolation, characterized in that: include: The data preprocessing module is used to obtain multi-source satellite precipitation products, multi-source environmental covariate data, and daily precipitation data at measured sites, and resample the multi-source precipitation products and environmental covariate data to the same temporal and spatial resolution; The weighted average module of multi-source precipitation products is used to obtain the time series of multi-source satellite precipitation products at each grid point on the watershed surface based on a unified spatiotemporal resolution, combine any four precipitation product sequences to construct a sample combination set, estimate the error variance and error covariance of the four precipitation sequences in each sample combination within each sliding window period based on the sliding quartet permutation analysis method, perform arithmetic averaging on the error variance and error covariance representing the same precipitation product in different sample combinations, obtain the error variance of the multi-source precipitation product and the error covariance between products, construct an error matrix based on the product error variance and the error covariance between products, calculate the weight coefficient of each product, and then calculate the satellite precipitation sequence after weighted averaging; wherein the sliding quartet permutation analysis method constructs a column vector according to the four precipitation sequences within the sliding window period, calculates the covariance between the vectors, constructs a covariance matrix, obtains the error characteristic matrix using the least squares solution, and calculates the error variance and error covariance of the four precipitation sequences in the sliding window period according to the error characteristic matrix; The random forest spatial interpolation model training module is used to construct a sample set on a grid with measured stations, taking the measured station precipitation as the dependent variable, and the weighted average satellite precipitation sequence, multi-source environmental covariate sequence, and the surrounding measured station precipitation sequence and station distance as independent variables, divide the sample set into training samples and validation samples, and train the random forest spatial interpolation model; The precipitation prediction module is used to predict the precipitation sequence of the grid point on the grid without measured sites based on the trained random forest spatial interpolation model, using the weighted average satellite precipitation sequence of the grid, the precipitation sequence of measured sites around the grid, the site distance, and the environmental covariate sequence as the input of the random forest spatial interpolation model, and integrate the predicted precipitation sequences of each grid point to finally form a new set of precipitation products.
9. A computer device, characterized in that: include: one or more processors; Memory; And one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and when the programs are executed by the processors, the steps of the multi-source satellite precipitation product fusion method combining sliding quartet permutation and random forest spatial interpolation as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multi-source satellite precipitation product fusion method combining sliding quartet permutation and random forest spatial interpolation as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Multi-source rainfall data fusion method
CN114861840A
Predicted wind speed correction method based on wind speed correction factor representation model
CN118861673A