A two-stage multi-source rainfall data fusion method and system based on three-source matching for data-scarce areas

By combining the ETCC method and the quantile mapping (QM) method, the accuracy problem of multi-source rainfall data fusion in data-scarce areas was solved, and high-precision rainfall datasets were generated, improving the data detection capability and correlation.

CN120951263BActive Publication Date: 2026-04-03ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing rainfall data fusion methods are difficult to effectively fuse multiple rainfall data in areas with scarce data, and the accuracy of the fused data depends on the quality of the reference data, which cannot meet the requirements for high accuracy.

Method used

Error analysis was performed using the ETCC method to calculate the fusion weights of multiple rainfall data, and bias correction was performed using the quantile mapping (QM) method to achieve two-stage fusion of multi-source rainfall data, reduce systematic errors, and improve data accuracy.

Benefits of technology

It improved the accuracy of rainfall data and the ability to detect rainfall events in areas with insufficient data, reduced the false alarm rate, and improved key success indices and correlation coefficients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120951263B_ABST
    Figure CN120951263B_ABST
Patent Text Reader

Abstract

This invention discloses a two-stage multi-source rainfall data fusion method and system based on three-source matching for data-scarce regions. The steps include: Step 1: Calculating fusion weights for the original rainfall dataset using the ETCC method, and weighting the original rainfall dataset according to the fusion weights to obtain a preliminary fused rainfall dataset; Step 2: Applying a bias correction method to the preliminary fused rainfall dataset to obtain a two-stage fused rainfall dataset. This invention addresses the two major error sources of rainfall data—random error and systematic error—by correcting these errors using the ETCC method and bias correction method, thereby improving the quality of the final rainfall dataset and enhancing its performance in terms of accuracy and rainfall event detection capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of hydrology and meteorology, and relates to a new method for generating higher quality rainfall data using existing rainfall data from multiple different sources. In particular, it relates to a two-stage multi-source rainfall data fusion method and system based on three-source matching for areas with insufficient data. Background Technology

[0002] Rainfall data is a crucial dataset for research in fields such as drought and flood disasters and runoff simulation. Traditional hydrological observation station data offers high accuracy but is limited by topography and other conditions, resulting in uneven spatial distribution. Other rainfall data also have their advantages and disadvantages. For example, radar and satellite rainfall data, while possessing good spatiotemporal continuity, have limited accuracy due to sensor performance and inversion algorithms. Further analysis of rainfall data reveals that its accuracy and uncertainty are influenced by the choice of data assimilation methods and the quality of the observation data. However, these rainfall datasets also contain information about the spatiotemporal distribution characteristics of rainfall data, which can be further improved through multi-source rainfall data fusion. However, existing rainfall data fusion methods suffer from difficulties in fusing multiple rainfall data sources, or the accuracy of the fused data heavily depends on the quality of reference and station data, making them unsuitable for rainfall data fusion in areas with limited data.

[0003] The extended three-source matching ETCC method provided by this invention is a referenceless error analysis method that can estimate the error of multiple rainfall datasets of different types and sources. Based on this, the weights of the rainfall data can be calculated and the data can be fused to combine the advantages of each dataset and obtain higher accuracy rainfall data.

[0004] The ETCC method is an extended TC fusion method. TC-based rainfall data fusion methods do not rely on station observation data or other reference datasets. They directly compare three types of input rainfall data to obtain estimates of the random error variance of each dataset compared to the "true value," and calculate the fusion weights of each rainfall data point by minimizing the estimated standard deviation of the fused rainfall data. The ETCC method, based on the error variance estimates of the TC method, aims to maximize the correlation coefficient between the fused rainfall data and the estimated "true value," calculating fusion weights and performing data fusion. The Quantile Mapping (QM) correction method can correct the overall bias of prediction or simulation results based on the cumulative empirical distribution function of rainfall data. Summary of the Invention

[0005] To address the shortcomings of existing data fusion methods, the present invention aims to provide a two-stage multi-source rainfall data fusion method and system based on three-source matching for data-scarce regions. This method provides a way to fuse multi-source rainfall data in data-scarce regions, obtain high-quality rainfall data, and promote related research and development.

[0006] To achieve the above objectives, a two-stage multi-source rainfall data fusion method based on three-source matching for data-scarce areas is proposed, comprising the following steps:

[0007] Step 1: Calculate the fusion weights of the original rainfall dataset using the ETCC method, and then weight the original rainfall dataset according to the fusion weights to obtain a preliminary fused rainfall dataset;

[0008] Step 2: The preliminary fused rainfall dataset is processed using a bias correction method to reduce the systematic error of the preliminary fused dataset, thereby obtaining a two-stage fused rainfall dataset.

[0009] Furthermore, in step 1, the original dataset includes three existing relatively independent rainfall datasets from different sources; the existing rainfall datasets include: the satellite rainfall dataset TRMM, PERSIANN, the reanalysis rainfall dataset ERA5, the bottom-up dataset SM2RAIN, and the land data assimilation dataset GLDAS.

[0010] Furthermore, in step 1, the initially fused rainfall dataset specifically includes:

[0011] (1) The ETCC method was used to preprocess the three types of rainfall datasets in the original rainfall dataset to unify the spatiotemporal range and resolution.

[0012] (2) Calculate the correlation coefficient estimates of the three rainfall datasets relative to the "true value";

[0013] (3) Based on the definition of the preliminary fused rainfall dataset and the definition of the correlation coefficient, the equation for the correlation coefficient estimate of the preliminary fused rainfall dataset is obtained by using the correlation coefficient estimate and the relationship between it and the covariance.

[0014] (4) Based on the equation, the optimization objective is to minimize the estimated correlation coefficient of the preliminary fused rainfall dataset, and the fusion weight is used as the optimization variable. A genetic algorithm is used to assign fusion weights to the three rainfall datasets. The three rainfall datasets are fused according to the fusion weights, and the preliminary fused rainfall dataset is obtained after fusion.

[0015] Furthermore, in step (2), the calculation of the correlation coefficient estimates of the three rainfall datasets relative to the "true value" is specifically as follows:

[0016]

[0017] Among them, CC i,t CC represents the estimated correlation coefficient between the i-th rainfall dataset and the "true value". i,j Let be the correlation coefficient between the i-th and j-th rainfall datasets, where i ≠ j;

[0018] In step (3), based on the definition of the initially fused rainfall dataset and the definition of the correlation coefficient, the equation for the estimated correlation coefficient of the initially fused rainfall dataset is obtained using the correlation coefficient estimate and its relationship with the covariance:

[0019]

[0020] Where, σ i 2 Let Cov be the variance of the i-th type of rainfall data. i,t Let σ be the covariance between the i-th rainfall dataset and the "true value". M 2 ω represents the variance of the merged rainfall data. i Let be the fusion weights for the i-th type of rainfall dataset;

[0021] In step (4), the formula for calculating the initially fused rainfall dataset is:

[0022]

[0023] In the formula, MTC represents the initially fused rainfall dataset, and ω i The fusion weights G of the i-th rainfall dataset i Let be the dataset for the i-th type of rainfall.

[0024] Furthermore, in step 2, the deviation correction method is one of LS, QM, and QDM.

[0025] Furthermore, in step 2, the QM method is specifically used for deviation correction, and the steps are as follows:

[0026] I. Select a rainfall dataset as the reference dataset for QM correction, and calculate its empirical cumulative distribution function eCDF by grid point;

[0027] II. For the initially fused rainfall dataset, calculate its empirical cumulative distribution function eCDF per grid point;

[0028] III. Calculate the frequency corresponding to each preliminary fused rainfall data value. Based on the eCDF of the reference dataset, find the eCDF quantile value of the reference dataset at each frequency. Use this value as the preliminary fused rainfall data value after correction at the corresponding frequency, i.e., the two-stage fused rainfall value. Organize the two-stage fused rainfall values ​​calculated at all frequencies to obtain the two-stage fused rainfall dataset.

[0029] Furthermore, in step III, the frequency corresponding to each initially fused rainfall data value is calculated. Based on the eCDF of the reference dataset, the eCDF quantile value of the reference dataset at each frequency is obtained. This quantile value is then used as the calculation formula for the initially fused rainfall data value after frequency correction, i.e., the two-stage fused rainfall value:

[0030]

[0031] In the formula, x represents the initial fused rainfall data value after correction at time t. m (t) represents the initially merged rainfall data value at time t, F m,q The empirical cumulative distribution function eCDF, F of the initially merged rainfall dataset o,q The empirical cumulative distribution function eCDF is used for the reference dataset.

[0032] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the two-stage multi-source rainfall data fusion methods based on three-source matching for data-scarce regions.

[0033] A two-stage multi-source rainfall data fusion system based on three-source matching for data-scarce areas, the system comprising:

[0034] One or more processors;

[0035] Memory, used to store one or more programs;

[0036] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the two-stage multi-source rainfall data fusion methods based on three-source matching for data-deficient areas as described in the one or more embodiments.

[0037] Beneficial effects:

[0038] This invention addresses the two main sources of error in rainfall data: random error and systematic error. It employs the ETCC method to improve the random error component of the original rainfall dataset and reduces the systematic error of the initially fused rainfall dataset through a bias correction method. Finally, it obtains a two-stage fused rainfall dataset, thereby improving the quality of the rainfall dataset and enhancing its performance in various aspects, such as accuracy and rainfall event detection capabilities. Attached Figure Description

[0039] Figure 1 This is a schematic diagram of the overall process of one embodiment of the present invention; Figure 2 This is an example of the spatial distribution of data fusion weights for the ETCC method in the Yarlung Tsangpo River Basin according to an embodiment of the present invention;

[0040] Figure 3 This is a box plot showing the accuracy evaluation results of the data before and after fusion in the Yarlung Tsangpo River Basin based on the CMFD dataset, according to one embodiment of the present invention.

[0041] Figure 4 This is an evaluation result of the detection capability of data before and after fusion in one embodiment of the present invention;

[0042] Figure 5 This is a spatial distribution diagram of the evaluation results of the correlation coefficient evaluation index of data before and after fusion, according to an embodiment of the present invention.

[0043] Figure 6 This is a spatial distribution diagram of the evaluation results of the root mean square error evaluation index of the data before and after fusion, according to an embodiment of the present invention. Detailed Implementation

[0044] The technical solution of the present invention will be further described in detail below through examples and with reference to the accompanying drawings.

[0045] Example:

[0046] like Figure 1 The diagram shows the entire data fusion, correction, and evaluation process in this embodiment.

[0047] The study area was first selected as the upper and middle reaches of the Yarlung Tsangpo River, with the Nuxia hydrological station as the outlet station, hereinafter referred to as the Yarlung Tsangpo River Basin. Three multi-source rainfall datasets were selected as the original rainfall datasets: the PERSIANN satellite remote sensing rainfall dataset, the GLDAS land data assimilation dataset, and the ERA5 reanalysis dataset. The CMFD rainfall dataset was used as the reference dataset for evaluating the fusion effect. Information on each rainfall dataset used is shown in Table 1.

[0048] Information on the rainfall dataset used in Table 1

[0049]

[0050] First, the three types of raw precipitation data were preprocessed, with a unified daily temporal resolution and a spatial resolution of 0.25°×0.25°, covering the period from 2000 to 2018. After processing, each precipitation data point within the study area had 365 grid points, spanning 6940 days. The ETCC method was used to estimate the error variance of the raw precipitation data, calculate the fusion weights, and fuse the precipitation data according to these weights to obtain the preliminary fused precipitation data, denoted as ETCC. The QM method was then used to perform overall correction on the preliminary fused precipitation data, resulting in the two-stage fused precipitation data, denoted as qmETCC.

[0051] This embodiment also calculates various evaluation indicators of the rainfall data before and after fusion, and evaluates the performance of the fused rainfall dataset in terms of accuracy and rainfall event detection capability, as follows:

[0052] 1) Accuracy indicators: correlation coefficient (CC), root mean square error (RMSE), Kling-Gupta efficiency coefficient (KGE)

[0053] 2) Rainfall event detection capability indicators: Point of Detection (POD), False Alarm Rate (FAR), Success Rate (SR), Bias Score (BS), Critical Success Index (CSI)

[0054] The specific calculation formulas, value ranges, and optimal values ​​for each indicator are shown in Table 2.

[0055] Table 2 Summary of Evaluation Indicators

[0056]

[0057] Among them, G i S i These are the rainfall values ​​for the reference dataset and the dataset to be evaluated, respectively. 1 represents the average rainfall values ​​of the reference dataset and the dataset to be evaluated, respectively; BR represents the bias ratio; RV represents the variation ratio; H represents the number of correctly predicted rainfall events; W represents the number of falsely reported rainfall events; and M represents the number of missed rainfall events.

[0058] Based on CMFD rainfall data, the accuracy (correlation coefficient CC, root mean square error RMSE, Kling-Gupta efficiency coefficient KGE) and rainfall event detection capability (hit rate POD, false alarm rate FAR, bias score BS, critical success index CSI) of raw rainfall data, preliminary fused rainfall data, and two-stage fused rainfall data were evaluated.

[0059] Table 3 summarizes the median evaluation results for each indicator. The bolded values ​​represent the best-performing results for that indicator. Statistical analysis shows that in this example, the two-stage fusion of precipitation data outperformed other precipitation data (including raw precipitation data and ETCC data) in both accuracy and precipitation event detection capability. Regarding accuracy, the median KGE of the qmETCC data was higher than that of the raw precipitation data and ETCC data (the median KGE of qmETCC data was 0.50, the best-performing raw precipitation data was 0.32, and the median KGE of ETCC data was 0.14). In terms of detection capability, the two-stage fusion method significantly reduced the FAR of the precipitation data and improved BS and CSI.

[0060] Table 3 Summary of Rainfall Data Assessment Results

[0061]

[0062]

[0063] Furthermore, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the two-stage multi-source rainfall data fusion method based on error analysis for data-deficient areas as described in any one of the claims.

[0064] This invention also provides a two-stage multi-source rainfall data fusion system based on error analysis for data-scarce areas, the system comprising:

[0065] One or more processors;

[0066] Memory, used to store one or more programs;

[0067] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the two-stage multi-source rainfall data fusion methods based on error analysis for data-deficient areas as described in the one or more embodiments.

[0068] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0069] The above description is merely an example implementation of the present invention and is not intended to limit the invention. The composition of the original rainfall dataset and the reference dataset, as well as the selection of evaluation indicators, can be specifically formulated according to different research problems. For researchers in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the scope of the claims of the present invention should be within the protection scope of the present invention.

Claims

1. A two-stage multi-source rainfall data fusion method based on three-source matching for data-scarce areas, characterized in that, Includes the following steps: Step 1: Calculate the fusion weights of the original rainfall dataset using the ETCC method, and then weight the original rainfall dataset according to the fusion weights to obtain a preliminary fused rainfall dataset; Step 2: Apply a bias correction method to the initially fused rainfall dataset to obtain a two-stage fused rainfall dataset; In step 1, the initially fused rainfall dataset specifically includes: (1) The ETCC method was used to preprocess the three types of rainfall datasets in the original rainfall dataset to unify the spatiotemporal range and resolution; (2) Calculate the estimated correlation coefficients of the three rainfall datasets relative to the "true value"; (3) Based on the definition of the preliminary fused rainfall dataset and the definition of the correlation coefficient, the equation for the correlation coefficient estimate of the preliminary fused rainfall dataset is obtained by using the correlation coefficient estimate and the relationship between it and the covariance. (4) Based on the above equation, the optimization objective is to minimize the estimated correlation coefficient of the preliminary fused rainfall dataset, and the fusion weight is used as the optimization variable. A genetic algorithm is used to assign fusion weights to the three rainfall datasets. The three rainfall datasets are fused according to the fusion weights, and the preliminary fused rainfall dataset is obtained after fusion. In step (2), the calculation of the correlation coefficient estimates of the three rainfall datasets relative to the "true value" is specifically as follows: (1) in, For the first The correlation coefficient estimates between the various rainfall datasets and the "true values" Let be the correlation coefficient between the i-th and j-th rainfall datasets. ; In step (3), based on the definition of the initially fused rainfall dataset and the definition of the correlation coefficient, the equation for the estimated correlation coefficient of the initially fused rainfall dataset is obtained using the correlation coefficient estimate and its relationship with the covariance: (2) in, For the first The variance of the rainfall data Let the covariance be the variance between the i-th rainfall dataset and the "true value". The variance of the merged rainfall data; Let be the fusion weights for the i-th type of rainfall dataset; In step (4), the formula for calculating the initially fused rainfall dataset is: (3) In the formula, MTC represents the initially fused rainfall dataset. The fusion weights of the i-th type of rainfall dataset, This is the dataset for the i-th type of rainfall; In step 2, the QM method is specifically used for deviation correction, and the steps are as follows: I. Select a rainfall dataset as the reference dataset for QM correction, and calculate its empirical cumulative distribution function eCDF by grid point; II. For the initially fused rainfall dataset, calculate its empirical cumulative distribution function eCDF per grid point; III. Calculate the frequency corresponding to each preliminary fused rainfall data value. Based on the eCDF of the reference dataset, find the eCDF quantile value of the reference dataset at each frequency. Use this value as the preliminary fused rainfall data value after correction at the corresponding frequency, i.e., the two-stage fused rainfall value. Organize the two-stage fused rainfall values ​​calculated at all frequencies to obtain the two-stage fused rainfall dataset.

2. The two-stage multi-source rainfall data fusion method based on three-source matching for data-scarce areas as described in claim 1, characterized in that, In step 1, the original dataset includes three existing rainfall datasets from different sources and which are relatively independent. The existing rainfall datasets include: the satellite rainfall dataset TRMM, PERSIANN, the reanalysis rainfall dataset ERA5, the bottom-up dataset SM2RAIN, and the land data assimilation dataset GLDAS.

3. The two-stage multi-source rainfall data fusion method based on three-source matching for data-scarce areas as described in claim 1, characterized in that, In step 2, the deviation correction method is one of LS, QM, and QDM.

4. The two-stage multi-source rainfall data fusion method based on three-source matching for data-scarce areas as described in claim 1, characterized in that, In step III, the frequency corresponding to each initially fused rainfall data value is calculated. Based on the eCDF of the reference dataset, the eCDF quantile value of the reference dataset at each frequency is obtained. This quantile value is then used as the calculation formula for the initially fused rainfall data value after frequency correction, i.e., the two-stage fused rainfall value: (4) In the formula, The value of the initially fused rainfall data after correction at time t. The rainfall data values ​​initially merged at time t. The empirical cumulative distribution function eCDF of the initially merged rainfall dataset. The empirical cumulative distribution function eCDF is used for the reference dataset.

5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the two-stage multi-source rainfall data fusion method based on three-source matching for data-deficient areas as described in any one of claims 1-4.

6. A two-stage multi-source rainfall data fusion system based on three-source matching for data-scarce areas, characterized in that, The system includes: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the two-stage multi-source rainfall data fusion method based on three-source matching for data-deficient areas as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Multi-source rainfall data fusion method based on mathematical uncertainty analysis

    CN115755221A

  • Comprehensive evaluation method and device for climate mode prediction data and storage medium

    CN119397179A