Method for predicting precision of remote sensing rainfall data of site-free area

By combining multiple accuracy assessment methods and deep neural network models, the consistency and interpretability issues of satellite remote sensing precipitation accuracy assessment were resolved, enabling reliable prediction of areas without monitoring stations and providing comprehensive and objective global satellite precipitation accuracy analysis.

CN122021982APending Publication Date: 2026-05-12CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNIV OF GEOSCIENCES (WUHAN)
Filing Date
2025-10-23
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing methods suffer from representativeness errors and poor consistency in evaluating the accuracy of satellite remote sensing precipitation. They also lack systematic comparison and fusion of multi-source evaluation results, resulting in low accuracy and practicality of predictions for areas without monitoring stations.

Method used

Preprocessing was performed using multi-source satellite remote sensing precipitation data and environmental auxiliary variable data. Accuracy was evaluated by combining the multiplicative triple combination analysis method, the extended dual-tool variable method, and the verification method based on actual measurement stations. A deep neural network model was constructed, and interpretability analysis was used to reveal the feature contribution, thus enabling the accuracy prediction of remote sensing precipitation data in areas without measurement stations.

Benefits of technology

It achieves reliable and accurate predictions in areas without monitoring stations, provides comprehensive and objective global satellite precipitation accuracy analysis, enhances the model's adaptability to complex environmental factors, improves the transparency and credibility of the prediction process, and provides a scientific basis for precipitation data selection and monitoring station deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122021982A_ABST
    Figure CN122021982A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of meteorology and hydrology, and particularly discloses a method for predicting precision of remote sensing rainfall data of a site-free area, and the method comprises the following steps: obtaining multi-source satellite remote sensing rainfall data and corresponding environment auxiliary variable data, and carrying out the preprocessing of the multi-source satellite remote sensing rainfall data and the corresponding environment auxiliary variable data; based on a multiplication triple matching analysis method, an extended double-tool variable method and a verification method based on an actual measurement site, performing precision evaluation on the multi-source satellite remote sensing rainfall data to obtain a corresponding precision evaluation index; constructing a deep neural network model, taking the precision evaluation index and the environment auxiliary variable data as input features, carrying out model training and cross validation, and revealing feature contributions by using an interpretability analysis method to obtain a trained deep neural network model; and using the trained deep neural network model to carry out precision prediction on the satellite remote sensing rainfall data of the site-free area, and outputting a precision prediction result. According to the method, reliable prediction of remote sensing rainfall data precision of a site-free area can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of meteorological and hydrological technology, and more specifically, relates to a method for predicting the accuracy of remote sensing precipitation data in areas without monitoring stations. Background Technology

[0002] Precipitation, as a core component of the water cycle, is a key element in meteorological and hydrological processes. Its highly variable and complex spatiotemporal distribution makes accurate measurement a significant challenge. Accurate estimation of satellite-sensed precipitation is crucial for understanding regional precipitation, preventing floods and droughts, and supporting eco-hydrological applications.

[0003] Currently, methods for assessing the accuracy of satellite remote sensing precipitation are mainly divided into point-scale and grid-scale methods. Point-scale methods calculate statistical indicators based on station observations, but are prone to representativeness errors due to uneven spatial distribution of stations. Grid-scale methods estimate accuracy by integrating multiple precipitation datasets in the absence of real values. However, different methods have poor consistency in assessment results under different geographical regions and environmental conditions, and lack systematic comparison and fusion of multi-source assessment results, which limits the comprehensiveness and reliability of accuracy assessment.

[0004] In recent years, machine learning techniques have been widely applied to satellite precipitation error modeling. These methods incorporate environmental variables such as topography, vegetation indices, and land use as input features and employ various algorithms to predict the spatial distribution of errors. However, existing methods still have significant shortcomings: insufficient coverage of global multi-source precipitation data assessments; inadequate analysis of accuracy differences between point-scale and grid-scale methods; failure to effectively integrate multiple accuracy assessment results with environmental factors as machine learning inputs; and weak model interpretability, resulting in low prediction accuracy and practicality in areas without monitoring stations.

[0005] Therefore, how to integrate multiple accuracy assessment methods and fuse multi-source environmental data to construct an interpretable machine learning model to achieve reliable prediction of the accuracy of remote sensing precipitation data in areas without monitoring stations is an urgent problem to be solved. Summary of the Invention

[0006] In view of the shortcomings of the existing technology, the purpose of this application is to provide a method for predicting the accuracy of remote sensing precipitation data in areas without monitoring stations, which can achieve reliable prediction of the accuracy of remote sensing precipitation data in areas without monitoring stations.

[0007] To achieve the above objectives, in a first aspect, this application provides a method for predicting the accuracy of remote sensing precipitation data in areas without monitoring stations, comprising the following steps:

[0008] S10: Acquire multi-source satellite remote sensing precipitation data and corresponding environmental auxiliary variable data, and preprocess the multi-source satellite remote sensing precipitation data and environmental auxiliary variable data to make all data have a unified spatiotemporal range, spatiotemporal resolution and projection coordinate system. S20. Based on the multiplicative triple combination analysis method, the extended dual-tool variable method, and the verification method based on actual measurement stations, the accuracy of multi-source satellite remote sensing precipitation data is evaluated to obtain the corresponding accuracy evaluation index; the accuracy evaluation index includes the correlation coefficient between remote sensing precipitation data and actual precipitation data. S30, Construct a deep neural network model, using the accuracy evaluation index and environmental auxiliary variable data as input features, perform model training and cross-validation, and use interpretability analysis methods to reveal feature contributions, so as to obtain a trained deep neural network model. S40 uses a trained deep neural network model to perform accuracy prediction on satellite remote sensing precipitation data in areas without monitoring stations, and outputs the accuracy prediction results.

[0009] The method for predicting the accuracy of remote sensing precipitation data in areas without monitoring stations provided in this application has the following advantages: By combining the multiplicative triple combination analysis method, the extended dual-tool variable method, and the verification method based on actual monitoring stations for comprehensive accuracy evaluation, it can effectively overcome the bias of a single method and provide more comprehensive and objective accuracy information. Simultaneously, by using multiple accuracy evaluation indicators and environmental auxiliary variable data as input features of the deep neural network model, the complementary information of multi-source data can be fully utilized, enhancing the model's adaptability to complex environmental factors. Furthermore, the introduction of interpretability analysis methods to reveal feature contributions can improve the transparency of the model's prediction process and the credibility of the results, ensuring reliable prediction in areas without monitoring stations. This provides a comprehensive, objective, and multi-scale global satellite precipitation accuracy analysis, effectively revealing common problems and regional differences in remote sensing precipitation accuracy, and providing a scientific basis for the selection of precipitation data, optimization of satellite data, and the deployment of precipitation monitoring stations.

[0010] As a further preferred embodiment, in step S10, the step of acquiring multi-source satellite remote sensing precipitation data and corresponding environmental auxiliary variable data also includes acquiring actual station-observed precipitation data, reanalyzed precipitation data, and precipitation data based on soil moisture inversion. The auxiliary environmental variable data include digital elevation model, drought index, leaf area index, soil texture, land use data, and climate zone classification data; The preprocessing includes time alignment, spatial resampling, and projection uniformity processing.

[0011] As a further preferred embodiment, in step S20, the correlation coefficient in the multiplicative triple combination analysis method... The calculation formula is:

[0012] In the formula, These represent three datasets that satisfy the assumptions of triple collocation analysis; Represented as a dataset Compared with the true value The correlation coefficient between them; This represents the covariance between data points.

[0013] As a further preferred embodiment, in step S20, in the extended dual-instrumental variable method, the correlation coefficient... The calculation formula is:

[0014] In the formula, This represents the error term of dataset X. Indicates variance; P This represents the actual precipitation value.

[0015] As a further preferred embodiment, in step S20, in the verification method based on actual test sites, the correlation coefficient... The calculation formula is:

[0016] In the formula, Represents the number of sample logs; This indicates the measured precipitation. This indicates the amount of precipitation sensed remotely.

[0017] As a further preferred embodiment, in step S30, the deep neural network model adopts a fully connected structure, including an input layer, an intermediate layer and an output layer. The nodes of the input layer correspond to multi-source features, the intermediate layer contains hidden layers and incorporates an attention mechanism, and the output layer corresponds to the accuracy prediction result. Furthermore, the deep neural network model is trained using the Adam optimizer, and five-fold cross-validation is used to reduce the risk of overfitting.

[0018] As a further preferred embodiment, the SHAP method is added to the deep neural network model to perform interpretability analysis on the results, and the Shapley value of each feature variable is calculated. The Shapley value represents the contribution of the feature variable to the model output.

[0019] As a further preferred embodiment, the formula for calculating the SHapley value is:

[0020] In the formula, Representation of features Shapley value; Indicates the number of features; Except The set of all features other than; express A feature set in a; A function representing the prediction model; The model's predicted output is represented as the sum of the baseline value and the contributions of each feature, calculated using the following formula:

[0021] In the formula, , where 1 indicates that the corresponding feature was selected in the subset; This represents the average predicted value of all observations; This represents the model's predicted value.

[0022] As a further preferred option, in step S40, a dense precipitation station area is selected, and its 0.25° large grid area is divided into a 1km×1km small grid. A deep neural network model is applied to obtain the small grid accuracy prediction results. Then, the prediction results are used to generate a spatial accuracy distribution map, enabling visual analysis and intuitive evaluation of the accuracy of remote sensing precipitation data in areas without monitoring stations.

[0023] Secondly, this application provides a system for implementing the method for predicting the accuracy of remote sensing precipitation data in non-site areas as described in any one of the above-mentioned methods, comprising: The data acquisition and preprocessing unit is used to acquire multi-source satellite remote sensing precipitation data and corresponding environmental auxiliary variable data, and to preprocess the multi-source satellite remote sensing precipitation data and environmental auxiliary variable data to make all data have a unified spatiotemporal range, spatiotemporal resolution and projection coordinate system. The accuracy assessment unit is used to assess the accuracy of multi-source satellite remote sensing precipitation data based on the multiplicative triple combination analysis method, the extended dual-tool variable method, and the verification method based on actual measurement stations, and to obtain the corresponding accuracy assessment index; the accuracy assessment index includes the correlation coefficient between remote sensing precipitation data and actual precipitation data; The model training unit is used to construct a deep neural network model. It takes the accuracy evaluation index and environmental auxiliary variable data as input features, performs model training and cross-validation, and uses interpretability analysis methods to reveal feature contributions in order to obtain a trained deep neural network model. The prediction unit is used to perform accuracy prediction on satellite remote sensing precipitation data in areas without monitoring stations using a trained deep neural network model, and outputs the accuracy prediction results.

[0024] It is understandable that the beneficial effects of the second aspect mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0025] Figure 1 This is a flowchart of the method for predicting the accuracy of remote sensing precipitation data in areas without monitoring stations, as provided in this application. Figure 2 This is a box plot of the global correlation coefficient R calculated by the MTC method provided in the embodiments of this application, which shows the accuracy of various remote sensing precipitation data; Figure 3 This is a box plot of the global correlation coefficient R calculated using the EIVD method provided in the embodiments of this application; Figure 4 This is a box plot comparing three methods—MTC, EIVD, and In-situ—provided in the embodiments of this application. Figure 5 This is a structural diagram of the DNN model provided in the embodiments of this application; Figure 6 This is a linear fit graph between the DNN model prediction results and the actual values ​​provided in the embodiments of this application; Figure 7 This is a ranking diagram of the importance of SHAP features in the DNN model provided in this application embodiment, showing the degree of contribution of each feature variable to the model output; Figure 8 This is a diagram showing the accuracy prediction results in a dense site grid area provided in an embodiment of this application. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0027] like Figure 1 As shown, this application provides a method for predicting the accuracy of remote sensing precipitation data in areas without monitoring stations, including steps S10 to S40, which are detailed below: Step S10: Acquire multi-source satellite remote sensing precipitation data and corresponding environmental auxiliary variable data, and preprocess the multi-source satellite remote sensing precipitation data and environmental auxiliary variable data to ensure that all data have a unified spatiotemporal range, spatiotemporal resolution and projection coordinate system.

[0028] In this application, step S10 eliminates the spatiotemporal inconsistencies between multi-source data by unifying the data format and scale, providing a consistent data foundation for subsequent accuracy assessment and model building. Specifically, it ensures that different data sources can be compared and analyzed under the same conditions.

[0029] Step S20: Based on the multiplicative triple combination analysis method, the extended dual-tool variable method, and the verification method based on actual measurement stations, the accuracy of multi-source satellite remote sensing precipitation data is evaluated to obtain the corresponding accuracy evaluation index; the accuracy evaluation index includes the correlation coefficient between remote sensing precipitation data and actual precipitation data.

[0030] In this application, step S20 integrates multiple independent accuracy assessment methods to comprehensively reflect the accuracy performance of precipitation data under different conditions. Specifically, it can overcome the bias that may exist in a single method, provide more comprehensive accuracy assessment results, and provide multi-dimensional features for model input.

[0031] Step S30: Construct a deep neural network model, using accuracy evaluation metrics and environmental auxiliary variable data as input features, perform model training and cross-validation, and use interpretability analysis methods to reveal feature contributions in order to obtain a trained deep neural network model.

[0032] In this application, step S30, by fusing multi-source features and interpretability analysis, can train a model capable of capturing complex nonlinear relationships, while clarifying the impact of each feature on the prediction results. Specifically, it can enhance the model's generalization ability and credibility in areas without sites, and avoid the uncertainty of black-box models.

[0033] Step S40: Using the trained deep neural network model, perform accuracy prediction on satellite remote sensing precipitation data in areas without monitoring stations, and output the accuracy prediction results.

[0034] In this application, step S40 can directly estimate the accuracy of precipitation data in areas without stations by applying the trained model. Specifically, it can generate a spatial accuracy distribution map, support the visualization analysis and scientific evaluation of remote sensing precipitation data, and provide a basis for data selection and station deployment.

[0035] The method for predicting the accuracy of remote sensing precipitation data in areas without monitoring stations provided in this application has the following advantages: By combining the multiplicative triple combination analysis method, the extended dual-tool variable method, and the verification method based on actual monitoring stations for comprehensive accuracy evaluation, it can effectively overcome the bias of a single method and provide more comprehensive and objective accuracy information. Simultaneously, by using multiple accuracy evaluation indicators and environmental auxiliary variable data as input features of the deep neural network model, the complementary information of multi-source data can be fully utilized, enhancing the model's adaptability to complex environmental factors. Furthermore, the introduction of interpretability analysis methods to reveal feature contributions can improve the transparency of the model's prediction process and the credibility of the results, ensuring reliable prediction in areas without monitoring stations. This provides a comprehensive, objective, and multi-scale global satellite precipitation accuracy analysis, effectively revealing common problems and regional differences in remote sensing precipitation accuracy, and providing a scientific basis for the selection of precipitation data, optimization of satellite data, and the deployment of precipitation monitoring stations.

[0036] In one embodiment, the technical solution to achieve the above objective can be as follows: This embodiment provides a method for predicting the accuracy of remote sensing precipitation data in areas without monitoring stations, including the following steps: The first step is data acquisition and preprocessing, including: Acquire mainstream global multi-source remote sensing precipitation data (IMERG, GSMAP, CMORPH, CHIRPS, PERSIANN-CDR), meteorological station data (GSOD), reanalysis data (ERA5), and soil moisture inversion precipitation data (SM2RAIN). Acquire environmental auxiliary data, including digital elevation model (DEM), drought index (AI), leaf area index (LAI), soil texture (ST), land use data, and climate zone classification data (Köppen classification), etc. All data were preprocessed uniformly with a grid scale of 0.25°, and the consistency of spatiotemporal range, spatiotemporal resolution, and projected coordinate system of all data was ensured.

[0037] Secondly, a comprehensive accuracy assessment method was used to analyze the accuracy of remote sensing precipitation data, including: The evaluation is based on the Multiplicative Triple Combination Analysis (MTC) method, which uses three independent data sources to estimate the error variance and precision. The R-value is calculated as follows:

[0038] In the formula This represents three datasets that satisfy the TC hypothesis. Represented as a dataset Compared with the true value The correlation coefficient between them This represents the covariance between data points.

[0039] The evaluation was performed using the Extended Bivariate Instrumental Method (EIVD), which estimates the error between the two datasets and the instrumental variables while considering common-mode error. The R-value was calculated as follows:

[0040] In the formula, For the error term, For variance, This represents the covariance of the data.

[0041] The evaluation was conducted using an in-situ validation method based on actual observations at weather stations. The correlation coefficient was calculated, and the R-value was calculated as follows:

[0042] In the formula For the sample logarithm, This is the measured precipitation (station data). For remotely sensed precipitation (satellite data or model data); By combining three accuracy assessment methods, we conduct accuracy analysis on various global satellite remote sensing precipitation data within the same framework. We evaluate the precipitation accuracy of each dataset at different dimensions, including global scale, different climate types (such as tropical, subtropical, and arid zones), and different land use types (such as forest, urban, farmland, and wasteland). We also compare the accuracy differences and regional differences among different precipitation data.

[0043] Furthermore, interpretable machine learning modeling is performed, including: The spatial representativeness of the auxiliary variable data is calculated and subsequently used as one of the features input into the model. Taking elevation as an example, we first divide a large grid cell (0.25° × 0.25°) into smaller grid cells (1km × 1km) (one large grid contains approximately 900 smaller grid cells), and calculate the station's location within the smaller grid cells and its location within the large grid cell based on latitude and longitude. Then, we obtain the elevation values ​​of the smaller grid cells containing the station and the elevation values ​​of all smaller grid cells within the large grid. These approximately 900 elevation values ​​are sorted from smallest to largest and divided into twenty categories based on quartiles. The number of smaller grid cells within the large grid cell that belong to the same elevation category as the station (let's call it X) is counted. The ratio of X to the total number of elevation grid cells (let's call it Y) is the representativeness of the elevation (representativeness = X / Y).

[0044] By integrating the three accuracy assessment results calculated above, as well as various auxiliary data, including elevation, drought index, leaf area index, soil texture, land cover type data, and their spatial representativeness, as input features of the machine learning model, the prediction accuracy is enhanced.

[0045] The deep neural network (DNN) model used in machine learning adopts a fully connected structure. The input layer nodes correspond to multi-source features, and the middle layer includes hidden layers and an attention mechanism. The output layer corresponds to the accuracy prediction result. The model is trained using the Adam optimizer and uses five-fold cross-validation to reduce the risk of overfitting.

[0046] The SHAP method is incorporated into the model to provide interpretability of the results, clarify the contribution of each feature variable to the model output, and reveal the impact of feature factors on prediction accuracy.

[0047] Finally, accuracy prediction is performed in areas without stations, including: A densely populated grid area was selected, and its large 0.25° grid area was divided into small grids of 1km×1km. The trained DNN model was then applied to obtain the prediction results of the small grid accuracy.

[0048] By using the prediction results to generate a spatial accuracy distribution map, we can achieve visual analysis and intuitive evaluation of the accuracy of remote sensing precipitation data in areas without monitoring stations.

[0049] The beneficial effects of the technical solution provided in this embodiment are: (1) By combining three accuracy assessment methods, the bias of a single method is overcome, the reliability of accuracy assessment is improved, and a comprehensive, objective, and multi-scale global satellite precipitation accuracy analysis is provided, revealing its common problems and regional characteristics. The significance lies in providing data users with key quality reference and pointing out the direction for improvement for data product developers.

[0050] (2) Establish accuracy assessment and accuracy prediction modeling driven by surface factors to achieve accuracy prediction of remote sensing precipitation data in areas without ground observation stations. It is interpretable and can reveal the contribution relationship of factors such as topography, climate, and land use to the accuracy of remote sensing precipitation, thereby improving the credibility of the model results.

[0051] The following is a specific implementation example of this application: This embodiment provides an interpretable machine learning model for predicting the accuracy of remote sensing precipitation data in areas without monitoring stations. It also integrates accuracy evaluation results from multiple methods to improve prediction accuracy. Specifically, it includes: 1. The correlation coefficient R value is obtained based on three accuracy evaluation methods, including MTC-R, EIVD-R and In-situ-R, and the specific calculation is shown in Equation (1), Equation (2) and Equation (3).

[0052] Equation (1) is the calculation method of MTC-R. MTC estimates the correlation between each data point and the true value through covariance matrix decomposition. Equation (2) is the calculation method of EIVD-R. The EIVD method assumes that the error of the instrumental variable is unrelated to other data by selecting instrumental variables and uses the covariance structure to estimate the true correlation.

[0053] Equation (3) is an evaluation method based on the actual measurement site (In-situ), which uses meteorological station observations as the benchmark to calculate the correlation coefficient.

[0054] 2. Feature construction of the DNN model, specifically including MTC-R, EIVD-R, and multi-source auxiliary variable data as feature inputs, with a uniform data resolution of 0.25°, and auxiliary variable data with a uniform resolution of 1km, located in each 0.25° grid, constructing feature vectors:

[0055] In the formula The spatial representativeness of the variable is represented by In-situ-R, which serves as the reference data for the model, treating it as the true numerical values ​​used for model training.

[0056] 3. Construction of an interpretable DNN model, specifically as follows:

[0057] In the formula This represents the model's prediction results. Representing each characteristic variable, For activation function, The weights represent the attention mechanism added to the DNN structure, which enhances the focus on key features.

[0058] Including SHAP analysis, the SHapley value is calculated as follows:

[0059] Representation of features Shapley value, It is the number of features. Except The set of all features except those mentioned above. yes A feature set, It is a function of the prediction model.

[0060] The model's predicted output can be represented as the sum of the baseline value and the contributions of each feature:

[0061] in , where 1 indicates that the corresponding feature was selected in the subset. It is the average predicted value of all observations. These are model predictions.

[0062] In one specific embodiment of this application, multiple accuracy assessment methods are applied to a remote sensing precipitation dataset, and the model is applied to a specific prediction experiment. The study area selected for this experiment is a global region, and the experimental data is shown in Table 1. The data preprocessing is standardized to the same spatial resolution of 0.25°, the same coordinate system WGS84, and the same spatial range when applying the assessment methods, which needs to be determined according to the actual data.

[0063] Table 1 Precipitation data

[0064] In this embodiment, the accuracy of five types of remote sensing precipitation data (IMERG, GSMAP, CMORPH, CHIRPS, and PERSIANN-CDR) was evaluated. The ERA5 reanalysis data, SM2RAIN soil moisture precipitation data, and remote sensing precipitation data were independent of each other, meeting the assumptions of TC. Therefore, ERA5, SM2RAIN, and any remote sensing data were used to form a triplet for calculation in the MTC method to obtain the correlation coefficient of the remote sensing data with the unknown true value. A total of eight triplets were formed for separate calculations. In EIVD, SM2RAIN and two remote sensing data were used as triplets for calculation, with data from the same series not being selected repeatedly. The correlation coefficient of the remote sensing data with the unknown true value was calculated similarly. The site-verified evaluation method directly matched the GSOD site precipitation data with each remote sensing data point to calculate the corresponding correlation coefficient. The results are as follows... Figure 2 and Figure 3 As shown, Figure 2 Box plot of the global correlation coefficient R calculated by the MTC method. Figure 3 Box plot of the global correlation coefficient R calculated using the EIVD method. Overall, the results are quite similar, with IMERG and GSMAP products showing superior performance.

[0065] In this embodiment, as Figure 4 As shown, Figure 4 This is a box plot comparing three methods: MTC, EIVD, and In-situ. MTC-R has the highest median and mean in most datasets, indicating that this method performs well in assessing data accuracy. In-situ-R more closely reflects actual observations but is susceptible to errors from individual stations or insufficient representativeness. EIVD-R values ​​typically fall between the first two, exhibiting some stability. EIVD-R has a more concentrated distribution and relatively neutral performance, closely resembling In-situ-R. Therefore, on a global scale, EIVD can be considered more suitable for accuracy assessment.

[0066] Similarly, the accuracy assessment methods for different precipitation data can be analyzed using climate type classification data and land use type classification data to compare the accuracy differences and regional variations. The GSMAP_mvkg and IMERG_final datasets show excellent overall performance, performing well across different climate zones and land use types. In subtropical climates, most datasets perform best, while most data perform poorly in polar climates. Forests and farmlands also perform best in most datasets, while wastelands perform poorly. In areas without monitoring stations, EIVD is a more robust accuracy assessment method; if high-quality monitoring stations are available, it should be combined for comparative analysis.

[0067] In this embodiment, as Figure 5 As shown, Figure 5 This is a diagram of the deep neural network (DNN) model structure in this embodiment. It adopts a fully connected structure, with input layer nodes corresponding to multi-source features, hidden layers in the middle and an attention mechanism added, and the output layer corresponding to the accuracy prediction result.

[0068]

[0069] In the formula This represents the model's prediction results. Representing each characteristic variable, For activation function, Represents weight.

[0070] To more intuitively evaluate the model's predictive performance, the metrics Ri and Rv are calculated. 2 The analysis of RMSE and MAE is performed using the following calculation method: R 2 The formulas for calculating RMSE, MAE, and R are as follows: (9) (10) (11) (12) In the formula: and These are the measured values ​​and the model predicted values, respectively. and R represents the average values ​​of the measured samples and the model-predicted samples, respectively. 2 The closer the R-value is to 1, the better the model's prediction results. The closer the RMSE and MAE are to 0, the better the model's prediction results.

[0071] In this embodiment, as Figure 6 As shown, Figure 6 The graph shows the linear fitting results between the DNN model's predictions and the actual values. The model constructed in this embodiment demonstrates better performance compared to the fitting results of TC-R, EIVD-R, and In-situ-R. The results are shown in the figure. 2 The results are mostly above 0.65, indicating that the model can explain more than 65% of the result variance. The R value is mostly above 0.8, indicating good fitting effect. This shows that DNN can effectively help understand the accuracy differences between MTC, EIVD and site verification methods.

[0072] In this embodiment, as Figure 7 As shown, Figure 7This is a ranking of the importance of SHAP features in the DNN model. Overall, it can be found that EIVD-R, soil and the correlation with the dataset R(-ERA5) and R(-SM2RAIN) are important features that contribute significantly to the model's output, especially the accuracy results calculated in the EIVD method. It can be seen that EIVD-R has a greater impact on In-situ-R, which also indicates that changes in EIVD-R are more related to In-situ-R.

[0073] In this embodiment, a densely populated grid area is selected for accuracy prediction. The selected region is located between longitude -63.75° and -63.5°, and latitude 44.5° and 44.75°, containing nine GSOD precipitation stations. The selected grid has a resolution of 0.25°, with each small grid having a resolution of 1km (one large grid contains 900 small grids). The input variables for each small grid within this region differ at 1km resolution but are the same at 0.25° resolution. Based on the preceding DNN model, the in-situ-R of each small grid is predicted, yielding the accuracy prediction result for the grid region.

[0074] like Figure 8 As shown, Figure 8 This is a map showing the accuracy prediction results in a densely gridded area of ​​precipitation data. Yellow triangles indicate the locations of GSOD stations within that grid area. The color gradient from white to dark blue represents the change in precipitation data accuracy from low to high; higher values ​​indicate higher remote sensing data accuracy and also higher spatial representativeness of the stations. This high-resolution spatial prediction map of precipitation accuracy provides more reliable accuracy assessment results for satellite precipitation data. This can provide data users with predictive quality references (such as avoiding low-accuracy areas) and a scientific basis for the selection and correction of satellite data in complex regions. Furthermore, high prediction values ​​also indicate that the point scale can more effectively represent the precipitation values ​​at the current grid scale. In other words, in addition to providing more reliable data accuracy assessment results, it can also provide reference opinions for deploying precipitation stations in highly representative locations.

[0075] In summary, the interpretable machine learning model proposed in this embodiment, by integrating three accuracy assessment methods and multi-source auxiliary environmental data, can predict the accuracy of remote sensing precipitation data in areas without monitoring stations. The model not only fully utilizes the features of existing data but also reveals the contribution of each feature to the prediction results through attention mechanisms and SHAP analysis, enhancing the model's interpretability and reliability. Experimental results show that the DNN model performs excellently in simulating global and local precipitation accuracy, effectively reflecting the accuracy differences between different datasets. Through small-grid accuracy prediction, the generated spatialized accuracy distribution map can provide users with a reference for avoiding low-accuracy areas and for the deployment of precipitation observation stations, achieving an intuitive, scientific, and operable assessment of the accuracy of precipitation data in areas without monitoring stations.

[0076] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for predicting the accuracy of remote sensing precipitation data in areas without monitoring stations, characterized in that, Includes the following steps: S10: Acquire multi-source satellite remote sensing precipitation data and corresponding environmental auxiliary variable data, and preprocess the multi-source satellite remote sensing precipitation data and environmental auxiliary variable data to make all data have a unified spatiotemporal range, spatiotemporal resolution and projection coordinate system. S20. Based on the multiplicative triple combination analysis method, the extended dual-tool variable method, and the verification method based on actual measurement stations, the accuracy of multi-source satellite remote sensing precipitation data is evaluated to obtain the corresponding accuracy evaluation index; the accuracy evaluation index includes the correlation coefficient between remote sensing precipitation data and actual precipitation data. S30, Construct a deep neural network model, using the accuracy evaluation index and environmental auxiliary variable data as input features, perform model training and cross-validation, and use interpretability analysis methods to reveal feature contributions, so as to obtain a trained deep neural network model. S40 uses a trained deep neural network model to perform accuracy prediction on satellite remote sensing precipitation data in areas without monitoring stations, and outputs the accuracy prediction results.

2. The method for predicting the accuracy of remote sensing precipitation data in areas without monitoring stations as described in claim 1, characterized in that, In step S10, the step of acquiring multi-source satellite remote sensing precipitation data and corresponding environmental auxiliary variable data also includes acquiring actual station observation precipitation data, reanalyzing precipitation data and precipitation data based on soil moisture inversion; The auxiliary environmental variable data include digital elevation model, drought index, leaf area index, soil texture, land use data, and climate zone classification data; The preprocessing includes time alignment, spatial resampling, and projection uniformity processing.

3. The method for predicting the accuracy of remote sensing precipitation data in areas without monitoring stations as described in claim 1, characterized in that, In step S20, in the multiplicative triple collocation analysis method, the correlation coefficient... The calculation formula is: In the formula, These represent three datasets that satisfy the assumptions of triple collocation analysis; Represented as a dataset Compared with the true value The correlation coefficient between them; This represents the covariance between data points.

4. The method for predicting the accuracy of remote sensing precipitation data in areas without monitoring stations as described in claim 1, characterized in that, In step S20, in the extended dual-instrument variable method, the correlation coefficient The calculation formula is: In the formula, This represents the error term of dataset X. Indicates variance; P This represents the actual precipitation value.

5. The method for predicting the accuracy of remote sensing precipitation data in areas without monitoring stations as described in claim 1, characterized in that, In step S20, in the verification method based on actual test sites, the correlation coefficient The calculation formula is: In the formula, Represents the number of sample logs; This indicates the measured precipitation. This indicates the amount of precipitation sensed remotely.

6. The method for predicting the accuracy of remote sensing precipitation data in areas without monitoring stations as described in claim 1, characterized in that, In step S30, the deep neural network model adopts a fully connected structure, including an input layer, an intermediate layer and an output layer. The nodes of the input layer correspond to multi-source features, the intermediate layer contains hidden layers and incorporates an attention mechanism, and the output layer corresponds to the accuracy prediction result. Furthermore, the deep neural network model is trained using the Adam optimizer, and five-fold cross-validation is used to reduce the risk of overfitting.

7. The method for predicting the accuracy of remote sensing precipitation data in areas without monitoring stations as described in claim 1, characterized in that, The SHAP method is added to the deep neural network model to perform interpretability analysis of the results, and the Shapley value of each feature variable is calculated. The Shapley value represents the contribution of the feature variable to the model output.

8. The method for predicting the accuracy of remote sensing precipitation data in areas without monitoring stations as described in claim 7, characterized in that, The formula for calculating the SHapley value is as follows: In the formula, Representation of features Shapley value; Indicates the number of features; Except The set of all features other than; express A feature set in a; A function representing the prediction model; The model's predicted output is represented as the sum of the baseline value and the contributions of each feature, calculated using the following formula: In the formula, , where 1 indicates that the corresponding feature was selected in the subset; This represents the average predicted value of all observations; This represents the model's predicted value.

9. The method for predicting the accuracy of remote sensing precipitation data in areas without monitoring stations as described in claim 1, characterized in that, In step S40, a dense precipitation station area is selected, and its 0.25° large grid area is divided into a 1km×1km small grid. A deep neural network model is applied to obtain the small grid accuracy prediction results. Then, the prediction results are used to generate a spatial accuracy distribution map, enabling visual analysis and intuitive evaluation of the accuracy of remote sensing precipitation data in areas without monitoring stations.

10. A system for achieving the accuracy of remote sensing precipitation data for predicting areas without monitoring stations as described in any one of claims 1 to 9, characterized in that, include: The data acquisition and preprocessing unit is used to acquire multi-source satellite remote sensing precipitation data and corresponding environmental auxiliary variable data, and to preprocess the multi-source satellite remote sensing precipitation data and environmental auxiliary variable data to make all data have a unified spatiotemporal range, spatiotemporal resolution and projection coordinate system. The accuracy assessment unit is used to assess the accuracy of multi-source satellite remote sensing precipitation data based on the multiplicative triple combination analysis method, the extended dual-tool variable method, and the verification method based on actual measurement stations, and to obtain the corresponding accuracy assessment index; the accuracy assessment index includes the correlation coefficient between remote sensing precipitation data and actual precipitation data; The model training unit is used to construct a deep neural network model. It takes the accuracy evaluation index and environmental auxiliary variable data as input features, performs model training and cross-validation, and uses interpretability analysis methods to reveal feature contributions in order to obtain a trained deep neural network model. The prediction unit is used to perform accuracy prediction on satellite remote sensing precipitation data in areas without monitoring stations using a trained deep neural network model, and outputs the accuracy prediction results.