Machine learning-based near-real-time satellite rainfall fusion method under sparse rainfall station
By constructing a near-real-time satellite precipitation fusion method based on machine learning, and using GSMAP and IMERG data combined with DEM, the model weight allocation was optimized, which solved the problems of high variability and spatial complexity of precipitation data from sparse rain gauge stations, and achieved higher accuracy precipitation prediction.
Patent Information
- Application Number
- CN202510169064.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-11-18
AI Technical Summary
Existing precipitation data merging methods cannot accurately predict precipitation data when rain gauge stations are sparse, especially in areas with complex terrain and extreme climates. Traditional methods struggle to handle the high variability and spatial complexity of precipitation data, limiting prediction accuracy and applicability.
A near-real-time satellite precipitation fusion method based on machine learning was adopted. The near-real-time satellite precipitation products GSMAP and IMERG were combined with geospatial DEM data. Artificial neural network, random forest, long short-term memory network and Transformer model were constructed, and precipitation level weight adjustment module was embedded to optimize the weight allocation of the model under different precipitation levels. Cross-validation and accuracy evaluation were carried out to select the best fused data.
It improves the accuracy of precipitation data at sparse rainfall stations, especially the prediction accuracy at different precipitation levels and spatial scales, and enhances the reliability and precision of precipitation event observation.
Smart Images

Figure CN120974394A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of meteorological precipitation, and in particular to a near real-time satellite precipitation fusion method based on machine learning in sparse rain gauge stations. BACKGROUND
[0002] Current precipitation data acquisition methods include rain gauges, ground weather radars and satellite precipitation products. Rain gauge networks can obtain high-precision data, but are limited by terrain heterogeneity and cost, resulting in uneven and sparse distribution at global and regional scales. Ground weather radars obtain atmospheric precipitation information by analyzing radar echo signals; however, the detection range of this method is limited, and its accuracy is greatly affected by environmental and climatic conditions. In recent years, satellite precipitation products based on visible light, infrared (IR), passive / active microwave (PMW) and multi-sensor joint inversion have been widely used in precipitation assessment. Commonly used satellite remote sensing-based precipitation data include Integrated Multi-satellite Retrievals for GPM (IMERG), Tropical Rainfall Measuring Mission (TRMM) Multi-satellite Precipitation Analysis (TMPA), Climate Hazards Group Infrared Precipitation (CHIRPS), Precipitation Estimation from Remotely Sensed Information using Artificial Neural Networks (PERSIANN) and PERSIANN-Climate Data Record (PERSIANN-CDR), etc. The most important feature of the above satellite precipitation products is global spatial coverage, but the accuracy is affected by atmospheric conditions and inversion algorithms.
[0003] Currently, appropriate algorithms have been established to combine data sets of different satellite precipitation products, including inverse distance weighted interpolation, spatial fusion, linear regression, geographically weighted regression, etc. Although the above precipitation data merging methods have application value under certain conditions, they have significant shortcomings in regions with complex terrain and extreme climate. Inverse distance weighted interpolation is not effective when data is unevenly distributed, spatial fusion methods ignore the interactive effects of terrain and climate, linear regression cannot capture non-linear and complex spatiotemporal relationships, and geographically weighted regression, although it considers spatial heterogeneity, still cannot accurately reflect the complex changes of precipitation in high-altitude areas. Therefore, these traditional methods are difficult to effectively handle the high variability and spatial complexity of precipitation data in areas with sparse rain gauge stations, especially in the case of sparse rain gauge stations, which limits their prediction accuracy and applicability. SUMMARY
[0004] In order to solve the problem that the existing precipitation data merging method cannot accurately predict the precipitation data in the case of sparse rain stations, the present application proposes a near real-time satellite precipitation fusion method based on machine learning under sparse rain stations, which uses near real-time satellite precipitation products GSMaP and IMERG under the condition of sparse rain stations, combines geographic spatial DEM data as input, uses four machine learning algorithms to construct a precipitation fusion model, evaluates the fused precipitation from the spatial and precipitation grade angles, optimizes the best fused precipitation data, and improves the precision of the precipitation data to solve the above problems.
[0005] The present application discloses a near real-time satellite precipitation fusion method based on machine learning under sparse rain stations, comprising the following steps:
[0006] S1, obtaining near real-time satellite precipitation data, DEM data and sparse rain station measured precipitation data;
[0007] S2, extracting real-time satellite precipitation data and DEM data at the scale of sparse rain stations, and establishing a plurality of data input combinations;
[0008] S3, constructing a precipitation fusion model, the precipitation fusion model comprising an artificial neural network, a random forest, a long short-term memory network and a Transformer, and embedding a precipitation grade weight adjustment module in all precipitation fusion models to adjust the weight of different precipitation grades in the model;
[0009] S4, using the plurality of data input combinations obtained in S2 as the input of the precipitation fusion model, using the sparse rain station measured precipitation data as the output, dividing the training set and the validation set through cross-validation, and training the precipitation fusion model;
[0010] S5, constructing a spatial fusion precipitation accuracy evaluation system and a fusion precipitation intensity accuracy evaluation system, comprehensively evaluating the accuracy of the fused precipitation of each precipitation fusion model, and obtaining the best fused precipitation data under the condition of sparse rain stations.
[0011] Preferably, the near real-time satellite precipitation data includes GSMaP GNRT daily scale data and IMERG Early daily scale data, and the spatial resolution is 0.1°.
[0012] Preferably, the plurality of data input combinations includes GSMaP+IMERG+DEM, GSMaP+DEM, IMERG+DEM and GSMaP+IMERG.
[0013] Preferably, the precipitation grade weight adjustment module assigns different weights to samples of different precipitation grades through a weighted mean square error loss function, specifically:
[0014] When the precipitation is greater than or equal to 0.1mm / day, the weight value is 5.0.
[0015] When the precipitation amount is <0.1 mm / day, the weight value is assigned as 1.0.
[0016] Preferably, the precipitation level weight adjustment module embedded in all precipitation fusion models comprises:
[0017] A weighted mean square error loss function is embedded in the loss function calculation stage of the artificial neural network. In each training iteration, the artificial neural network assigns weights according to the specific conditions of the target value, which is used to calculate the weighted mean square error;
[0018] A weighted mean square error loss function is embedded in the calculation of the random forest decision tree node splitting standard. Each time the node splits, the weighted mean square error loss function dynamically adjusts the weight according to the target value, so as to more preferentially split the high-weight samples;
[0019] A weighted mean square error loss function is embedded in the loss function calculation stage of the long short-term memory network. For time series data, the weight of the sample is dynamically calculated and the loss value is adjusted each time the prediction value is updated by the time step;
[0020] A weighted mean square error loss function is embedded in the loss function calculation stage of the Transformer. By dynamically adjusting the weight, the loss function gives greater penalty value to strong precipitation events, and dynamically adjusts the attention distribution to ensure that the Transformer captures the global dependence of the sequence while guiding the self-attention mechanism to focus on key spatiotemporal features.
[0021] Preferably, the precipitation level is divided based on the quantile method as follows:
[0022] L1 level, precipitation value: 0.1 mm / day-0.7 mm / day;
[0023] L2 level, precipitation value: 0.7 mm / day-2.9 mm / day;
[0024] L3 level, precipitation value: 2.9 mm / day-10.8 mm / day;
[0025] L4 level, precipitation value: > 10.8 mm / day.
[0026] Preferably, the cross-validation adopts 3-fold cross-validation, including the following steps:
[0027] The measured precipitation data of the sparse rainfall station is randomly divided into three folds, and each time two folds are selected as the training set and the remaining one fold is selected as the validation set;
[0028] The training and validation are repeated three times to ensure that all rainfall stations participate in training and validation.
[0029] Preferably, S5 comprises the following steps:
[0030] By the spatial fusion precipitation accuracy evaluation system, the precipitation accuracy index of each precipitation fusion model at each rain gauge station is calculated, and the evaluation results of all rain gauge stations are counted to evaluate the overall performance of each precipitation fusion model in the spatial dimension;
[0031] By the intensity fusion precipitation accuracy evaluation system, the precipitation is classified according to intensity, the precipitation accuracy index of each grade is calculated and compared, and the ability of each precipitation fusion model to capture different precipitation intensities is evaluated.
[0032] Combined with the results of the spatial fusion precipitation accuracy evaluation system and the intensity fusion precipitation accuracy evaluation system, the weighted average is performed from the spatial accuracy and the intensity accuracy, and the final fusion precipitation data under the sparse rain gauge station is selected.
[0033] Preferably, the precipitation accuracy index includes a probabilistic index and a statistical index, the probabilistic index includes a detection rate and a false alarm rate, and the statistical index includes a correlation coefficient, a root mean square error and a relative deviation.
[0034] The beneficial effects of the present application are:
[0035] (1) Fusion precipitation accuracy under different input combinations: The traditional method generally uses multiple satellite precipitation data, and the redundancy of these data has certain doubts on the improvement of model accuracy. In the present application, three kinds of input data are combined as the input of the model, and the precipitation fusion is carried out to provide fusion precipitation data under multiple input combinations.
[0036] (2) Precipitation accuracy evaluation under different grades: There are few fusion precipitation accuracy evaluation systems considering different precipitation grades in the prior art. In the present application, the regional precipitation grade is divided, the accuracy analysis of the multi-source precipitation fusion data is completed, and the accuracy of the precipitation under different precipitation grades is improved.
[0037] (3) Precipitation accuracy evaluation under different spatial scales: The relatively sparse rain gauge station measured data is used as a reference, DEM is used as input data in precipitation fusion, the geographical spatial characteristics of different altitudes in the region are considered, precipitation fusion is carried out, and the reliability and accuracy of the observation of regional precipitation events are improved. BRIEF DESCRIPTION OF DRAWINGS
[0038] Figure 1 The present application is an embodiment of a machine learning-based near real-time satellite precipitation fusion method under sparse rain gauge stations;
[0039] Figure 2 The present application is an embodiment of a Taylor diagram of a different input combination model;
[0040] Figure 3 Box plots of precipitation detection indices fused by different input combination models in this embodiment of the invention;
[0041] Figure 4 This is a line graph showing the annual average precipitation of different precipitation fusion models in this invention.
[0042] Figure 5 Box plots of precipitation accuracy indices for different precipitation fusion models in embodiments of the present invention;
[0043] Figure 6 This is a spatial distribution map of precipitation accuracy indices for different precipitation fusion models in this invention embodiment;
[0044] Figure 7 This is a schematic diagram comparing the histograms of precipitation accuracy indices between the model and the rain gauge dataset at different altitudes according to an embodiment of the present invention.
[0045] Figure 8 Box plots showing precipitation accuracy indices for different models at different altitudes in this invention embodiment. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided with reference to the accompanying drawings and embodiments.
[0048] This application discloses a near-real-time satellite precipitation fusion method based on machine learning for sparse rain gauge stations, the process of which is as follows: Figure 1 As shown, it includes the following steps:
[0049] S1. Acquire near-real-time satellite precipitation data, DEM data, and sparse rain gauge measured precipitation data. Near-real-time satellite precipitation data includes GSMAP GNRT diurnal data and IMERG Early diurnal data, with a spatial resolution of 0.1°.
[0050] In this study, the Tibetan Plateau was used as the precipitation fusion research area, employing 78 monitoring stations over a time series from January 1, 2001 to December 31, 2022. The Tibetan Plateau covers approximately 2.5 million km², and based on the formula for calculating rainfall station density, each station on the Tibetan Plateau serves approximately 32,051 km². 2 This value is far greater than the standard recommended by the World Meteorological Organization (WMO) (i.e., 250 km). 2 and 575km 2 The distance between the two points indicates that the meteorological stations in this area are sparsely distributed, belonging to a region with sparse rainfall stations.
[0051] Global Satellite Mapping of Precipitation (GSMaP) adjusted by rain gauges is a hybrid microwave and infrared precipitation product of the GPM mission. GSMaP near real-time versions (GSMaP-NRT and GSMaP-Gauge-NRT) provide precipitation estimates at 0.1° spatial resolution and hourly temporal resolution, covering the latitudinal range from 60°N to 60°S. In this embodiment, the near real-time GSMaP GNRT dataset (version 06) is adopted, with the time range from January 2001 to December 2022.
[0052] IMERG integrates data from the Global Precipitation Measurement (GPM) Core Satellite (GPM CO) and various satellite sensors, providing global precipitation estimates at 0.1° resolution and 30-minute intervals, covering the latitudinal range from 90°N to 90°S. In this embodiment, the near real-time IMERG dataset is used, with the time range from January 2001 to December 2022, providing timely precipitation estimates for real-time applications.
[0053] SRTM Digital Elevation Model (DEM) data covers a wide range of global areas, including land and ocean areas. The DEM data used in this embodiment has a resolution of 90 meters.
[0054] The World Meteorological Organization (WMO) has certain recommended standards for the density of weather station networks, especially in terms of precipitation monitoring. Sparse rain gauge stations generally refer to those located in areas with low density of weather station networks. These areas may be difficult to establish more weather monitoring stations due to geographical reasons (such as mountainous areas, deserts, etc.), or due to economic or technical reasons.
[0055] S2, extract real-time satellite precipitation data at the scale of sparse rain gauge stations, DEM data, and establish multiple data input combinations.
[0056] Multiple data input combinations include GSMaP + IMERG + DEM (GID), GSMaP + DEM (GD), IMERG + DEM (ID), and GSMaP + IMERG (GI).
[0057] S3, construct a precipitation fusion model, the precipitation fusion model includes an artificial neural network (ANN), a random forest (RF), a long short-term memory network (LSTM), and a Transformer, and embed a precipitation grade weight adjustment module in all precipitation fusion models to adjust the weight of different precipitation grades in the model.
[0058] Because the spatial distribution of precipitation on the Tibetan Plateau is extremely uneven in this embodiment, with most daily precipitation concentrated between 0.1 and 5 mm, moderate to heavy rain, especially extreme precipitation events, are relatively rare in this region. This makes traditional precipitation classification methods less applicable. Therefore, determining appropriate thresholds for classifying precipitation levels in this region is a worthwhile issue to explore. The quantile method is a commonly used method in precipitation classification, particularly suitable for analyzing distribution characteristics and creating categories based on different precipitation levels. Given the unique precipitation patterns of the Tibetan Plateau, precipitation is divided into four levels, as shown in Table 1.
[0059] Table 1 Precipitation Classification Thresholds
[0060]
[0061] In this embodiment, the precipitation level weight adjustment module assigns differentiated weights to samples of different precipitation levels using a weighted mean squared error loss function (weighted_mse_loss), helping the model to pay more attention to heavy precipitation events during training. Regions with precipitation exceeding a threshold (0.1) are given a larger weight (5.0), while regions with low precipitation are given a smaller weight (1.0). That is, when precipitation ≥ 0.1 mm / day, a weight value of 5.0 is assigned; when precipitation < 0.1 mm / day, a weight value of 1.0 is assigned. This approach makes the precipitation fusion model more sensitive to heavy precipitation, optimizing its prediction performance in these key regions.
[0062] This weighting method is particularly suitable for situations where precipitation data is unevenly distributed, as precipitation amounts typically exhibit strong spatial and temporal heterogeneity. By reducing overfitting to low-precipitation areas, precipitation fusion models can better capture the changing trends of heavy precipitation events, thereby improving the overall accuracy of predictions, especially when dealing with extreme weather phenomena.
[0063] In this embodiment, the Artificial Neural Network (ANN) further optimizes its performance in precipitation data modeling by embedding a weighted mean squared error loss function (weighted_mse_loss) module. For areas with precipitation exceeding a threshold (0.1 mm), the loss function assigns larger weights (5.0) to these data, while low-precipitation areas are given smaller weights (1.0). During the ANN's training process, the weighted mean squared error loss function is embedded in the loss function calculation stage. In each training iteration, the ANN assigns weights based on specific conditions of the target value (such as precipitation intensity) to calculate the weighted mean squared error (MSE). The weight logic directly affects the error between the predicted and actual values, ensuring that the model is more sensitive to samples with heavy precipitation.
[0064] A weighted mean squared error (MSE) loss function is introduced into the training process of Random Forest (RF) to better adapt to the modeling needs of precipitation events of different intensities. By introducing precipitation level weights, the MSE loss function makes the RF model focus more on the prediction accuracy of areas with heavy precipitation during training. In the RF algorithm, the MSE loss function is incorporated into the calculation of the decision tree node splitting criteria. Each time a node splits, the loss function dynamically adjusts the weights based on the target value, thus favoring more accurate splits on high-weight samples. This adjustment continues throughout the growth process of all trees, thereby improving RF's ability to learn specific categories (heavy precipitation events).
[0065] Embedding a weighted mean squared error (MSE) loss function into a Long Short-Term Memory (LSTM) network assigns differentiated weights based on the intensity of precipitation events, allowing the LSTM to give greater attention to heavy precipitation events during training. In the LSTM model, the MSE loss function is applied during the loss function calculation phase. For time-series data, the sample weights are dynamically calculated and the loss value is adjusted each time the predicted value is updated with a time step. This weight allocation makes the model focus more on key features related to heavy precipitation in the time series. Furthermore, the introduction of the MSE loss function further optimizes the LSTM's ability to model the temporal dependence of heavy precipitation events in time series.
[0066] By embedding a weighted mean squared error (MSE) loss function into the Transformer model, and assigning higher weights to heavy precipitation events, the self-attention mechanism is guided to prioritize data points in areas of heavy precipitation. During training, the MSE loss function is integrated into the loss function calculation stage. The model dynamically adjusts the weights, causing the loss function to impose a larger penalty on heavy precipitation events. This dynamic adjustment of the attention distribution ensures that while capturing global dependencies in the sequence, the model indirectly guides the self-attention mechanism to focus on key spatiotemporal features, exhibiting higher sensitivity to key regions and extreme precipitation events.
[0067] S4. Using the various data input combinations obtained in S2 as input to the precipitation fusion model, and the measured precipitation data from sparse rain gauge stations as output, the model is trained by dividing the training and validation sets through cross-validation. Different precipitation data input combinations and precipitation fusion models are shown in Table 2.
[0068] Table 2. Different combinations of precipitation data input and precipitation fusion models
[0069]
[0070]
[0071] Cross-validation employs a 3-fold cross-validation method: the data from 78 rainfall stations is randomly divided into three folds (Fold 1, Fold 2, Fold 3), ensuring that each fold contains approximately one-third of the stations and that the station distribution is representative. In each iteration, data from two folds of stations is selected as the training set, and data from the remaining fold is used as the validation set. This training and validation process is repeated three times to ensure that all rainfall stations participate in both training and validation.
[0072] S5. Construct a spatial fusion precipitation accuracy evaluation system and a fusion precipitation intensity accuracy evaluation system to comprehensively evaluate the accuracy of precipitation fusion of various precipitation fusion models and obtain the best precipitation fusion data under the condition of sparse rain gauge stations.
[0073] After fusing the validation set data of each station output by the precipitation fusion model, unlike conventional precipitation evaluation indicators, this embodiment further establishes two system evaluation frameworks to comprehensively assess the accuracy of the fused precipitation: a spatial fused precipitation accuracy evaluation framework and a fused precipitation intensity accuracy evaluation framework. Through the spatial fused precipitation accuracy evaluation framework, the precipitation accuracy index of each precipitation fusion model at each rain gauge station is calculated, and the evaluation results of all rain gauge stations are statistically analyzed to assess the overall performance of each precipitation fusion model in the spatial dimension. Then, through the fused precipitation intensity accuracy evaluation framework, precipitation is classified according to intensity level (light rain, moderate rain, heavy rain, and torrential rain), and the precipitation accuracy index of each level is calculated and compared to evaluate the ability of each precipitation fusion model to capture different precipitation intensities. Combining the results of the spatial fused precipitation accuracy evaluation framework and the fused precipitation intensity accuracy evaluation framework, a weighted average is calculated from both spatial accuracy and precipitation intensity accuracy dimensions, and a comprehensive ranking is performed. Finally, the best precipitation fusion data is selected under the condition of sparse rain gauge stations.
[0074] Quantitative validation of the precipitation fusion model largely depends on carefully selected performance indices. In this embodiment, precipitation accuracy indices include probabilistic and statistical indices. Probabilistic indices include detection rate (POD), false alarm rate (FAR), and critical success index (CSI), while statistical indices include correlation coefficient (CC), root mean square error (RMSE), and relative bias (RB). Specific calculations are shown in Table 3.
[0075] Table 3 Precipitation Accuracy Indicators
[0076]
[0077] Where n is the sample size, S i For the i-th value of the satellite data, O i To verify the i-th value of the data, For S i The average value, For O i The average value is given by H, where H is the number of precipitation events correctly detected by satellite precipitation products (SPPs), M is the number of precipitation events missed by satellite precipitation products, and F is the number of false detections by satellite precipitation products.
[0078] In different spatial fusion precipitation accuracy evaluation systems, the principle of numerical selection is to assess the spatial accuracy of the precipitation fusion model by comprehensively evaluating the performance of indicators at different spatial locations. For high precipitation intensity areas within a region, a higher POD value (>0.7) indicates a more ideal performance, while the FAR value should be kept as low as possible (<0.2) to avoid too many false precipitation events. To comprehensively evaluate spatial fusion accuracy, a weighted approach is used to combine the POD, FAR, CC, RMSE, and RB values of each region, assigning different weights to different regions. The comprehensive score can be obtained through weighted averaging or fuzzy logic methods to ensure the consistency of various indicators, thereby obtaining a comprehensive evaluation of spatial precipitation accuracy.
[0079] In different fusion precipitation intensity accuracy evaluation systems, the main focus is on the ability and accuracy of the precipitation fusion model to capture heavy precipitation events (such as rainstorms and torrential rain). When assessing precipitation, the requirements for Point of Observation (POD) are relatively high, while the Expected Average Response (FAR) value should be as low as possible. The POD value should ideally be greater than 0.8, while the FAR should be kept below 0.2. To comprehensively evaluate precipitation intensity accuracy, separate scores need to be given for different precipitation intensity categories, and then the final evaluation result is synthesized using a weighted average method. Weights can be set for different precipitation intensities (light precipitation, moderate precipitation, heavy precipitation, and rainstorms), adjusting the weights according to the impact and frequency of these events. Heavy precipitation and rainstorm events can be given higher weights to emphasize the model's performance in these key events. The final comprehensive score will be obtained through a weighted average method to ensure that heavy precipitation events occupy an important position in the evaluation.
[0080] The comprehensive evaluation system combines spatial accuracy and intensity accuracy to fully assess the performance of the fused precipitation model under conditions of sparse rainfall stations. In the comprehensive evaluation of spatial and intensity accuracy, numerical selection and weighting are crucial in determining the final score. Regarding spatial accuracy, the combined scores of POD, FAR, CC, RMSE, and RB determine the performance of the fused precipitation model in different regions, while regarding intensity, the scores of POD, FAR, CC, RMSE, and RB focus more on the performance during heavy precipitation and extreme precipitation events.
[0081] Spatial accuracy and intensity accuracy are evaluated in groups, each categorized into different scoring dimensions. For each dimension, the evaluation criteria adjust the weights based on the characteristics of precipitation events and regional differences. For areas with a high density of rain gauges, the weight of spatial accuracy is used to enhance the assessment importance of these areas; for areas with frequent extreme precipitation events, the weight of intensity accuracy is used to enhance the assessment impact of heavy precipitation. The scores for spatial accuracy and intensity accuracy are combined according to a certain weighting ratio. By adjusting the weighting ratio of spatial accuracy and intensity accuracy, the evaluation process is further optimized to ensure that the accuracy of precipitation fusion models under different conditions is fully evaluated, providing a basis for the final selection of the optimal precipitation fusion method.
[0082] In one specific embodiment, the Tibetan Plateau was selected as the study area for precipitation fusion. Rainfall monitoring stations are sparsely distributed across the Tibetan Plateau, especially in the central and western regions where station density is relatively low. Stations are more concentrated on the edges of the Tibetan Plateau, particularly near cities and densely populated areas.
[0083] Four precipitation data sets—GSMaP+IMERG+DEM(GID), GSMAP+DEM(GD), IMERG+DEM(ID), and GSMAP+IMERG(GI)—were combined and input into four machine learning models. The results were plotted as a Taylor plot, as shown below. Figure 2 As shown, Figure 2 In the middle (a), the result of the ANN model is shown. Figure 2 In the middle (b), the results of the RF model are shown. Figure 2 (c) represents the result of the LSTM model. Figure 2 In the diagram (d), the results represent the Transformer model. Overall, all precipitation fusion models output precipitation data with indices superior to those of GSMAP and IMERG satellite precipitation products. There was little difference between precipitation fusion models with different input combinations.
[0084] Combined Taylor diagram ( Figure 2 ) and 3 precipitation event detection indicators ( Figure 3 From the perspective of the combined precipitation, the precipitation after fusion is significantly better than the original precipitation. In order to obtain a better precipitation fusion effect among different combinations, this embodiment selects the GID combination as the input of the precipitation fusion model.
[0085] Using GSMAP+IMERG+DEM(GID) as input to the precipitation fusion model and measured station observation data as output, the results are used for statistical analysis of annual average precipitation. The results are as follows: Figure 4 As shown. In Figure 4In the diagram, the gray lines represent station observations. GSMAP and IMERG data are located at the top and bottom, respectively. GSMAP significantly overestimates the annual precipitation on the Tibetan Plateau, while IMERG significantly underestimates it. Among the precipitation fusion models, the Transformer model slightly overestimates precipitation, while the other three models underestimate it. The Transformer and ANN models are closest to the actual station observations, followed by the RF model, while the LSTM model significantly underestimates the annual precipitation. Looking at the interannual precipitation trends, the GSMAP product's trend line fluctuates the most. In 2015, the trends of the GSMAP product, IMERG product, and RF model all showed a decline followed by an increase, while the actual observed values at this time point showed an increase followed by a decline. This characteristic was captured by the ANN, Transformer, and LSTM models. In terms of the overall trend, the annual average precipitation change trends of the ANN and Transformer models are closest to the actual observed values, while the LSTM model's change characteristics before and after the time points of 2008 and 2010 differ from the actual observed trend.
[0086] Using 0.1 mm / day as the threshold for daily precipitation events on the Tibetan Plateau, the precipitation capture capabilities of four precipitation fusion models, GSMAP products, and IMERG products were statistically analyzed, and box plots were drawn as follows. Figure 5 As shown in the figure. Among the four machine learning models, ANN and RF models had the highest hit rates in the POD index, while Transformer and LSTM models had slightly lower hit rates. For the FAR index, the four models and the two satellite precipitation products were not significantly different. In the CSI index, the median of the Transformer model was similar to that of the IMERG product, while the CSI index of the LSTM model was slightly better than that of the Transformer model and the two satellite precipitation products, but lower than that of the ANN and RF models.
[0087] Precipitation typically exhibits temporal and spatial characteristics, showing significant spatial variations on the Tibetan Plateau. To improve the accuracy of precipitation fusion on the Tibetan Plateau, this embodiment uses the DEM (Digital Elevation Model) as input data in precipitation fusion. Figure 6 Spatial distributions of precipitation event capture indices from four precipitation fusion models, GSMAP and IMERG products, were plotted. Overall, the spatial distribution of all precipitation indices shows poor precipitation event capture capabilities at stations on the northern and northeastern edges of the Tibetan Plateau. These factors collectively affect the reliability and accuracy of precipitation event observations in this region.
[0088] To further evaluate the accuracy of the four precipitation fusion models on the Tibetan Plateau at different altitudes, this method divides meteorological stations into four levels based on their altitude. The spatial distribution of stations at different altitudes is shown below.Figure 7 As shown. Figure 7 In the figure (a), the correlation coefficient (CC) is represented. Figure 7 In the middle (b), the root mean square error (RMSE) is represented. Figure 7 In the middle (c), the relative bias (RB) is represented. The merged precipitation data from the GSMAP and IMERG products and the outputs of four precipitation fusion models are categorized according to different altitudes. Three indicators for each altitude are plotted in a bar chart as shown below. Figure 7 As shown, overall, the merged precipitation data in all models gradually decreases with increasing altitude. In the 3000-3900m altitude range, the precipitation fusion model achieves high accuracy. Among the different models, ANN, RF, and Transformer models fit the precipitation trend of the Tibetan Plateau well, while the LSTM model shows poor fitting accuracy.
[0089] Box plots were created using precipitation classification indices derived from the fused precipitation data of GSMAP and IMERG products and four precipitation fusion models, categorized by different altitudes. Figure 8 As shown, the CSI indices of all four precipitation fusion models are higher than those of the two satellite precipitation products. For the ANN, RF, and LSTM models, the CSI indices initially increase and then decrease with increasing altitude, reaching their optimal level in the 3900-4300m range. Above 4300m, the CSI indices decrease further. In contrast, the Transformer model's CSI indices increase with altitude. Looking at different models, the CSI indices of the ANN and RF models are consistently higher than those of the LSTM and Transformer models across different altitude regions. This same pattern applies to the overall CSI index. Therefore, for precipitation fusion models in the Tibetan Plateau study area, the ANN and RF models offer better accuracy.
[0090] In summary, this invention uses near-real-time satellite precipitation products GSMAP and IMERG, combined with geospatial DEM data as model input, in the case of sparse rainfall stations. It employs four machine learning algorithms (ANN, RF, LSTM, and Transformer) to construct a precipitation fusion model, evaluates the fused precipitation from spatial and precipitation level perspectives, selects the best fused precipitation data, and improves the accuracy of precipitation data.
[0091] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A near-real-time satellite precipitation fusion method based on machine learning for sparse rain gauge stations, characterized in that, Includes the following steps: S1. Acquire near-real-time satellite precipitation data, DEM data, and sparse rain gauge measured precipitation data; S2. Extract real-time satellite precipitation data and DEM data at the scale of sparse rain gauge stations, and establish multiple data input combinations; S3. Construct a precipitation fusion model, which includes artificial neural networks, random forests, long short-term memory networks, and Transformers. Embed a precipitation level weight adjustment module in all precipitation fusion models to adjust the weights of different precipitation levels in the model. S4. The various data inputs obtained in S2 are combined as inputs to the precipitation fusion model, and the measured precipitation data from sparse rain gauge stations are used as outputs. The training set and validation set are divided through cross-validation to train the precipitation fusion model. S5. Construct a spatial fusion precipitation accuracy evaluation system and a fusion precipitation intensity accuracy evaluation system to comprehensively evaluate the accuracy of precipitation fusion of various precipitation fusion models and obtain the best precipitation fusion data under the condition of sparse rain gauge stations.
2. The near-real-time satellite precipitation fusion method based on machine learning for sparse rain gauges according to claim 1, characterized in that, The near-real-time satellite precipitation data includes GSMAP GNRT diurnal data and IMERG Early diurnal data, with a spatial resolution of 0.1°.
3. The near-real-time satellite precipitation fusion method based on machine learning for sparse rain gauges according to claim 2, characterized in that, The various data input combinations include GSMAP+IMERG+DEM, GSMAP+DEM, IMERG+DEM, and GSMAP+IMERG.
4. The near-real-time satellite precipitation fusion method based on machine learning for sparse rain gauges according to claim 3, characterized in that, The precipitation level weight adjustment module assigns differentiated weights to samples of different precipitation levels using a weighted mean square error loss function, specifically: When the precipitation is ≥0.1mm / day, a weight value of 5.0 is assigned. When the precipitation is less than 0.1 mm / day, a weight value of 1.0 is assigned.
5. The near-real-time satellite precipitation fusion method based on machine learning for sparse rain gauges according to claim 4, characterized in that, The method of embedding a precipitation level weight adjustment module in all precipitation fusion models includes: In the loss function calculation stage of the artificial neural network, a weighted mean square error loss function is embedded. In each training iteration, the artificial neural network assigns weights according to specific conditions of the target value to calculate the weighted mean square error. The weighted mean squared error loss function is embedded in the calculation of the node splitting criteria of the random forest decision tree. Each time a node splits, the weighted mean squared error loss function will dynamically adjust the weights according to the target value, thus making it more inclined to split high-weight samples accurately. In the loss function calculation stage of the Long Short-Term Memory Network, a weighted mean squared error loss function is embedded. For time series data, the weight of the sample is dynamically calculated and the loss value is adjusted each time the predicted value is updated by time step. By embedding a weighted mean squared error loss function in the loss function calculation stage of Transformer, the weights are dynamically adjusted to give the loss function a larger penalty value for heavy precipitation events. The attention distribution is also dynamically adjusted to ensure that Transformer captures global dependencies of sequences while guiding the self-attention mechanism to focus on key spatiotemporal features.
6. The near-real-time satellite precipitation fusion method based on machine learning for sparse rain gauges according to claim 5, characterized in that, The precipitation levels are classified based on the quantile method as follows: Level L1, precipitation: 0.1 mm / day - 0.7 mm / day; Level L2, precipitation: 0.7 mm / day - 2.9 mm / day; Level L3, precipitation: 2.9 mm / day - 10.8 mm / day; Level L4, precipitation value: >10.8mm / day.
7. The near-real-time satellite precipitation fusion method based on machine learning for sparse rain gauges according to claim 6, characterized in that, The cross-validation uses 3-fold cross-validation and includes the following steps: The measured precipitation data from sparse rain gauge stations were randomly divided into three sections. In each iteration, two sections were selected as the training set, and the remaining section was used as the validation set. The training and validation process was repeated three times to ensure that all rainfall stations participated in the training and validation.
8. The near-real-time satellite precipitation fusion method based on machine learning for sparse rain gauges according to claim 7, characterized in that, S5 includes the following steps: By using a spatial fusion precipitation accuracy evaluation system, the precipitation accuracy index of each precipitation fusion model at each rain gauge station is calculated, and the evaluation results of all rain gauge stations are statistically analyzed to assess the overall performance of each precipitation fusion model in the spatial dimension. By integrating the precipitation intensity accuracy evaluation system, precipitation is classified according to intensity level, and the accuracy indicators of precipitation at each level are calculated and compared to evaluate the ability of each precipitation fusion model to capture different precipitation intensities. By combining the results of the spatial fusion precipitation accuracy evaluation system and the fusion precipitation intensity accuracy evaluation system, a weighted average is calculated and ranked from the two dimensions of spatial accuracy and precipitation intensity accuracy. Finally, the best precipitation fusion data is selected under the condition of sparse rain gauge stations.
9. The near-real-time satellite precipitation fusion method based on machine learning for sparse rain gauges according to claim 8, characterized in that, The precipitation accuracy indicators include probabilistic indicators and statistical indicators. The probabilistic indicators include detection rate, false alarm rate, and key success index, while the statistical indicators include correlation coefficient, root mean square error, and relative deviation.