Rainfall data set construction method and device, equipment and storage medium
By integrating ERA5 and MSWEP V2 datasets and adopting Swin-UNet architecture and FSS models, the problem of taking into account both temporal and spatial resolution of rainfall datasets is solved, and a high-precision global hourly resolution rainfall benchmark dataset is constructed, which improves the accuracy and efficiency of rainfall prediction.
Patent Information
- Application Number
- CN202510568910.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-05
AI Technical Summary
When constructing existing rainfall datasets for deep learning rainfall prediction, it is difficult to take into account both temporal and spatial resolution, resulting in low accuracy of rainfall prediction.
By integrating the ERA5 dataset and the MSWEP V2 dataset, a rainfall correction dataset was constructed, and a model based on the Swin-UNet architecture and score skill score (FSS) was used for data correction, and appropriate input features were selected to form a global hourly resolution deep learning rainfall benchmark dataset GHPD.
It significantly improved the global overall rainfall accuracy, rainfall accuracy in complex terrain areas and heavy rainfall events, and improved the accuracy of rainfall prediction. The data set size was reduced by 37.5%, but the accuracy was reduced by only 3%-8%.
Smart Images

Figure CN120429614A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a rainfall dataset construction method, apparatus, equipment and storage medium. Background Art
[0002] Based on the data collection source, rainfall data can be divided into rain gauge measurements, satellite observation data, and reanalysis data. Rain gauge measurements are obtained using tools such as rain gauges and radar. These tools directly measure precipitation and can measure rainfall at surface sites. Satellite observation data uses sensors onboard satellites to obtain rainfall information. Its advantages lie in its wide coverage and high temporal and spatial resolution, which can compensate for the shortcomings of ground-based observations in oceans and remote areas. Reanalysis rainfall data is generated by a reanalysis system, which combines observational data with models that encompass physical and dynamic processes to generate an estimate of the system's state on a pre-defined grid.
[0003] Many excellent rainfall datasets have been developed based on rain gauge rainfall data and satellite observations. Schneider et al. collected rainfall data from over 85,000 stations worldwide and constructed the GPCC Full Data, a rainfall dataset covering the period from 1901 to the present. Chen et al. constructed the CPC dataset by interpolating land-based rain gauge observations and reconstructing historical ocean observations using empirical orthogonal functions, reconstructing global monthly precipitation data. Ashouri et al. used an artificial neural network model to estimate global rainfall values based on GridSat-B1 infrared window data measured by satellites, generating the PERSIANN-CDR dataset.
[0004] However, rain gauge and satellite observation data are insufficient for constructing high-quality datasets for deep learning rainfall prediction. Specifically, rain gauge data suffer from a small number of observation sites and uneven distribution. Coverage of oceanic and sparsely populated areas is difficult, resulting in gridded interpolated data that is unreliable and unrepresentative. Satellite rainfall data, on the other hand, is indirectly estimated and fails to reflect the topographic dependence of rainfall, resulting in significant uncertainty.
[0005] Therefore, current high-precision deep learning rainfall prediction models are generally trained based on reanalysis datasets. However, existing reanalysis datasets all have certain limitations. Although the ERA5 dataset has a high temporal and spatial resolution, it has significant errors in rainfall estimation, especially in high-altitude areas of my country, where its data accuracy is significantly insufficient. The MSWEP V2 dataset performs well in rainfall accuracy and has a high spatial resolution (up to ), however, its 3-hour temporal resolution is slightly inferior to ERA5’s 1-hour resolution. The spatial resolution of the NCEP2 dataset is too low, only , it is difficult to accurately capture the detailed characteristics of the rainfall process.
[0006] Existing rainfall datasets have achieved certain results in their respective application scenarios, but there is a problem of difficulty in balancing temporal and spatial resolution, and their accuracy lacks verification.
[0007] Therefore, how to construct a high-quality benchmark dataset specifically for deep learning rainfall prediction to improve the accuracy of rainfall prediction is a technical problem that needs to be solved urgently.
[0008] The above content is only used to assist in understanding the technical solution of the present invention and does not constitute an admission that the above content is prior art. Summary of the Invention
[0009] The main purpose of the present invention is to provide a rainfall dataset construction method, device, equipment and storage medium, aiming to solve the technical problem in the prior art that high-quality benchmark datasets used for deep learning rainfall prediction are difficult to balance temporal resolution and spatial resolution, resulting in low accuracy of rainfall prediction.
[0010] To achieve the above object, the present invention provides a method for constructing a rainfall dataset, the method comprising: In addition, the present invention also proposes a rainfall dataset construction device, which includes: The ERA5 dataset and the MSWEP V2 dataset for the preset years were integrated to construct a rainfall correction dataset. based on The model performs data correction on the rainfall correction data set; The ERA5 rainfall data before and after correction were evaluated from three dimensions: global overall rainfall accuracy, rainfall accuracy in complex terrain areas, and accuracy of heavy rainfall events. The input features in the rainfall correction dataset are screened to obtain a target rainfall dataset.
[0011] Preferably, the MSWEP V2 dataset for the preset year period is used as input, and the ERA5 dataset for the preset year period is used as the true value.
[0012] Preferably, the step of integrating the ERA5 dataset and the MSWEP V2 dataset for a preset year to construct a rainfall correction dataset includes: The ERA5 dataset and MSWEP V2 dataset for the preset year are forward grouped according to the preset time interval, and the MSWEP V2 dataset and ERA5 dataset are aligned according to the time tags to construct the rainfall correction dataset.
[0013] Preferably, the Before the model performs data correction on the rainfall correction dataset, the model further includes: The model adopts the PrecipUNet architecture, embeds an improved skip connection with a residual block, and adopts an attention downsampling module and an attention upsampling module; According to different loss functions Model evaluation and training Model.
[0014] Preferably, The model incorporates the latitude-weighted composite loss function FSS of the fractional skill score. FSS is expressed as: ; Where N represents the total number of prediction points or grid points, x i Represents the predicted value of the i-th position, y i Represents the true value of the i-th position.
[0015] Preferably, screening the input features in the rainfall correction dataset comprises: Optimal selection of input variables and input pressure layers in the rainfall correction dataset.
[0016] In addition, to achieve the above-mentioned purpose, the present invention further proposes a rainfall dataset construction device, the rainfall dataset construction device comprising: A construction module is used to integrate the ERA5 dataset and the MSWEP V2 dataset for a preset year to construct a rainfall correction dataset; Correction module for The model performs data correction on the rainfall correction data set; An evaluation module is used to evaluate ERA5 rainfall data before and after correction based on three dimensions: global overall rainfall accuracy, rainfall accuracy in complex terrain areas, and accuracy of heavy rainfall events; The screening module is used to screen the input features in the rainfall correction dataset to obtain a target rainfall dataset.
[0017] In addition, to achieve the above objectives, the present invention also proposes a rainfall dataset construction device, on which a rainfall dataset construction program is stored. When the rainfall dataset construction program is executed by a processor, the steps of the rainfall dataset construction method described above are implemented.
[0018] In addition, to achieve the above object, the present invention also proposes a storage medium, which stores a rainfall dataset construction program. When the rainfall dataset construction program is executed by a processor, the steps of the rainfall dataset construction method described above are implemented.
[0019] The present invention has the following advantages and positive effects: 1. To address the problem of insufficient rainfall data accuracy in the ERA5 dataset, this paper first constructs a high-precision rainfall correction dataset based on the ERA5 and MSWEP V2 datasets for a preset period of time. On this basis, a rainfall data correction model based on the Swin-UNet architecture and fractional skill scoring (FSS) is proposed. After the model was trained on the calibration dataset for a preset year, it was successfully applied to calibrate ERA5 rainfall data for other years.
[0020] 2. The corrected data obtained by the present invention has achieved significant improvements in the accuracy of global overall rainfall, rainfall accuracy in complex terrain areas, and accuracy of heavy rainfall events.
[0021] 3. Based on the variable information recorded by ERA5, this paper screened out the input variables and pressure layer data suitable for the rainfall benchmark dataset, and successfully constructed the global hourly resolution deep learning rainfall benchmark dataset GHPD. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a structural diagram of a method for constructing a rainfall dataset in a hardware operating environment according to an embodiment of the present invention; Figure 2 This is a flow chart of a first embodiment of a method for constructing a rainfall dataset according to the present invention; Figure 3 In the embodiment of the present invention, a Flowchart for producing a global hourly resolution deep learning rainfall benchmark dataset; Figure 4 is a data visualization diagram of a rainfall correction data set in an embodiment of the present invention; Figure 5 This is a structural diagram of the PrecipUNet model in an embodiment of the present invention; Figure 6 1 is a comparison diagram of rainfall between FSS and DFSS in an embodiment of the present invention; Figure 7 In the embodiment of the present invention 、 and Model training process diagram; Figure 8 This is a comparison diagram of the effects of three models in the embodiment of the present invention (with as a benchmark); Figure 9 The comparison of various indicators before and after the rainfall data correction in the embodiment of the present invention (with the data before calibration as the benchmark); Figure 10 : is a graph showing the relative errors of monthly rainfall data before and after correction compared to the true value in the monsoon and westerly regions of the Qinghai-Tibet Plateau in an embodiment of the present invention; Figure 11 : is a correlation coefficient graph of monthly rainfall data before and after correction in the Qinghai-Tibet Plateau monsoon and westerly regions compared to the true value in an embodiment of the present invention; Figure 12 This is a comparison chart of the average annual rainfall at each station before and after correction and the true value in the embodiment of the present invention; Figure 13 1 is a comparison diagram before and after correction of a cyclone event in an embodiment of the present invention; Figure 14 This is a graph showing the percentage growth of WRMSE for each model when a single variable is missing in an embodiment of the present invention; Figure 15 This is a graph showing the percentage growth of WRMSE for each model under different high-level schemes in an embodiment of the present invention; Figure 16 This is a structural block diagram of the first embodiment of the rainfall dataset construction device of the present invention.
[0023] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0024] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0025] Reference Figure 1 , Figure 1 A schematic diagram of a device structure for constructing a rainfall dataset in a hardware operating environment according to an embodiment of the present invention.
[0026] like Figure 1As shown, the rainfall dataset construction device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display. Optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. In the present invention, the wired interface of the user interface 1003 may be a USB interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a wireless fidelity (WI-FI) interface). The memory 1005 may be a high-speed random access memory (RAM) or a non-volatile memory (NVM), such as a disk drive. Optionally, the memory 1005 may be a storage device independent of the processor 1001.
[0027] Those skilled in the art will understand that Figure 1 The structure shown in the figure does not constitute a limitation on the rainfall dataset construction device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0028] like Figure 1 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a rainfall dataset construction program.
[0029] exist Figure 1 In the rainfall dataset construction device shown, the network interface 1004 is mainly used to connect to the backend server and communicate data with the backend server; the user interface 1003 is mainly used to connect to the user device; the rainfall dataset construction device calls the rainfall dataset construction program stored in the memory 1005 through the processor 1001 and executes the rainfall dataset construction method provided in the embodiment of the present invention.
[0030] Based on the above hardware structure, an embodiment of a rainfall dataset construction method of the present invention is proposed.
[0031] Reference Figure 2 , Figure 2 1 is a flow chart of the first embodiment of the method for constructing a rainfall dataset of the present invention, and provides the first embodiment of the method for constructing a rainfall dataset of the present invention.
[0032] Specifically, in the first embodiment, the rainfall dataset construction method includes the following steps: Step S10: Integrate the ERA5 dataset and the MSWEP V2 dataset for a preset year to construct a rainfall correction dataset.
[0033] In a specific implementation, the execution subject of this embodiment is the rainfall dataset construction device, wherein the rainfall dataset construction device can be an electronic device such as a personal computer or a server, and this embodiment does not limit this. Figure 3 , Figure 3 In the embodiment of the present invention, a Flowchart of a global hourly resolution deep learning rainfall benchmark dataset for a model. The preset year period can be the period from 2001 to 2018, or any other year period, and this embodiment does not limit this. This embodiment is illustrated by taking the preset year period from 2001 to 2018 as an example. The ERA5 dataset and the MSWEP V2 dataset for the period from 2001 to 2018 are integrated to construct a rainfall correction dataset. The MSWEP V2 data is used as input, and the ERA5 data is used as a true value label. Furthermore, in this embodiment, the MSWEP V2 dataset for the preset year period is used as input, and the ERA5 dataset for the preset year period is used as the true value.
[0034] Step S20: Based on The model performs data correction on the rainfall correction dataset.
[0035] It is understood that step S20 specifically includes the following steps: S21: Selecting an MSWEP V2 dataset for a second predetermined year period, for example, data from 1979-2000, to calibrate the ERA5 rainfall data for the same period. First, MSWEP V2 data from 1979-2000 is collected. Second, ERA5 data from the same period is collected and forward grouped according to predetermined time intervals. The predetermined time intervals can be three-hour intervals (e.g., a group labeled 08:00 UTC includes data from 08:00, 09:00, and 10:00 UTC), or other time intervals, which are not limited to this embodiment. Finally, the MSWEP V2 data and ERA5 data are aligned according to their time tags. The resulting ERA5 rainfall correction dataset contains a large number of records, for example, 64,288 records. Figure 4 The data set shows a sample of records at 00:00 on September 1, 1992 in a visual way. In the records, the input is a single channel ( ) of MSWEP V2 rainfall data, output as three channels ( ) using ERA5 rainfall data.
[0036] S22: Based on Swin-Unet, PrecipUNet is further proposed, and model evaluation is performed based on different loss functions. Its overall structure is as follows Figure 5 As shown in the figure, the Attention Down-Sampling Module (ADSM) and the Attention Up-Sampling Module (AUSM) are core modules in PercipUNet, replacing the Patch Merging and Patch Expanding modules in Swin-UNet. In the PrecipUNet architecture, traditional skip connections are replaced with improved skip connections embedded in residual blocks. This design leverages the core features of residual blocks: First, the residual structure of residual blocks is functionally consistent with the skip connections of U-Net, both enabling cross-layer information transfer; second, during backpropagation, the residual structure ensures efficient gradient propagation to shallow layers; third, the convolutional layers within residual blocks extract and transform features from multiple dimensions, significantly improving the model's feature capture capabilities. Therefore, integrating residual blocks into U-Net's skip connections not only enables fast information transfer but also further enhances model performance.
[0037] Furthermore, in this embodiment, the Before the model performs data correction on the rainfall correction dataset, the model further includes: The model adopts the PrecipUNet architecture, embeds an improved skip connection with a residual block, and adopts an attention downsampling module and an attention upsampling module; According to different loss functions Model evaluation and training Model.
[0038] Furthermore, S22 is specifically: S221: To verify the rationality of the PrecipUNet model structure, we selected the R50 U-Net, R50 ViT, and SwinUNet models and compared them with the PrecipUNet model on the ACDC dataset. Table 1 shows the comparison results. DSC (DiceSimilarity Coefficient) is the dice similarity coefficient, which measures the overlap between the predicted result and the true value. Its value ranges from 0 to 1, with values closer to 1 indicating better model performance. The table shows that PrecipUNet performs best.
[0039] Table 1 Comparison of DSC indicators on ADCD dataset
[0040] S222: Latitude-weighted composite loss function incorporating fractional skill scoring. Previous models for error correction in gridded rainfall data have mostly used pixel-level metrics, such as the Mean Separation Estimation (MSE), as loss functions to guide model training and evaluate performance. However, directly using the MSE results in an overly biased bias toward polar regions. Therefore, a weighting factor, as shown in the following formula, is added to balance the impact of errors in different latitudes, ensuring more balanced model training and evaluation.
[0041]
[0042] lat(j): represents the latitude of the jth position, in radians. The cosine of the latitude reflects the area ratio of different latitudes on the Earth's surface. cos(lat(j)): represents the cosine of the latitude of the jth position, used to calculate the weight. The higher the latitude (closer to the pole), the smaller the cosine value; the lower the latitude (closer to the equator), the larger the cosine value. N lat : represents the total number of latitude positions, that is, the number of latitudes involved in the entire calculation. ∑i=1N lat : Represents the sum of the cosine values of all latitudes, which is used to standardize the weights to ensure that all weights add up to 1 (normalization).
[0043] Since the cosine function ranges from 0 to The interval of is monotonically decreasing, so the latitude weighting factor can reduce the impact of polar errors on the results, thereby balancing the polar and equatorial regions. The weighted mean square error at this time is: (2) Where j represents the latitude of the sample point, and k represents the longitude of the sample point.
[0044] N lat Indicates the total number of latitudes, that is, the number of latitude grid points involved in the entire calculation. N lon Represents the total number of longitudes, that is, the number of longitude grid points involved in the entire calculation. w(j) represents the weight factor associated with latitude j. The weight is usually calculated by normalizing the cosine value of the latitude (see the previous weight formula) to reflect the area ratio of different latitudes on the earth's surface. j,k represents the true value (observed value or target value) at the j-th latitude and k-th longitude. Represents the predicted value (model output value) at the j-th latitude and k-th longitude. Represents the square of the error at the jth latitude and the kth longitude, which is used to measure the deviation between the model prediction and the true value. The normalization coefficient in the formula is 1 / N lat N lonIt is to calculate the average error of each grid point.
[0045] WMSE is essentially still a pixel-by-pixel evaluation indicator. To address this limitation, this paper introduces the Fraction Skill Score (FSS) based on the pixel-by-pixel indicator and constructs a composite loss function. FSS is a common spatial validation indicator in atmospheric science that measures the similarity of spatial patterns between grid data. Its value ranges from 0 (no match) to 1 (perfect match). In order to use FSS, we first need to define a rainfall intensity threshold. , which is the percentile of rainfall intensity. It is important to note that the threshold for each image must be calculated separately. The input image and the target image are converted into binary images, that is, the rainfall value is 1 if it is greater than the threshold value, and 0 if it is less than the threshold value. The specific calculation process can be expressed as:
[0046]
[0047] In the formula, Q represents the percentile of rainfall intensity, Percentile represents the calculation of the Qth quantile of input data, x0 represents the rainfall intensity threshold, Represents a binary image, where the rainfall value is 1 if it is greater than the threshold value, and 0 if it is less than the threshold value.
[0048] The FSS value can then be obtained by applying the following formula to the two binary images, which can be used to measure the degree of overlap of rainfall peaks.
[0049]
[0050] FSS represents the final result of the score, and its value range is [0,1]. N represents the total number of prediction points or grid points, that is, the total number of data points in the grid. i Represents the predicted value (model output value) of the i-th position. i Represents the true value (observed value or target value) of the i-th position. It represents the sum of squared errors between all predicted values and true values, and is used to measure the deviation between the prediction and the true data. It represents the sum of squares of all predicted values and true values, and is used to normalize the squared error so that the score range is fixed in [0,1].
[0051] Furthermore, in this embodiment, The model incorporates the dimensionally weighted composite loss function FSS that incorporates fractional skill scoring.
[0052] However, the FSS value obtained by the above formula cannot be directly used as the loss function of the deep learning model. The following two adjustments are required. First, the FSS needs to be flipped so that when FSS=1, it means that there is no match between the input image and the target image, thus generating the maximum gradient (i.e., FSS=1-FSS). Second, the hard threshold operation is mathematically expressed as a non-differentiable step function, as shown in formula (3). There are significant jumps at the threshold, which can lead to gradient breakage, loss oscillation, and optimization stagnation during model training. To address this, this paper replaces the hard threshold with the arctangent function shown below. The replaced result is then smoothed using a Gaussian filter, and the smoothed value is finally normalized to the range [0, 1]. In this example, the modified FSS is referred to as DFSS (Differentiable Fraction Skill Score).
[0053]
[0054] Percentile(input,Q) represents the rainfall threshold, actan is the inverse tangent function, input is the input data, and BinaryMap represents the binary image.
[0055] Comparison of rainfall maps obtained after applying FSS and DFSS, such as Figure 6 As shown. Finally, the composite loss function used for model training in this paper is as follows:
[0056] In the formula, Loss represents the compliance loss function, DFSS0.7 represents the 70th percentile threshold of each image, and so on. WMSE is the weighted mean square error, which is used to adjust the pixel-level deviation. The FSS indicator includes four different percentile thresholds, namely the 70th, 80th, 90th and 95th percentile values of each image, and these four values are accumulated. This allows the model to adjust the spatial distribution of rainfall peaks of different intensities rather than targeting a single peak. The weight of the FSS term is 0.6, which is the optimal value obtained by searching in the range of [0.1,1] with an increment of 0.1. In the composite loss function proposed in the present invention, the FSS term is used to adjust the spatial position of the rainfall peak, and the WMSE term is used to adjust the pixel-level deviation.
[0057] S222: In order to evaluate the effect of adding FSS index into the loss function on adjusting the spatial distribution of rainfall data, this paper designed a set of comparative experiments using 、 and These three different loss functions train three PrecipUNet models and are denoted as 、 and The structural hyperparameters and training hyperparameters of these three models are the same. The training curves of these three models are as follows Figure 7 Table 2 shows the three models 、 and Scores of 11 indicators on the test set. The bold numbers indicate the highest scores. Figure 8 The radar chart shows Model and The model is equivalent to The percentage improvement in the model.
[0058] Table 2 、 and Comparison of results
[0059] S223: Based on training completion The model uses the rainfall data from the MSWEP V2 dataset from 1979 to 2000 as input to generate the corrected ERA5 rainfall data. Furthermore, in this embodiment, the ERA5 dataset and the MSWEP V2 dataset for the preset years are integrated to construct the corrected rainfall dataset, including: The ERA5 dataset and MSWEP V2 dataset for the preset year are forward grouped according to the preset time interval, and the MSWEP V2 dataset and ERA5 dataset are aligned according to the time tags to construct the rainfall correction dataset.
[0060] Step S30: Evaluate the ERA5 rainfall data before and after correction from three dimensions: global overall rainfall accuracy, rainfall accuracy in complex terrain areas, and accuracy of heavy rainfall events.
[0061] It should be noted that step S30 is specifically as follows: S31: Comparison of global rainfall accuracy based on MSWEP V2. ERA5 rainfall data has been The time resolution of the corrected data is the first preset time, which can be 1 hour or other values, and is not limited in this embodiment. The time resolution of the MSWEP V2 rainfall data is the second preset time, which can be 3 hours or other values, and is not limited in this embodiment. This embodiment takes the first preset time of 1 hour and the second preset time of 3 hours as an example for detailed description. Therefore, in this embodiment, the ERA5 rainfall data before and after correction from 1979 to 2000 are forward aggregated for 3 hours (i.e., the rainfall at 00:00 UTC refers to the rainfall between 00:00 UTC and 02:59 UTC), and the gap between the aggregated data and the MSWEP V2 dataset in 11 indicators is observed. Table 4 shows the scores of the ERA5 rainfall data before and after correction on the 11 indicators. Figure 9 The comparison of rainfall data before and after correction is shown in Table 3 and Figure 9 The results show that compared with the original ERA5 rainfall data, the corrected data shows advantages in all 11 indicators.
[0062] Table 3 Comparison of ERA5 rainfall data before and after correction based on MSWEP V2 data
[0063] S32: Comparison of Rainfall Accuracy in Complex Terrain Regions Using the Qinghai-Tibet Plateau as an Example. To verify the effectiveness of the corrected ERA5 rainfall data in complex terrain regions, this example evaluates the Qinghai-Tibet Plateau. Research by Sun H et al. has shown that ERA5 rainfall data overestimates rainfall values in the Qinghai-Tibet Plateau. To assess the quality of the data before and after correction, this example compares the data before and after calibration on monthly and annual scales for the Qinghai-Tibet Plateau.
[0064] Figure 10 and Figure 11 This monthly dataset compares the relative errors and correlation coefficients of rainfall data relative to the true rainfall values for the monsoon- and westerly-dominated basins over the Qinghai-Tibet Plateau before and after correction. The bar graphs represent the average relative error and correlation coefficient of monthly rainfall values relative to the true rainfall values for the 22 years from 1979 to 2000. The upper limit of the error bar represents the maximum relative error or correlation coefficient ever observed, while the lower limit represents the minimum relative error or correlation coefficient ever observed. Figure 8 and Figure 9 The results show that the corrected data reduced the relative error of rainfall on a monthly scale from 75%-143% to 5%-23%, and increased the correlation coefficient from 0.32-0.75 to 0.80-0.99.
[0065] Figure 12Results show that, on an annual scale, pre-calibration rainfall data exhibit poor correspondence and high deviations from rain gauge observations in monsoon- and westerly-dominated basins over the Tibetan Plateau, with correlation coefficients of 0.43 and 0.46 and relative deviations of 116% and 98%. The post-calibration rainfall data provide a better fit, increasing the correlation coefficients to 0.95 and 0.92. This also successfully mitigates the overestimation of rainfall in the ERA5 dataset before calibration, reducing the relative deviations to 13% and 16%. Furthermore, the weighted root mean square error (RMS) of annual rainfall at these stations was reduced from 635 mm and 338 mm to 136 mm and 105 mm.
[0066] S33: In order to verify the improvement of the corrected data during heavy rainfall events, the present invention uses cyclone data for evaluation. Taking the latitude and longitude coordinates of the cyclone as the benchmark, an area with a side length of 1,000 kilometers is cropped on the ERA5 data before correction, the ERA5 data after correction, and the MSWEP V2 data. The MSWEP V2 data is taken as the true value, and for the convenience of comparison, the ERA5 data before and after correction are forward aggregated. A total of 2,349 cyclone records were obtained between 1979 and 2000. Subsequently, the present invention compared and evaluated the ERA5 data before and after correction with the MSWEP V2 data based on four indicators: WMAE, WMSE, SSIM, and peak spacing. Table 4 shows the comparison results. It can be seen that the corrected data exceeds the pre-correction data in terms of both pixel-by-pixel accuracy and spatial similarity.
[0067] Table 4 Comparison of ERA5 cyclone events before and after correction
[0068] Figure 13 This visualization shows a cyclone record where the true value is the MSWEP data. The figure shows that the corrected data is better than the uncorrected data.
[0069] Step S40: Screening the input features in the rainfall correction dataset to obtain a target rainfall dataset.
[0070] It should be understood that screening the input features of the rainfall dataset, excluding rainfall values, primarily involves optimizing the selection of input variables and input pressure layers, providing reliable data support for the rainfall prediction model described later. Furthermore, in this embodiment, screening the input features of the rainfall correction dataset includes optimizing the selection of input variables and input pressure layers in the rainfall correction dataset.
[0071] Step S40 is specifically as follows: S41: Dataset Input Variable Screening. The ERA5 dataset consists of upper-air variables distributed across 37 pressure layers and 16 surface variables. Currently, 12 of these variables are commonly used for rainfall prediction. To select appropriate input variables, this paper selected eight representative deep learning models for rainfall prediction: ResNet, U-Net, ConvLSTM, PredRNN, ViT, SwimT, Pangu, and FuXi. Experiments were conducted to compare model performance before and after variable reduction. The performance metric used in this comparison was the weighted root mean squared error (WRMSE).
[0072] Table 5 shows the reference baseline, which is the 24-hour average WRMSE accuracy of each model using all variables as input. Table 6 shows the 24-hour average WRMSE values of each model when individual variables are missing. Figure 14 This is the WRMSE percentage growth heat map calculated based on Tables 5 and 6.
[0073] Table 5 WRMSE baselines of each model
[0074] Table 6 24-hour average WRMSE of each model when a single variable is missing
[0075] from Figure 14 As can be seen, the lack of specific humidity has the greatest impact on model performance, causing WRMSE to increase by 26.63%-37.5%. The lack of meridional and zonal winds also significantly impacts model performance, increasing WRMSE by 20%-28.92% and 19.18%-28.79%, respectively. The lack of land and sea masks and mean sea level pressure has little impact on model results. Therefore, mean sea level pressure and land and sea masks were removed from the dataset input variables.
[0076] S42: Dataset Input Pressure Layer Screening. Currently, mainstream deep learning rainfall prediction models retain 13 pressure layers at high altitudes: 50hPa, 100hPa, 150hPa, 200hPa, 250hPa, 300hPa, 400hPa, 500hPa, 600hPa, 700hPa, 850hPa, 925hPa, and 1000hPa. However, a large number of redundant pressure layers exist because meteorological elements in adjacent pressure layers exhibit correlation and continuity in the vertical direction of the atmosphere. To select appropriate pressure layers as input variables, eight representative deep learning models for rainfall prediction were selected: ResNet, U-Net, ConvLSTM, PredRNN, ViT, SwimT, Pangu, and FuXi. Comparative experiments were conducted on the eight pressure layer schemes listed in Table 7.
[0077] Table 7 High-level retention scheme
[0078] The comparative experimental results are shown in Table 8. At this time, the land and sea mask and the mean sea level pressure are still retained in the ground variables of the input data. Figure 13 This is the WRMSE percentage growth heat map calculated based on Tables 5 and 8.
[0079] Table 8 24-hour average WRMSE of each model under different schemes
[0080] from Figure 15 From the results, it can be seen that the model performance of Schemes 1, 2, 6, and 7 has declined significantly. Through analysis, it can be found that these models only retain a small amount of the bottom atmosphere. The bottom atmosphere is closest to the ground and can represent the temperature, humidity, and wind field characteristics of the ground, which are closely related to the rainfall intensity. Therefore, the lack of the bottom atmosphere will have a greater impact on the model results. Therefore, the following conclusion can be drawn: the bottom atmosphere needs to be preserved as much as possible. On the basis of preserving all the bottom atmosphere, the present invention conducted multiple experiments to select the middle and upper atmosphere. It was finally determined to retain the following pressure layers: 50hPa, 250hPa, 400hPa, 600hPa, 700hPa, 850hPa, 925hPa, and 1000hPa.
[0081] The average WRMSE of each model under this final solution is shown in Table 9. As can be seen from Table 9, all models under this solution perform well. Compared with the baseline values of each model recorded in Table 5, the 24-hour average WRMSE only decreases by 2%-6%.
[0082] Table 9 24-hour average WRMSE of each model on the final pressure layer scheme
[0083] Based on the conclusions from input variable selection and pressure layer selection, the invented GHPD dataset ultimately contains 10 variables and 8 pressure layers. Five upper-altitude variables are distributed across the eight pressure layers: 50 hPa, 250 hPa, 400 hPa, 600 hPa, 700 hPa, 850 hPa, 925 hPa, and 1000 hPa. These five upper-altitude variables are gravity potential, temperature, specific humidity, meridional wind, and zonal wind. The other five ground variables are 2-meter temperature, 10-meter meridional wind, 10-meter zonal wind, topography, and precipitation. Precipitation serves as both an input and the output to be predicted. Under this scheme, the 24-hour average WRMSE of each model is shown in Table 10.
[0084] Table 10 24-hour average WRMSE of each model on the GHPD dataset
[0085] The final GHPD dataset, with a size of 62.5% of typical input, only experienced a 3%-8% drop in accuracy. The 24-hour average WRMSE for six of the eight models decreased within 5%, with only U-Net and ViT experiencing drops exceeding 5%, to 8.4% and 5.5%, respectively.
[0086] The specific values of the time periods and related times in this embodiment are examples and may also be other values, which are not limited in this embodiment.
[0087] 1. To address the problem of insufficient rainfall data accuracy in the ERA5 dataset before 2001, this paper first constructs a high-precision rainfall correction dataset based on the ERA5 and MSWEP V2 datasets from 2001 to 2018. On this basis, a rainfall data correction model based on the Swin-UNet architecture and fractional skill scoring (FSS) is proposed. After completing training on the 2001-2018 calibration dataset, the model was successfully applied to the calibration of the 1979-2000 ERA5 rainfall data.
[0088] 2. The corrected data obtained by the present invention have achieved significant improvements in the accuracy of global overall rainfall, rainfall accuracy in complex terrain areas, and accuracy of heavy rainfall events. Specifically, in the comparison of global 3-hour aggregated data, compared with the data before correction, the WMAE of the corrected data was reduced by 13.7%, the WMSE was reduced by 18.8%, the SSIM was improved by 4.8%, the 3-hour rainfall error was reduced by 226 mm, and the offset of the highest rainfall peak was reduced by 9.8%. In the comparison of rainfall data on the Qinghai-Tibet Plateau, compared with the data before correction, the relative error of rainfall on the monthly scale was reduced from 75%-143% to 5%-23%, and the monthly correlation coefficient was increased from 0.32-0.73 to 0.80-0.99. On the annual scale, the correlation coefficients of the monsoon region and the westerly region increased from 0.43 and 0.46 to 0.95 and 0.92, respectively, the relative deviations were reduced from 116% and 98% to 13% and 16%, and the annual average WRMSE was reduced by 68.9% and 78.6%. In the comparison of cyclone data, the WMAE of the corrected data decreased by 19.6%, the WMSE decreased by 23.5%, the SSIM increased by 10.4%, and the offset of the maximum rainfall peak decreased by 9.1%.
[0089] 3. Based on the variable information recorded by ERA5, this paper selected input variables and pressure layer data suitable for a rainfall benchmark dataset, successfully constructing a global hourly resolution deep learning rainfall benchmark dataset (GHPD). Compared to currently common input datasets, the GHPD reduces the data volume by 37.5%, but the prediction accuracy of this dataset only decreases by 3%-8%.
[0090] In addition, an embodiment of the present invention further provides a storage medium on which a rainfall dataset construction program is stored. When the rainfall dataset construction program is executed by a processor, the steps of the rainfall dataset construction method described above are implemented.
[0091] In addition, refer to Figure 16 The embodiment of the present invention further provides a rainfall dataset construction device, the rainfall dataset construction device comprising: A construction module 10 is used to integrate the ERA5 dataset and the MSWEP V2 dataset during a preset year to construct a rainfall correction dataset; Correction module 20, for The model performs data correction on the rainfall correction data set; Evaluation module 30 is used to evaluate the ERA5 rainfall data before and after correction from three dimensions: global overall rainfall accuracy, rainfall accuracy in complex terrain areas, and accuracy of heavy rainfall events; The screening module 40 is used to screen the input features in the rainfall correction dataset to obtain a target rainfall dataset.
[0092] Other embodiments or specific implementations of the rainfall dataset construction device of the present invention can refer to the above-mentioned method embodiments and will not be described in detail here.
[0093] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.
[0094] The serial numbers of the embodiments of the present invention are for descriptive purposes only and do not represent superiority or inferiority of the embodiments. In a unit claim that lists several means, several of these means may be embodied by the same item of hardware. The use of the terms first, second, and third, etc., does not denote any order and should be construed as identifiers.
[0095] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as a magnetic disk or optical disk) and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0096] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for constructing a rainfall dataset, characterized in that: The rainfall dataset construction method includes: The ERA5 dataset and the MSWEP V2 dataset for the preset years were integrated to construct a rainfall correction dataset. Based on PrecipUNet Compound The model performs data correction on the rainfall correction data set; The ERA5 rainfall data before and after correction were evaluated from three dimensions: global overall rainfall accuracy, rainfall accuracy in complex terrain areas, and accuracy of heavy rainfall events. The input features in the rainfall correction dataset are screened to obtain a target rainfall dataset.
2. The method for constructing a rainfall dataset according to claim 1, wherein: The MSWEPV2 dataset during the preset year is used as input, and the ERA5 dataset during the preset year is used as the true value.
3. The method for constructing a rainfall dataset according to claim 1, wherein: The ERA5 dataset and the MSWEP V2 dataset for a preset year are integrated to construct a rainfall correction dataset, including: The ERA5 dataset and MSWEP V2 dataset for the preset year are forward grouped according to the preset time interval, and the MSWEP V2 dataset and ERA5 dataset are aligned according to the time tags to construct the rainfall correction dataset.
4. The method for constructing a rainfall dataset according to claim 1, wherein: The based Before the model performs data correction on the rainfall correction dataset, the model further includes: The model adopts the PrecipUNet architecture, embeds an improved skip connection with a residual block, and adopts an attention downsampling module and an attention upsampling module; According to different loss functions Model evaluation and training Model.
5. The method for constructing a rainfall dataset according to claim 4, wherein: exist The model incorporates the latitude-weighted composite loss function FSS of the fractional skill score. FSS is expressed as: ; Where N represents the total number of prediction points or grid points, x i Represents the predicted value of the i-th position, y i Represents the true value of the i-th position.
6. The method for constructing a rainfall dataset according to any one of claims 1 to 5, wherein: Screening the input features in the rainfall correction dataset includes: Optimal selection of input variables and input pressure layers in the rainfall correction dataset.
7. A rainfall dataset construction device, characterized in that: The rainfall dataset construction device comprises: A construction module is used to integrate the ERA5 dataset and the MSWEP V2 dataset for a preset year to construct a rainfall correction dataset; Correction module for The model performs data correction on the rainfall correction data set; An evaluation module is used to evaluate ERA5 rainfall data before and after correction based on three dimensions: global overall rainfall accuracy, rainfall accuracy in complex terrain areas, and accuracy of heavy rainfall events; The screening module is used to screen the input features in the rainfall correction dataset to obtain a target rainfall dataset.
8. A rainfall dataset construction device, characterized in that: The rainfall dataset construction device stores a rainfall dataset construction program, and when the rainfall dataset construction program is executed by a processor, the steps of the rainfall dataset construction method according to any one of claims 1 to 6 are implemented.
9. A storage medium, characterized in that: The storage medium stores a rainfall dataset construction program, which, when executed by a processor, implements the steps of the rainfall dataset construction method according to any one of claims 1 to 6.