A method for generating CMIP daily-scale precipitation data pattern ensemble based on machine learning
By constructing random forest regression and classification models using machine learning methods, we solved the problem of low simulation accuracy of CMIP daily precipitation data using the multi-model averaging method, and achieved the generation of high-precision spatiotemporal precipitation data, especially the accurate distinction between rainy and non-rainy conditions.
Patent Information
- Application Number
- CN202411493893.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2044-10-24
AI Technical Summary
The existing multi-model averaging method has low simulation accuracy when generating CMIP daily-scale precipitation data and fails to effectively consider the spatial differences between different models, especially in distinguishing between rainy and non-rainy conditions.
Using machine learning methods, by collecting and interpolating CMIP multi-model and observational precipitation data, constructing training samples and training random forest regression and classification models, we calculate monthly and daily precipitation under future scenarios and generate a high-precision precipitation data pattern set.
The simulation accuracy of CMIP daily precipitation data has been improved, which can more accurately reflect the spatiotemporal heterogeneity and accurately distinguish between rainy and non-rainy conditions on a daily scale, generating a high-temporal-resolution future scenario precipitation dataset.
Smart Images

Figure CN119377676B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of hydrological technical analysis, and in particular relates to a method for generating a CMIP daily-scale precipitation data pattern set based on machine learning. Background Art
[0002] Global change and climate change have become one of the most pressing issues worldwide in recent years. The IPCC's Sixth Assessment Report indicates that with further global warming, the spatiotemporal distribution of precipitation and extreme precipitation events will undergo significant changes. The Coupled Model Intercomparison Project (CMIP) is a crucial tool for studying future climate change. However, there is significant variability in model results from different institutions, necessitating the use of model ensembles to correct for individual model biases. Existing studies often use multi-model averaging to generate model ensembles. Simulation results from different models can vary significantly across time and space, particularly in precipitation data. For example, some models perform well in the flood season but poorly in the dry season; others perform well in mountainous areas but poorly in plains. While multi-model averaging is simple and can mitigate some of these biases, it cannot effectively improve the accuracy of ensemble results and fails to account for spatial variability among models. Furthermore, on a daily scale, the number of rainy days varies across models, leading to varying degrees of precipitation on a daily basis, making it difficult to distinguish between rainy and dry periods. Summary of the Invention
[0003] The purpose of the present invention is to provide a method for generating a CMIP daily-scale precipitation data pattern set based on machine learning, which can generate a CMIP daily-scale precipitation dataset with higher simulation accuracy and spatiotemporal differences compared with the multi-model averaging method.
[0004] To achieve the above object, the present invention provides the following technical solutions:
[0005] The present invention discloses a method for generating a CMIP daily-scale precipitation data pattern set based on machine learning, the method comprising the following steps:
[0006] Step 1: Data collection: Collect historical CMIP multi-model daily precipitation data and daily observation precipitation data in the study area, and interpolate precipitation data of different spatial resolutions to a unified spatial resolution;
[0007] Step 2: Construct training samples: CMIP multi-model precipitation data are bias-corrected based on historical observed precipitation data and sampled to the monthly scale. Annual precipitation distribution characteristics are extracted grid-by-grid based on the monthly observed precipitation data. For each grid, the central grid is considered and the correlation between the central grid and its surrounding grids is calculated. Peripheral grids with correlation coefficients greater than 0.95 are selected. The CMIP multi-model monthly precipitation data and annual precipitation distribution characteristics of the central grid and the selected surrounding grids are used as training sample features, and the corresponding monthly observed precipitation data are used as training sample labels to construct training samples.
[0008] Step 3: Model training: The random forest regression model is trained grid-by-grid using the training samples constructed in step 2. Based on the trained model, monthly precipitation under different future scenarios is calculated. The random forest classification model is trained grid-by-grid using historical daily rainfall data as training samples. Daily rainfall data under different future scenarios are calculated.
[0009] Step 4. Generate a daily precipitation data pattern set: Based on the daily rain data under different future scenarios, calculate the monthly precipitation weight of each rainy day; based on the monthly precipitation under different future scenarios and the monthly precipitation weight of rainy days, generate the daily precipitation under different future scenarios, and then obtain the daily precipitation data pattern set.
[0010] Furthermore, the step 1 involves collecting historical CMIP multi-model daily precipitation data and daily observed precipitation data for the study area, and interpolating precipitation data of different spatial resolutions to a uniform spatial resolution. Specifically, the collected data include CMIP6 multi-model precipitation data and CN05.1 grid-based observed precipitation data.
[0011] The CMIP6 multi-model precipitation data was obtained from the Internet and interpolated to a uniform spatial resolution of 0.25° using bilinear interpolation. The formula for the bilinear interpolation method is:
[0012]
[0013] Where: f(x,y) is the coordinate point to be interpolated and its corresponding value; Q 11 (x1,y1),Q 21 (x2,y1),Q 21 (x2,y1),Q 12 (x1, y2) are the coordinates of four known points and their corresponding values;
[0014] The CN05.1 grid observation precipitation data is a daily observation data with a spatial resolution of 0.25° obtained by interpolation and superposition of more than 2,400 national stations of the National Meteorological Information Center using the thin disk spline function method and the angular distance weight method.
[0015] Furthermore, the bias correction of the CMIP multi-model precipitation data based on the historical observed precipitation data described in step 2 is specifically performed by using the transfer cumulative probability step-by-step method to perform the bias correction. Assuming that there is a transfer function T in the historical period, the cumulative probability distribution function of the historical observed precipitation data can be related to the cumulative probability distribution function of the historical CMIP multi-model precipitation data, as shown in the following formula:
[0016] T(F MH (X))=F OH (X) (2)
[0017] First, define u=F MH (X), then The value range of u is [0,1]. Substituting X into the equation, the transfer function T is:
[0018]
[0019] Assuming that the transfer function still holds true in the future, the corrected result is:
[0020] T(F MH (X))=F OF (X) (4)
[0021]
[0022] Where: F OH (X) is the cumulative probability distribution function of historical observed precipitation data; F MH (X) is the cumulative probability distribution function of historical CMIP multi-model precipitation data; F OF (X) is the cumulative probability distribution function of future observed precipitation; F CF (X) is the cumulative probability distribution function of future CMIP multi-model precipitation data after bias correction.
[0023] Furthermore, the step 2 of extracting the annual distribution characteristics of precipitation based on the monthly scale observed precipitation data grid by grid is as follows: given the monthly scale observed precipitation data P m (x,y), where m represents the month m = 1, 2, ..., 12, (x,y) represents the grid position, and for each grid (x,y), calculate the multi-year average monthly precipitation To determine the month with the highest precipitation, label the annual distribution characteristics; define the feature labeling matrix T(x,y), whose size is the same as the raster data; extract and label the annual distribution characteristics of precipitation grid by grid according to the following rules: label the month with the highest precipitation as 4, the second and third months as 3, the fourth and fifth months as 2, and the remaining months as 0;
[0024] Multi-year average monthly precipitation The calculation formula is:
[0025]
[0026] Where: It represents the precipitation in month m of year n; N is the number of observation years.
[0027] Furthermore, the calculation of the correlation between the central grid and its surrounding grids in step 2 is specifically as follows: for each grid (x, y), it is regarded as the central grid, and its correlation with the surrounding 8 grids is calculated. The correlation calculation uses the Pearson correlation coefficient PCC, and the calculation formula is:
[0028]
[0029] Where: X is the monthly precipitation data of the central grid; Y is the monthly precipitation data of the surrounding grids.
[0030] Furthermore, in step 3, the random forest regression model is trained grid-by-grid using the training samples constructed in step 2. Based on the trained model, the monthly precipitation under different future scenarios is calculated as follows: using the training samples constructed in step 2, the CMIP multi-model monthly precipitation data and the annual distribution characteristics of precipitation of the central grid and the filtered surrounding grids are used as training sample features, the corresponding monthly observed precipitation data are used as training sample labels, the random forest regression model is trained grid-by-grid, and the grid search algorithm is used for hyperparameter optimization;
[0031] The random forest regression model is:
[0032]
[0033] Where: represents the predicted value of the random forest model at the grid (x, y); f k (x, y; θ) represents the k-th decision tree model; θ represents the model parameter set;
[0034] The grid search algorithm formula is:
[0035]
[0036] Where: θ *is the best hyperparameter combination; GridNum is the total number of grids; Represents the predicted value of the model at the grid (x, y); P obs (x,y) represents the actual observation value of the grid (x,y);
[0037] Based on the trained random forest regression model for each grid, the monthly precipitation under different scenarios is calculated by taking the bias-corrected CMIP multi-model monthly precipitation data and annual distribution characteristics as input.
[0038] Furthermore, in step 3, the random forest classification model is trained grid-by-grid using the historical daily-scale rain data as training samples to calculate the daily rain data under different future scenarios. Specifically, the historical CMIP multi-model daily-scale precipitation data and the daily-scale observed precipitation data are classified into two types: rain and no rain, and labeled as 0 and 1 respectively; the rain or no rain data of the multi-model daily-scale precipitation data are used as training sample features, and the rain or no rain data of the observed precipitation data are used as training sample labels, and the random forest classification model is trained grid-by-grid, and the grid search algorithm is used for hyperparameter optimization;
[0039] The two-category labeling formula is:
[0040]
[0041] Where: τ is the rainfall threshold (0.5 mm), which is used to distinguish between rain and no rain; is the precipitation of the kth mode on the dth day;
[0042] Based on the trained random forest classification model, the bias-corrected future CMIP multi-model daily-scale rain-or-no rain characteristics were used as input to calculate the daily rain-or-no rain data under different future scenarios.
[0043] Furthermore, the step 4 of calculating the monthly precipitation weight of each rainy day based on the daily rainy data under different future scenarios is specifically as follows: based on the calculated daily rainy data under different future scenarios, the daily precipitation of the rainy day is calculated using a multi-model averaging method, and the monthly precipitation weight of each rainy day is calculated;
[0044] The calculation formula for precipitation weight is:
[0045]
[0046] Where: W d,m represents the precipitation weight of day d in month m, reflecting the proportion of the precipitation on that day in the total precipitation of the whole month; P d (x,y) is the model average precipitation on rainy day d; P m(x,y) is the precipitation on rainy days in the month P d The sum of (x,y);
[0047] Based on the monthly precipitation under different scenarios and the monthly precipitation weights of rainy days, the daily precipitation under different scenarios is generated, and the daily precipitation data pattern set is obtained as follows:
[0048] Based on the calculated monthly precipitation under different future scenarios and the calculated monthly precipitation weights of rainy days, the monthly precipitation data are spread to the daily scale to generate daily precipitation under different future scenarios, and then a set of daily precipitation data patterns is obtained.
[0049] The calculation formula for daily precipitation is:
[0050]
[0051] Where: W d,m (x,y) is the precipitation weight of day d in month m; is the monthly precipitation of month m on the grid (x, y); P d,ensemble (x,y) is the daily precipitation on the dth day on the grid (x,y) under different future scenarios.
[0052] The beneficial effects of the present invention are as follows: a method for generating a CMIP daily-scale precipitation data pattern set based on machine learning is proposed in the present invention, which achieves a better data set result than the pattern averaging method by training a random forest regression model on a monthly scale by introducing the annual distribution characteristics of precipitation. The method of the present invention takes into account the performance differences between different patterns, can more accurately reflect the spatiotemporal heterogeneity, and effectively improves the simulation accuracy of the results; secondly, by training a random forest classification model on a daily scale, it is judged whether there will be rain in the future, and the distribution of rainy days in a month under future scenarios is calculated. The daily-scale precipitation data is further spread with the monthly-scale precipitation data as a constraint, thereby forming a high-temporal-resolution future scenario daily-scale precipitation dataset.
[0053] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 This is a flow chart of the method of the present invention;
[0055] Figure 2 This is an overview of the study area in Example 1;
[0056] Figure 3 This is a schematic diagram of the CMIP multi-model precipitation data interpolation process in Example 1;
[0057] Figure 4Box plots of annual precipitation before and after bias correction for the four watershed units in Example 1;
[0058] Figure 5 This is a tree structure diagram of the random forest regression model after training in Example 1;
[0059] Figure 6 Taylor diagrams of the CMIP models and model sets for precipitation in the four watershed units of Example 1;
[0060] Figure 7 This is a binary classification result diagram of whether there is rain or not in random months in Example 1;
[0061] Figure 8 This is the precipitation weight map for rainy days in Example 1. DETAILED DESCRIPTION
[0062] The present invention discloses a method for generating a CMIP daily-scale precipitation data pattern set based on machine learning, such as Figure 1 As shown, the method includes the following steps:
[0063] Step 1: Data collection: Collect historical CMIP multi-model daily precipitation data and daily observation precipitation data in the study area, and interpolate precipitation data of different spatial resolutions to a unified spatial resolution.
[0064] Among them, the collected data include CMIP6 (Sixth International Coupled Model Intercomparison Project) multi-model precipitation data and CN05.1 grid observation precipitation data.
[0065] CMIP6 multi-model precipitation data can be obtained online at https: / / esgf-data.dkrz.de / search / cmip6-dkrz / . Because the spatial resolutions of the collected CMIP6 multi-model precipitation data vary, bilinear interpolation is used to bring them to a uniform spatial resolution of 0.25° for subsequent calculations. The formula for bilinear interpolation is:
[0066]
[0067] Where: f(x,y) is the coordinate point to be interpolated and its corresponding value; Q 11 (x1,y1),Q 21 (x2,y1),Q 21 (x2,y1),Q 12 (x1, y2) are the coordinates of four known points and their corresponding values.
[0068] The CN05.1 grid precipitation observation data are daily observation data with a spatial resolution of 0.25° obtained by interpolating and superimposing the thin disk spline function method (ANUSPLIN) and the angular distance weighting method (ADW) at more than 2,400 national stations of the National Meteorological Information Center.
[0069] Step 2: Construct training samples: bias-correct the CMIP multi-model precipitation data based on historical observed precipitation data and sample the precipitation data to the monthly scale. Extract the annual distribution characteristics of precipitation grid by grid based on the monthly observed precipitation data. For each grid, treat it as the central grid, calculate the correlation between the central grid and its surrounding grids, and select surrounding grids with a correlation coefficient greater than 0.95. Use the CMIP multi-model monthly precipitation data and annual distribution characteristics of precipitation of the central grid and the selected surrounding grids as training sample features, and the corresponding monthly observed precipitation data as training sample labels to construct training samples.
[0070] The specific steps include:
[0071] Step 21: Perform bias correction on the CMIP multi-model precipitation data based on historical observed precipitation data and resample the precipitation data to the monthly scale. Specifically, the bias correction is performed using the cumulative probability distribution function (CDF-t) method. The main idea of CDF-t is to assume that there is a transfer function T in the historical period, which can establish a relationship between the cumulative probability distribution function (CDF) of the historical observed precipitation data and the cumulative probability distribution function of the historical CMIP multi-model precipitation data, as shown in the following formula:
[0072] T(F MH (X))=F OH (X) (2)
[0073] First, define u=F MH (X), then The value range of u is [0,1]. Substituting X into the equation, the transfer function T is:
[0074]
[0075] Assuming that the transfer function still holds true in the future, the corrected result is:
[0076] T(F MH (X))=F OF (X) (4)
[0077]
[0078] Where: F OH (X) is the cumulative probability distribution function of historical observed precipitation data; F MH(X) is the cumulative probability distribution function of historical CMIP multi-model precipitation data; F OF (X) is the cumulative probability distribution function of future observed precipitation; F CF (X) is the cumulative probability distribution function of future CMIP multi-model precipitation data after bias correction.
[0079] Step 22: Extract the annual distribution characteristics of precipitation based on the monthly scale observation precipitation data grid by grid. Specifically, given the monthly scale observation precipitation data P m (x,y), where m represents the month m = 1, 2, ..., 12, (x,y) represents the grid position, and for each grid (x,y), calculate the multi-year average monthly precipitation To determine the month with the highest precipitation, annotate the annual distribution characteristics. Define a feature annotation matrix T(x,y) with the same size as the raster data. Extract and annotate the annual distribution characteristics of precipitation on a grid-by-grid basis according to the following rules: label the month with the highest precipitation as 4, the second and third months as 3, the fourth and fifth months as 2, and the remaining months as 0.
[0080] Multi-year average monthly precipitation The calculation formula is:
[0081]
[0082] Where: It represents the precipitation in month m of year n; N is the number of observation years.
[0083] Step 23: Form a training sample. For each grid (x, y), consider it as the central grid and calculate its correlation with the surrounding eight grids. The correlation calculation uses the Pearson correlation coefficient PCC, and the calculation formula is:
[0084]
[0085] Where: X is the monthly precipitation data of the central grid; Y is the monthly precipitation data of the surrounding grids.
[0086] For each central grid (x, y), if the PCC correlation coefficient between a surrounding grid and the central grid is greater than 0.95, the surrounding grid is added to the training sample. The intra-annual precipitation distribution characteristics refer to the corresponding values (0, 2, 3, or 4) in the feature labeling matrix T(x, y) defined in step 22. Using a portion of historical CMIP6 precipitation data and historical observational data as the training set, the following training sample is constructed: feature X {CMIP multi-model monthly precipitation, monthly labeling features}, label Y {observed monthly precipitation}. The training set contains the central grid and the filtered surrounding grid data. Repeat the above process until all grids are processed to form the final training sample set.
[0087] Step 3: Model training: The random forest regression model is trained grid-by-grid using the training samples constructed in step 2. Based on the trained model, the monthly precipitation under different future scenarios is calculated. The random forest classification model is trained grid-by-grid using the historical daily rainfall data as training samples. The daily rainfall data under different future scenarios are calculated.
[0088] The specific steps include:
[0089] Step 31: Using the training samples constructed in step 2, the CMIP multi-model monthly precipitation data and annual precipitation distribution characteristics of the central grid and the filtered peripheral grids are used as training sample features, and the corresponding monthly observed precipitation data are used as training sample labels. A random forest regression model is trained grid by grid, and a grid search algorithm is used for hyperparameter optimization.
[0090] The random forest regression model is:
[0091]
[0092] Where: represents the predicted value of the random forest model at the grid (x, y); f k (x, y; θ) represents the kth decision tree model; θ represents the model parameter set.
[0093] The grid search algorithm formula is:
[0094]
[0095] Where: θ * is the best hyperparameter combination; GridNum is the total number of grids; Represents the predicted value of the model at the grid (x, y); P obs (x,y) represents the actual observed value of the grid (x,y).
[0096] Step 32: Based on the random forest regression model of each grid trained in step 31, the bias-corrected future CMIP multi-model monthly precipitation data and annual distribution characteristics are used as input to calculate the monthly precipitation under different scenarios in the future.
[0097] Step 33: Binary classify the historical CMIP multi-model daily precipitation data and the daily observed precipitation data into two categories: rain and no rain, labeled 0 and 1, respectively. A random forest classification model is trained grid by grid, using the presence or absence of rain in the multi-model daily precipitation data as training sample features and the presence or absence of rain in the observed precipitation data as training sample labels. A grid search algorithm is used for hyperparameter optimization.
[0098] The two-category labeling formula is:
[0099]
[0100] Where: τ is the rainfall threshold (0.5 mm), which is used to distinguish between rain and no rain; is the precipitation of the kth mode on the dth day.
[0101] Step 34: Based on the random forest classification model trained in step 33, the bias-corrected future CMIP multi-model daily-scale rain-or-no rain features are used as input to calculate the daily rain-or-no rain data under different future scenarios.
[0102] Step 4. Generate a daily precipitation data pattern set: Based on the daily rain data under different future scenarios, calculate the monthly precipitation weight of each rainy day; based on the monthly precipitation under different future scenarios and the monthly precipitation weight of rainy days, generate the daily precipitation under different future scenarios, and then obtain the daily precipitation data pattern set.
[0103] The specific steps include:
[0104] Step 41: Based on the daily rain data under different future scenarios obtained in step 34, the daily precipitation on rainy days is calculated using a multi-mode averaging method, and the monthly precipitation weight of each rainy day is calculated.
[0105] The calculation formula for precipitation weight is:
[0106]
[0107] Where: W d,m represents the precipitation weight of day d in month m, reflecting the proportion of the precipitation on that day in the total precipitation of the whole month; P d (x,y) is the model average precipitation on rainy day d; P m (x,y) is the precipitation on rainy days in the month P d The sum of (x,y).
[0108] Step 42: Based on the monthly precipitation under different future scenarios calculated in step 32 and the monthly precipitation weights of rainy days calculated in step 41, the monthly precipitation data is spread to the daily scale to generate daily precipitation under different future scenarios, thereby obtaining a daily precipitation data pattern set.
[0109] The calculation formula for daily precipitation is:
[0110]
[0111] Where: W d,m (x,y) is the precipitation weight of day d in month m; is the monthly precipitation of month m on the grid (x, y); P d,ensemble (x,y) is the daily precipitation on the dth day on the grid (x,y) under different future scenarios.
[0112] The present invention is based on bias-corrected multi-model CMIP precipitation data, using the CN05.1 precipitation dataset as observation data. A random forest method with temporal features is used to perform grid-by-grid regression analysis. This method considers the intra-annual distribution and spatial variation characteristics during model training, and uses a binary classification based on the presence or absence of rain to determine the distribution of daily-scale data. This generates a daily-scale CMIP precipitation dataset with higher simulation accuracy and spatial and temporal differences than the multi-model averaging method.
[0113] Example 1
[0114] This embodiment is an application example of the above method.
[0115] This embodiment discloses a method for generating a CMIP daily precipitation data pattern set based on machine learning. The above method is verified by taking the Tianshan region as the research area. The research area is as follows: Figure 2 The Tianshan region is subdivided into four third-level drainage units, namely the Eastern Tianshan (ETM), Southern Tianshan (STM), Central Tianshan (CTM) and Northern Tianshan (NTM).
[0116] The method described in this embodiment includes the following steps:
[0117] Step 1: Data collection: Collect historical CMIP multi-model daily precipitation data and daily observation precipitation data for the study area, and interpolate precipitation data of different spatial resolutions to a uniform spatial resolution. The detailed information of the CMIP6 models collected in this example is shown in Table 1. The bilinear interpolation method is used to interpolate multi-model precipitation data of different spatial resolutions to a uniform spatial resolution of 0.25°, as shown in Table 1. Figure 3 shown.
[0118] Table 1 CMIP6 model information
[0119]
[0120] Step 2: Construct training samples: bias-correct the CMIP multi-model precipitation data based on historical observed precipitation data and sample the precipitation data to the monthly scale. Extract the annual distribution characteristics of precipitation grid by grid based on the monthly observed precipitation data. For each grid, treat it as the central grid, calculate the correlation between the central grid and its surrounding grids, and select surrounding grids with a correlation coefficient greater than 0.95. Use the CMIP multi-model monthly precipitation data and annual distribution characteristics of precipitation of the central grid and the selected surrounding grids as training sample features, and the corresponding monthly observed precipitation data as training sample labels to construct training samples.
[0121] like Figure 4 As shown, the box plots of annual precipitation in the four basin units before and after bias correction are displayed. Before bias correction, the original CMIP models seriously underestimated the precipitation in the NTM, ETM, and MTM regions. After bias correction, the simulation effect was greatly improved, and its mean was close to the observed value.
[0122] Step 3: Model training: The random forest regression model is trained grid-by-grid using the training samples constructed in step 2. Based on the trained model, the monthly precipitation under different future scenarios is calculated. The random forest classification model is trained grid-by-grid using the historical daily rainfall data as training samples. The daily rainfall data under different future scenarios are calculated.
[0123] Randomly select a coordinate point to display the tree structure of the trained random forest regression model, such as Figure 5 shown. Figure 6 Taylor diagrams for each CMIP model and model ensemble for precipitation in the four basin units ETM, NTM, STM, and CTM are presented. Before ensembling, the correlation coefficients for each sub-unit were around 0.4, with a maximum value not exceeding 0.6. The variances of the individual models were close to those of the observed data, indicating that the constructed models can effectively capture the variability and fluctuation range of the observed data. After ensembling, the correlation coefficients increased to above 0.6, significantly reducing the uncertainty in precipitation simulations.
[0124] Step 4. Generate a daily precipitation data pattern set: Based on the daily rain data under different future scenarios, calculate the monthly precipitation weight of each rainy day; based on the monthly precipitation under different future scenarios and the monthly precipitation weight of rainy days, generate the daily precipitation under different future scenarios, and then obtain the daily precipitation data pattern set.
[0125] Figure 7 The first 25 columns are the rainy or rainy days in different patterns, and the last column is the rainy day judgment after random forest classification and regression. Figure 8 As shown in Figure 2, the range is between 0.008 and 0.12. Then, a daily precipitation data pattern set is generated based on the precipitation weight and monthly precipitation.
[0126] Finally, it should be noted that the above description is only used to illustrate the technical solution of the present invention and is not intended to limit it. Although the present invention has been described in detail with reference to the preferred arrangement scheme, those skilled in the art should understand that the technical solution of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.
Claims
1. A method for generating a CMIP daily-scale precipitation data pattern set based on machine learning, characterized in that: The method comprises the following steps: Step 1: Data collection: Collect historical CMIP multi-model daily precipitation data and daily observation precipitation data in the study area, and interpolate precipitation data of different spatial resolutions to a unified spatial resolution; Step 2: Construct training samples: CMIP multi-model precipitation data are bias-corrected based on historical observed precipitation data and sampled to the monthly scale. Annual precipitation distribution characteristics are extracted grid-by-grid based on the monthly observed precipitation data. For each grid, the central grid is considered and the correlation between the central grid and its surrounding grids is calculated. Peripheral grids with correlation coefficients greater than 0.95 are selected. The CMIP multi-model monthly precipitation data and annual precipitation distribution characteristics of the central grid and the selected surrounding grids are used as training sample features, and the corresponding monthly observed precipitation data are used as training sample labels to construct training samples. Step 3: Model training: The random forest regression model is trained grid-by-grid using the training samples constructed in step 2. Based on the trained model, monthly precipitation under different future scenarios is calculated. The random forest classification model is trained grid-by-grid using historical daily rainfall data as training samples. Daily rainfall data under different future scenarios are calculated. Step 4: Generate a daily precipitation data pattern set: Based on the daily rainfall data under different future scenarios, calculate the monthly precipitation weight of each rainy day; based on the monthly precipitation under different future scenarios and the monthly precipitation weight of rainy days, generate the daily precipitation under different future scenarios, and then obtain the daily precipitation data pattern set; The method of calculating the monthly precipitation weight of each rainy day based on the daily rainy data under different future scenarios is as follows: based on the calculated daily rainy data under different future scenarios, the daily precipitation on the rainy day is calculated using a multi-model averaging method, and the monthly precipitation weight of each rainy day is calculated; The calculation formula for precipitation weight is: Where: W d,m represents the precipitation weight of day d in month m, reflecting the proportion of the precipitation on that day in the total precipitation of the whole month; P d (x,y) is the model average precipitation on rainy day d; P m '(x,y) is the precipitation on rainy days in the month P d The sum of (x,y); Based on the monthly precipitation under different scenarios and the monthly precipitation weights of rainy days, the daily precipitation under different scenarios is generated, and the daily precipitation data pattern set is obtained as follows: Based on the calculated monthly precipitation under different future scenarios and the calculated monthly precipitation weights of rainy days, the monthly precipitation data are spread to the daily scale to generate daily precipitation under different future scenarios, and then a set of daily precipitation data patterns is obtained. The calculation formula for daily precipitation is: Where: W d,m (x,y) is the precipitation weight of day d in month m; is the monthly precipitation of month m on the grid (x, y); P d,ensemble (x,y) is the daily precipitation on the dth day on the grid (x,y) under different future scenarios.
2. The method for generating a CMIP daily-scale precipitation data pattern set based on machine learning according to claim 1, characterized in that: Step 1 involves collecting historical CMIP multi-model daily precipitation data and daily observed precipitation data for the study area, and interpolating precipitation data of different spatial resolutions to a uniform spatial resolution. Specifically, the collected data include CMIP6 multi-model precipitation data and CN05.1 grid-based observed precipitation data. The CMIP6 multi-model precipitation data was obtained from the Internet and interpolated to a uniform spatial resolution of 0.25° using bilinear interpolation. The formula for the bilinear interpolation method is: Where: f(x,y) is the coordinate point to be interpolated and its corresponding value; Q 11 (x1,y1),Q 21 (x2,y1),Q 21 (x2,y1),Q 12 (x1, y2) are the coordinates of four known points and their corresponding values; The CN05.1 grid observation precipitation data is a daily observation data with a spatial resolution of 0.25° obtained by interpolating and superimposing more than 2,400 national stations of the National Meteorological Information Center using the thin disk spline function method and the angular distance weight method.
3. The method for generating a CMIP daily-scale precipitation data pattern set based on machine learning according to claim 1, characterized in that: Step 2 describes the bias correction of CMIP multi-model precipitation data based on historical observed precipitation data, specifically: using the transfer cumulative probability step-by-step method to perform bias correction.
4. The method for generating a CMIP daily-scale precipitation data pattern set based on machine learning according to claim 1, characterized in that: The specific steps of extracting the annual distribution characteristics of precipitation based on monthly observed precipitation data grid by grid in step 2 are as follows: given monthly observed precipitation data P m (x,y), for each grid (x,y), calculate the multi-year average monthly precipitation To determine the month with the highest precipitation and mark the distribution characteristics within the year; Define a feature labeling matrix T(x,y) with the same size as the raster data. Extract and label the annual distribution characteristics of precipitation grid by grid according to the following rules: label the month with the highest precipitation as 4, the second and third months as 3, the fourth and fifth months as 2, and the remaining months as 0. Multi-year average monthly precipitation The calculation formula is: Where: represents the precipitation in the mth month of the nth year; m represents the month, m = 1, 2, …, 12, (x, y) represents the grid position; N is the number of observation years.
5. The method for generating a CMIP daily-scale precipitation data pattern set based on machine learning according to claim 1, characterized in that: The calculation of the correlation between the central grid and its surrounding grids in step 2 is as follows: for each grid (x, y), it is regarded as the central grid, and its correlation with the eight surrounding grids is calculated. The correlation calculation uses the Pearson correlation coefficient PCC, and the calculation formula is: Where: X is the monthly precipitation data of the central grid; Y is the monthly precipitation data of the surrounding grids.
6. The method for generating a CMIP daily-scale precipitation data pattern set based on machine learning according to claim 1, characterized in that: In step 3, the random forest regression model is trained grid-by-grid using the training samples constructed in step 2. Based on the trained model, the monthly precipitation under different future scenarios is calculated as follows: using the training samples constructed in step 2, the CMIP multi-model monthly precipitation data and the annual distribution characteristics of precipitation of the central grid and the filtered surrounding grids are used as training sample features, and the corresponding monthly observed precipitation data are used as training sample labels. The random forest regression model is trained grid-by-grid, and the grid search algorithm is used for hyperparameter optimization. Based on the trained random forest regression model for each grid, the monthly precipitation under different future scenarios was calculated using the bias-corrected future CMIP multi-model monthly precipitation data and annual distribution characteristics as input.
7. The method for generating a CMIP daily-scale precipitation data pattern set based on machine learning according to claim 1, characterized in that: In step 3, the random forest classification model is trained grid-by-grid using the historical daily-scale rain data as training samples to calculate the daily rain data under different future scenarios. Specifically, the historical CMIP multi-model daily-scale precipitation data and the daily-scale observed precipitation data are classified into two types: rain and no rain, and labeled as 0 and 1 respectively; the rain or no rain data of the multi-model daily-scale precipitation data are used as training sample features, and the rain or no rain data of the observed precipitation data are used as training sample labels, and the random forest classification model is trained grid-by-grid, and the grid search algorithm is used for hyperparameter optimization; The two-category labeling formula is: Where: τ is the rainfall threshold (0.5 mm), which is used to distinguish between rain and no rain; is the precipitation of the kth mode on the dth day; Based on the trained random forest classification model, the bias-corrected future CMIP multi-model daily-scale rain-or-no rain characteristics were used as input to calculate the daily rain-or-no rain data under different future scenarios.