Analysis and Attribution Identification Method for Spatiotemporal Variation of Multiple Precipitation Elements under Climate Change
Through multi-factor spatiotemporal change analysis method and GA hyperparameter optimization GBDT model, the accuracy of the analysis of the characteristics of descent in climate change is solved, and an in-depth understanding of the laws of precipitation change and influencing factors is achieved, and a scientific basis for hydrological resource management is provided.
Patent Information
- Application Number
- CN202310732938.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-20
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-06-20
AI Technical Summary
The existing technology is difficult to fully reveal the internal mechanism and influencing factors of climate change descent, and it is unable to effectively utilize multi-source data and multi-scale information, resulting in insufficient accuracy in the analysis and prediction of precipitation changes characteristics.
The multi-factor spatiotemporal change analysis method is adopted to divide the research areas and periods, build precipitation feature clustering and driver factor models, combine with GA hyperparameter-optimized GBDT model to attribution identification of precipitation changes, and use a variety of data processing and analysis technologies such as trend analysis, clustering analysis and neural network analysis to identify multi-angle characteristics and influencing factors of precipitation.
A comprehensive analysis of the spatial and temporal changes of precipitation has been achieved, the accuracy of the laws and trends of precipitation changes has been improved, and the influencing factors and their mechanisms can be more accurately identified, providing a scientific basis for hydrological resource management.
Smart Images

Figure CN116701974B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to precipitation analysis technology, especially a method for spatio-temporal variation analysis and attribution identification of multi-factor precipitation under climate change. Background Art
[0002] Precipitation is an important part of the hydrological cycle and also an important factor affecting surface water resources, ecological environment and human social activities. The identification of spatio-temporal variation characteristics of precipitation and their influencing factors is of great significance in revealing the hydrological process under climate change, evaluating the sustainable utilization of water resources, preventing and reducing disasters, etc.
[0003] At present, the following methods are mainly used for the research on spatio-temporal variation characteristics of precipitation and their influencing factors: methods based on statistical analysis, such as trend analysis, regression analysis, etc.; although the methods based on statistical analysis can reflect the variation characteristics of precipitation, they cannot reveal their internal mechanisms and influencing factors. Methods based on cluster analysis, such as hierarchical clustering method, K-means clustering method, fuzzy clustering method, etc., are mainly used to divide the spatial distribution regions and types of precipitation and analyze their spatial distribution patterns and location rules; the disadvantage of this method is that it cannot consider the mutual relationships between precipitation characteristics and the coupling effects between influencing factors. Methods based on model simulation, such as neural network model, support vector machine model, etc., are mainly used to establish the mapping relationship between precipitation and influencing factors and predict the future changes of precipitation. Although the mapping relationship between precipitation and influencing factors can be established, multi-source data and multi-scale information cannot be fully utilized, and there are problems such as parameter optimization and insufficient generalization ability.
[0004] Therefore, further research and innovation are needed. Summary of the Invention
[0005] Object of the Invention: To provide a method for spatio-temporal variation analysis and attribution identification of multi-factor precipitation under climate change to solve the above problems existing in the prior art.
[0006] Technical Solution: Provide a method for spatio-temporal variation analysis and attribution identification of multi-factor precipitation under climate change, including the following steps:
[0007] Step S1, delimit the scope of the research area and the scope of the research period, collect research data and extract precipitation information and the mapping relationship between clustering and sub-watersheds; each research area includes at least two sub-watersheds, and each research period includes at least two research time periods;
[0008] Step S2, perform clustering based on the precipitation characteristics of each sub-watershed, read each sub-watershed in the research area and generalize it into a point, construct a point pattern analysis method, and perform spatial distribution pattern analysis;
[0009] Step S3: Construct a set of driving factors for precipitation change and a GBDT model optimized by GA hyperparameters. For each research period in the study area, screen sensitive driving factors for precipitation change as key driving factors;
[0010] Step S4: Conduct a cluster analysis on the sub-basins in the study area based on the precipitation change impact factors, obtain the spatial distribution characteristics of each precipitation impact factor, and conduct an attribution analysis on the evolution of the precipitation spatial distribution pattern.
[0011] According to one aspect of the present application, the step S1 is further as follows:
[0012] Step S11: Read the geographical information of the study area, extract watershed data, and divide the study area into a sub-basins; collect and preprocess the research data of each sub-basin in the study area to form time series data of precipitation information; a is a natural number greater than 2;
[0013] Step S12: Construct a set of trend analysis methods. For each sub-basin, use the trend analysis methods one by one to conduct trend tests and mutation tests on the time series data to obtain trend data and mutation point data, and divide the research period into b research periods according to the mutation point data; b is a natural number greater than 2.
[0014] According to one aspect of the present application, the step S11 is further as follows:
[0015] Step S11a: Obtain the digital elevation model data of the study area, and use the ArcGIS analysis tool to preprocess the digital elevation model until the quality of the digital elevation model data meets the standard;
[0016] Step S11b: Call the pre-configured hydrological analysis tool to extract watershed data from the digital elevation model data, determine the boundary and outlet points of the study area according to the terrain characteristics and water flow direction, and generate a watershed layer; conduct sub-basin division on the watershed layer to generate a sub-basin layer;
[0017] Step S11c: Construct a set of precipitation characteristics of the study area, and the precipitation characteristics include extremeness, intensity, rainfall amount, time, and space; collect the precipitation information of each station in each sub-basin of the study area, and use the GIS module or the pre-configured statistical analysis module for preprocessing;
[0018] Step S11d: Use the GIS module or the interpolation module to perform spatial interpolation on the precipitation information to generate a precipitation raster layer. Based on the precipitation raster layer, calculate the precipitation of each sub-basin according to different time scales through the GIS or time series analysis module to form time series data of precipitation information.
[0019] According to one aspect of the present application, the step S1b is further as follows:
[0020] Read digital elevation model data, generate a map of the study area, aggregate adjacent and similar pixels in the study area map into one area, convert them into superpixels, and use them as nodes of the neural network analysis module;
[0021] Determine the connection relationship between each node according to the spatial distance, color similarity and gradient direction between superpixels, and construct the edges of the neural network analysis module;
[0022] Calculate the new feature vectors according to the adjacency matrix and the feature vectors of the previous layer, so as to determine and update the feature vectors of each node and edge;
[0023] Construct the loss function and optimization method of the neural network analysis module, form the neural network analysis module, analyze and obtain the output result, and obtain the watershed layer.
[0024] According to one aspect of the present application, the step S12 is further:
[0025] Step S12a, construct a set of trend analysis methods, and the trend analysis methods include MK, prewhitened MK, detrended prewhitened MK, variance corrected MK, bootstrap MK, TS analysis method and innovative trend analysis method;
[0026] Step S12b, for each sub-watershed, use each trend analysis method to conduct trend test and mutation test on each precipitation feature in the time series data, and form a trend data set and a mutation point data set of precipitation features;
[0027] Step S12c, synthesize the trend data of each precipitation feature obtained by each method to generate the final trend data and mutation point data of each precipitation feature;
[0028] Step S12d, segment the study period according to the mutation point data, and judge whether the sequence length of adjacent study periods is less than the threshold. If it is less, conduct cluster analysis on each precipitation feature of adjacent study periods, and decide whether to merge them into one study period according to whether the similarity of each precipitation feature of adjacent study periods is greater than the threshold; conduct trend analysis on each precipitation feature of each clustered study period to determine the trend direction and amplitude.
[0029] According to one aspect of the present application, the step S2 is further:
[0030] Step S21, read the boundary data of each sub-watershed in the study area, calculate the centroid coordinates of each sub-watershed, and generalize each sub-watershed into a point;
[0031] Step S22: Select precipitation characteristics as clustering indicators, perform standardization processing on each sub-basin to eliminate the dimensional differences between indicators; use the hierarchical clustering method to cluster each sub-basin, select the Euclidean distance as the clustering criterion, select the complete linkage method as the inter-class distance calculation method, draw the hierarchical clustering diagram, determine the number of categories, and obtain the sub-basins included in each category.
[0032] Step S23: Use the R-O function method to analyze the spatial distribution pattern of the sub-basins in each category, select the distance scale range and interval, calculate the R-O function values of each category at different distance scales, draw the R-O function curve of random distribution and the actual R-O function curve, judge the spatial distribution type of each category, and obtain the position law of spatial distribution.
[0033] According to one aspect of the present application, the step S3 is further as follows:
[0034] Step S31: Use the meta-analysis method to collect the driving factors of precipitation change in the study area and construct a set of driving factors for precipitation change.
[0035] Step S32: For each study period in the study area, construct a GBDT model optimized by GA hyperparameters, use the driving factors of precipitation change as input and precipitation parameters as output to simulate the precipitation in the study area, obtain the sensitive factors of precipitation change in the study area, and use the sensitive factors as key driving factors.
[0036] Step S33: Construct a comprehensive discrimination criterion to quantify and rank the key driving factors in descending order.
[0037] According to one aspect of the present application, the step S4 is further as follows:
[0038] Step S41: Based on the values of k key driving factors in each sub-basin as the coordinate values of the sub-basin in the k-dimensional space, calculate the distances between each sub-basin through the Euclidean distance, and cluster the sub-basins; k is a natural number.
[0039] Step S42: Examine the correlation between the precipitation spatial distribution status and the key driving factor spatial distribution status, use the driving factor and precipitation mapping relationship model to simulate the precipitation change degree under the individual or combined action of each key driving factor, cluster each sub-basin according to the precipitation change degree, and calculate the proportion of each category affected by the key driving factors.
[0040] Step S43: If the degree of correlation is higher than the threshold, then combine the spatial distribution of the key influencing factors, and conduct an attribution analysis on the precipitation evolution with different sizes and different change degrees in view of the spatial differences existing in the precipitation evolution process.
[0041] According to one aspect of the present application, the process of attributing the evolution of precipitation of different size levels specifically includes:
[0042] In the i th time period, the precipitation characteristics of each sub-basin are divided into l categories, and the difference between each category lies in the different precipitation amounts. Attribution analysis is carried out on the characteristics presented in the evolution pattern of l categories of precipitation respectively. Using the area affected by the dominant influence as a measurement index, the influence degree of the influencing factors on it is determined.
[0043] ,
[0044] In the formula: α i C,p is the proportion of the precipitation of the i th size level in the p th category that is dominated by the key influencing factor among the meteorological factors in the subsequent precipitation evolution process; A j is the area of the j th sub-basin; n is the number of sub-basins in the basin; m is the number of time periods included in the research period; l is the number of sub-basin clusters based on precipitation size levels in the i th time period, i = 1, 2,..., m - 1, p = 1, 2,..., l.
[0045] According to one aspect of the present application, the process of attributing the evolution of precipitation with different degrees of change specifically includes:
[0046] From the i th time period to the i +1th time period, the precipitation change characteristics of each sub-basin are divided into k categories, and the difference between each category lies in the different degrees of precipitation change. Attribution analysis is carried out on the characteristics presented in the evolution pattern of k categories of precipitation respectively. Using the area affected by the dominant influence as a measurement index, the influence degree of the influencing factors on it is determined.
[0047] ,
[0048] β i C,q is the proportion of the precipitation of the i th degree of change from the i th time period to the q +1th time period that is dominated by a certain key influencing factor in its evolution process; A j is the area of the j th sub-basin; nis the number of sub - basins within the basin; m is the number of time periods included in the study period; k is from the i time period to the i +1 time period, the number of sub - basin clusters based on the degree of precipitation change, i = 1, 2, …, m - 1, q = 1, 2, …, k.
[0049] According to one aspect of the present application, the step S12a is further as follows:
[0050] Obtain trend analysis methods and applicability information for each trend analysis method;
[0051] Analyze the characteristics of the research data of each sub - basin for each research time period one by one, establish a mapping relationship among the sub - basin, research time period, and trend analysis method based on the applicability information, and assign weights to the trend analysis methods for each period of each sub - basin;
[0052] Read the test data set, verify the mapping relationship, and obtain the trend analysis methods for each sub - basin for each research time period and the overall trend analysis method for the research period;
[0053] Construct a set of trend analysis methods based on the trend analysis methods for each sub - basin for each research time period and the overall trend analysis method for the research period.
[0054] According to one aspect of the present application, the step S22 further includes:
[0055] Step S22a: Read the precipitation data of each station in each sub - basin of the research area and analyze to obtain the precipitation center and precipitation boundary of each precipitation process;
[0056] Step S22b: Obtain the distances of each station located within the precipitation boundary from the precipitation center and from the precipitation boundary;
[0057] Step S22c: Take the precipitation data of the station closest to the precipitation center as the standard data, calculate the ratio of the precipitation data of each station within the precipitation boundary to the standard data, and draw a distribution change diagram of similar precipitation processes.
[0058] Beneficial effects: This application can comprehensively analyze multiple characteristics of precipitation, such as extremity, intensity, rainfall amount, time, and space, etc., and reflect the spatio-temporal variation characteristics of precipitation from multiple perspectives; utilize multi-source data and multi-scale information, such as digital elevation model data, precipitation data, impact factor data, etc., to reveal the influencing factors of precipitation from multiple levels; adopt a variety of data processing and analysis techniques, such as trend analysis, clustering analysis, R-O function method, neural network analysis module, GBDT model based on GA hyperparameter optimization, etc., to identify the attribution factors of precipitation from multiple methods. Through multi-angle clustering, the attribution of the basin is carried out, greatly improving the accuracy. Description of the Drawings
[0059] Figure 1 is the flowchart of this application.
[0060] Figure 2 is the flowchart of step S1 of this application.
[0061] Figure 3 is the flowchart of step S2 of this application.
[0062] Figure 4 is the flowchart of step S3 of this application.
[0063] Figure 5 is the flowchart of step S4 of this application.
[0064] Figure 6 is the schematic diagram of the research object considering spatial differences of this application. Detailed Implementation Manner
[0065] As Figure 1 shown, a method for spatio-temporal variation analysis and attribution identification of multiple precipitation elements under climate change is provided, including the following steps:
[0066] Step S1: Define the scope of the research area and the scope of the research period, collect research data and extract the mapping relationship between precipitation characteristics, clustering, and sub-basins; where each research area includes at least two sub-basins, and each research period includes at least two research time periods;
[0067] Among them, precipitation characteristics include extremity, intensity, rainfall amount, time, and space. Among them, extremity includes: maximum daily precipitation at a single station, maximum cumulative precipitation at a single station, and maximum daily precipitation in the precipitation area; intensity includes average daily precipitation in the precipitation area and average daily precipitation in the sub - regions; rainfall amount includes cumulative precipitation in the precipitation area and cumulative precipitation in the sub - regions; time includes number of precipitation days and precipitation months; space includes areas above 50mm, areas above 100mm, areas above 250mm, precipitation center location, and precipitation concentration degree. In this embodiment, the precipitation center location is determined by dividing the study area into N×N grids in a grid - making manner, denoted as WG = 1, 2, …, N×N respectively. The precipitation center location WG is equal to the number of the grid where the precipitation center falls.
[0068] Step S2: Cluster based on the precipitation characteristics of each sub - basin, read each sub - basin in the study area and generalize it into a point, construct a point pattern analysis method, and conduct spatial distribution pattern analysis; Using the point pattern analysis method can better reflect the spatial relationship and pattern among the sub - basins in the study area and provide basic data support for subsequent attribution recognition.
[0069] Step S3: Construct a precipitation change driving factor set and a GBDT model optimized by GA hyperparameters. For each research period in the study area, screen sensitive precipitation change driving factors as key driving factors; Using the GBDT model optimized by GA hyperparameters can more accurately mine and screen out important factors affecting precipitation change and provide reliable data support for subsequent attribution recognition. The genetic algorithm can be used to adaptively optimize the model parameters to improve the accuracy and stability of the model.
[0070] Step S4: Conduct cluster analysis on the sub - basins in the study area based on the precipitation change influencing factors, obtain the spatial distribution characteristics of each precipitation influencing factor, and conduct attribution analysis on the evolution of the precipitation spatial distribution pattern. By conducting attribution analysis on the evolution of the precipitation spatial distribution pattern, it is possible to more accurately identify the influence degree and action mechanism of different factors on precipitation change. Conducting attribution analysis on the evolution of the precipitation spatial distribution pattern provides a scientific basis for hydrological water resource management and decision - making.
[0071] In summary, by adopting the multi - factor spatio - temporal change analysis method, it is possible to comprehensively and systematically study the precipitation characteristics and their influencing factors under climate change, which helps to deeply understand the precipitation change law and trend. By conducting cluster analysis on the sub - basins and extracting the mapping relationship between the sub - basins, it is possible to more accurately depict the differences and connections between different sub - basins in the study area, and reflect the spatial distribution pattern and evolution law of precipitation.
[0072] According to one aspect of the present application, the step S1 is further:
[0073] Step S11: Read the geographical information of the research area, extract the watershed data, and divide the research area into a sub-watersheds; collect and preprocess the research data of each sub-watershed in the research area to form time series data of precipitation information; a is a natural number greater than 2;
[0074] Step S12: Construct a set of trend analysis methods. For each sub-watershed, use the trend analysis methods one by one to perform trend tests and mutation tests on the time series data to obtain trend data and mutation point data, and divide the research period into b research periods according to the mutation point data; b is a natural number greater than 2.
[0075] In this embodiment, the precipitation evolution characteristics are not only reflected in the time course change, but also in the spatial difference. The meteorological conditions in different parts of the research area are different, so the precipitation characteristics also vary. According to the watershed, the sub-watersheds are divided, and each sub-watershed is used as a spatial unit to study the precipitation characteristics of each sub-watershed. There are two types of research objects that need to consider spatial differences. One is the spatial distribution of precipitation characteristics within each period, and the other is the spatial distribution of precipitation change characteristics between each period.
[0076] For example, there are a sub-watersheds in the research area, and the research period is divided into b periods. The precipitation characteristics of the j-th sub-watershed (j = 1, 2,..., a) in the i-th period (i = 1, 2,..., b) are denoted as S j i , and the precipitation characteristics of the j-th sub-watershed (j = 1, 2,..., a) between the i-th period and the (i + 1)-th period (i = 1, 2,..., b - 1) are denoted as ΔS j (i→i+1) . Then, the spatial distribution characteristics of precipitation in the i-th period (i = 1, 2,..., b) can be denoted as G i (S1 i , S2 i ,... S a i ), and the spatial distribution characteristics of the degree of precipitation change from the i-th period to the (i + 1)-th period (i = 1, 2,..., b - 1) can be denoted as H i,i+1 (ΔS1 (i→i+1) , S2 (i→i+1), ... S a (i→i+1) ). G i (i = 1, 2,..., b) and H i,i+1 (i = 1, 2,..., b - 1) are the research objects that need to consider spatial differences when analyzing the precipitation evolution law of the watershed.
[0077] According to one aspect of the present application, the step S11 is further as follows:
[0078] Step S11a: Obtain the digital elevation model data of the research area, and use the ArcGIS analysis tool to preprocess the digital elevation model until the quality of the digital elevation model data meets the standard;
[0079] Step S11b: Invoke the pre-configured hydrological analysis tool to extract watershed data from the digital elevation model data, determine the boundary and outlet points of the research area according to the terrain features and water flow directions, and generate a watershed layer; perform sub-watershed division on the watershed layer to generate a sub-watershed layer;
[0080] Step S11c: Construct a precipitation feature set for the research area, where the precipitation features include extremeness, intensity, rainfall amount, time, and space; collect precipitation data of each station in each sub-watershed of the research area, and use the GIS module or the pre-configured statistical analysis module for preprocessing;
[0081] Step S11d: Use the GIS module or the interpolation module to perform spatial interpolation on the precipitation data to generate a precipitation raster layer. Based on the precipitation raster layer, calculate the precipitation amount of each sub-watershed according to different time scales through the GIS or time series analysis module to form time series data of precipitation information.
[0082] According to one aspect of the present application, the step S1b is further:
[0083] Read the digital elevation model data, generate a research area map, aggregate adjacent and similar pixels in the research area map into one area, convert it into superpixels, and use it as the nodes of the neural network analysis module;
[0084] Determine the connection relationship between each node according to the spatial distance, color similarity, and gradient direction between the superpixels, and construct the edges of the neural network analysis module;
[0085] Calculate the new feature vectors according to the adjacency matrix and the feature vectors of the previous layer, so as to determine the updated feature vectors of each node and edge;
[0086] Construct the loss function and optimization method of the neural network analysis module to form the neural network analysis module, analyze and obtain the output result, and obtain the watershed layer.
[0087] According to one aspect of the present application, the step S12 is further:
[0088] Step S12a: Construct a set of trend analysis methods, where the trend analysis methods include MK, pre-whitened MK, detrended pre-whitened MK, variance-corrected MK, bootstrap MK, TS analysis method, and innovative trend analysis method;
[0089] Step S12b: For each sub - basin, each precipitation characteristic in the time - series data is subjected to trend test and mutation test one by one using each trend analysis method, forming a trend data set and a mutation point data set of precipitation characteristics;
[0090] Step S12c: Synthesize the trend data of each precipitation characteristic obtained by each method to generate the final trend data and mutation point data of each precipitation characteristic;
[0091] Step S12d: Segment the research period according to the mutation point data, and determine whether the sequence length of adjacent research periods is less than the threshold. If it is less than, perform cluster analysis on the precipitation characteristics of adjacent research periods. Decide whether to merge them into one research period according to whether the similarity of precipitation characteristics of adjacent research periods is greater than the threshold; perform trend analysis on each precipitation characteristic of each clustered research period to determine the trend direction and amplitude.
[0092] In this embodiment, seven methods, namely Mann - Kendall (MK), Pre - Whitening MK (PW - MK), Trend - free Pre - Whitening MK (TFPW - MK), Modified Mann - Kendall (MMK), Bootstrap MK, Theil - Sen, and Innovative Trend Analysis (ITA), are used to analyze the annual - scale change trends of the above - mentioned precipitation characteristic elements in the study area. The M - K test detects multiple precipitation mutation points. For a long enough precipitation time series, time - series analysis methods can sometimes detect multiple mutation points according to their change characteristics, and each stage shows different precipitation characteristics. The present invention uses the M - K test to perform trend test and mutation test on the above - mentioned precipitation elements respectively, and divides the research period into different time periods according to the mutation test of each element. Each element in each time period only contains one change trend state. When each sub - division stage is short (less than 10), cluster adjacent stages according to multiple elements to increase the sequence length of the sub - division time period, meet the minimum length requirement for the sequence change test of the sub - division time period, and at the same time reduce the unnecessary workload caused by too many sub - division time periods.
[0093] The step S12a is further as follows:
[0094] Obtain trend analysis methods and the applicability information of each trend analysis method;
[0095] Analyze the characteristics of the research data of each research period in each sub - basin one by one, establish the mapping relationship between the sub - basin, research period and trend analysis method based on the applicability information, and assign weights to the trend analysis methods for each period of each sub - basin;
[0096] Read the test data set, verify the mapping relationship, and obtain the trend analysis methods for each sub-basin in each research period, as well as the overall trend analysis method for the entire research period.
[0097] Based on the trend analysis methods for each sub-basin in each research period and the overall trend analysis method for the entire research period, construct a set of trend analysis methods.
[0098] In some studies, it is found that different methods have different usage conditions and scenarios. For example, TFPW-MK can only reflect the overall rainfall change trend and cannot directly reflect the trend changes of large or small rainfall amounts. ITA can detect the trend changes of low, medium, and high rainfall amounts simultaneously and reflect the rainfall change range, but this method cannot quantify the overall change trend. Affected by climate differences and geographical variability, the applicable methods for solving the problem of serial autocorrelation in different regions are also different. MMK can overcome the problem of the change in the dispersion degree of the variance distribution (S) caused by the correlation of time series. Therefore, if the same trend analysis method is used for different sub-basins and different research periods in the same research area, it will cause inaccurate analysis.
[0099] In order to more comprehensively analyze the precipitation change law, the above embodiments are given.
[0100] Specifically, in the above embodiment, first, the literature research method is adopted to collect various trend analysis methods, and list the usage scope, usage conditions, and data on advantages and disadvantages of each trend analysis method, that is, obtain its applicability information. On this basis, according to the above embodiment, the research area is divided into several research sub-basins, and the research period is divided into several research periods. For each period of each sub-basin, analyze the characteristics of the research data, and then establish the mapping relationship of the trend analysis methods for each research period of the sub-basin. For example, there are a total of 11 sub-basins, and each sub-basin has 4 research periods (in other embodiments, the research periods of each sub-basin can be different). Then, for each period of each sub-basin, establish a trend analysis method. For example, in the second period of the third sub-basin, there are 3 feasible trend analysis methods. Analyze the accuracy of each trend analysis method in this sub-basin and this research period according to the existing research data or through test data, and set weights for the trend analysis methods according to the accuracy calculation data. If the neural network method is used for weighting, the weights of each trend analysis method can be set as initial values, and then trained through the training data set to obtain the corresponding weights. In this embodiment, after obtaining the weights, verify the mapping relationship and weights through the test data set, determine the set of trend analysis methods corresponding to each research period of each sub-basin, and use the trend analysis method for analysis in the subsequent process.
[0101] It should be noted that since some trend analysis methods are more suitable for the analysis of overall trends, a trend analysis method is given for the entire research period to construct a set of overall trend analysis methods. Through the above process, a total set of trend analysis methods is formed.
[0102] In the subsequent analysis process, for each research period of each sub-basin, the analysis results are calculated by the weighting method and the analysis results are normalized, so as to give the final conclusion of the trend analysis. This avoids the defects of various analysis methods and the defect that the analysis results are different for different analysis methods in the same scenario, improving the accuracy of the analysis. The mapping relationship and weights are verified using a validation set, mainly to address the difference between the data in the current usage scenario and the data in the literature. The trend analysis method selected according to experience is not applicable to the current scenario, and there will be certain problems with its accuracy. Therefore, through calibration, the applicability and accuracy of the method can be improved.
[0103] In summary, in the above embodiments, first, according to the data characteristics of different sub-basins and research periods, the most suitable trend analysis method is selected to improve the analysis efficiency and accuracy. Second, by assigning weights, the importance of the trend analysis methods for different sub-basins and research periods in the overall trend analysis is reflected, increasing the flexibility and reliability of the analysis. Third, through the verification test data set, the rationality and effectiveness of the mapping relationship and weight settings are evaluated to optimize the analysis parameters and results. Finally, by constructing a set of trend analysis methods, the advantages of multiple trend analysis methods are comprehensively utilized to improve the analysis accuracy and robustness.
[0104] According to one aspect of the present application, in the above embodiments, the process of reading the test data set and verifying the mapping relationship specifically includes:
[0105] Collect the research data of a predetermined research period of a predetermined sub-basin as the standard data;
[0106] Based on the standard data, generate abnormal data through an abnormal data generation method;
[0107] Verify the trend analysis method using the standard data and the abnormal data.
[0108] The abnormal data generation method includes:
[0109] ①. Select the research data with different trends in another sub-basin and randomly insert it into the research data of the predetermined research period of the predetermined sub-basin.
[0110] ②. Divide the research data of the predetermined research period of the predetermined sub-basin into several segments and select one segment to be replaced with a constant.
[0111] ③. Select a segment of data from the research data of a predetermined study period in a predetermined sub-basin, reverse the order, and then re-insert it.
[0112] Through this method, not only can the applicability of the trend analysis method and the accuracy of trend calculation be improved, but also the flexibility of trend analysis can be enhanced, and it can be verified whether the relevant trend analysis method has the ability to detect outliers.
[0113] According to one aspect of the present application, the step S2 is further as follows:
[0114] Step S21: Read the boundary data of each sub-basin in the research area, calculate the centroid coordinates of each sub-basin, and generalize each sub-basin into a point;
[0115] Step S22: Select precipitation characteristics as the clustering index, perform standardization processing on each sub-basin to eliminate the dimensional differences between the indexes; use the hierarchical clustering method to cluster each sub-basin, select the Euclidean distance as the clustering criterion, select the longest distance method as the method for calculating the distance between clusters, draw a hierarchical clustering diagram, determine the number of categories, and obtain the sub-basins included in each category;
[0116] The basic idea of clustering is to first classify each sub-basin into a separate category, merge the two categories with the shortest distance into a new category, then calculate the distance between the new category and the remaining categories, and merge the two categories with the shortest distance into a new category again, and so on, until all are merged into one large category. When calculating the distance between categories, other methods that can be used include the shortest distance method, the longest distance method, the centroid method, the group average method, and the within-group sum of squares method.
[0117] The class Y composed of several sub-basins a and the class Y b The distance D ab can be calculated according to the following formula:
[0118] D ab = max (d ij ) X i ∈Y_a,X j ∈Y b ;
[0119] In the formula: X i and X j respectively represent the i-th and j-th sub-basins, which belong to the class Y a and the class Y b ; d ij represents the distance between X i and X j .
[0120] It can be seen that the distance D abThat is, it is the maximum value of the distances between sub-basins in two categories. Calculate the distances between categories in a loop according to the process and perform category merging. When the number of categories is 1, draw the hierarchical clustering diagram of each sub-basin, determine the number of categories in the final result according to the graph structure, and count the final results of the sub-basins included in each category to complete the hierarchical clustering.
[0121] Step S23: Use the R-O function method to analyze the spatial distribution pattern of sub-basins in each category. Select the distance scale range and interval, calculate the R-O function values of each category at different distance scales, draw the R-O function curves of random distribution and the actual R-O function curves, judge the spatial distribution types of each category, and obtain the position law of spatial distribution. The R-O function represents the combination of the Ripley L and O-ring functions. In other embodiments, the Ripley L or O-ring function can also be used alone, or the Ripley's K function can be used.
[0122] After hierarchical clustering of sub-basins, multiple categories with different precipitation characteristics are obtained, and each category contains several sub-basins with similar precipitation characteristics. Describe the spatial difference of precipitation in units of categories, perform spatial distribution pattern analysis on each category with different precipitation characteristics respectively, and then summarize and generalize the position law of spatial distribution based on its spatial distribution pattern.
[0123] Generalize each sub-basin into a point. Then the problem of the spatial distribution of basin precipitation can be abstracted into a point pattern problem, and the spatial distribution pattern is described by the method of point pattern analysis. First, according to the geometric shape of each sub-basin, determine its centroid position, generalize each sub-basin into a point at the centroid, and then select a suitable point pattern analysis method to perform spatial distribution pattern analysis on each type of sub-basin respectively.
[0124] Perform spatial distribution pattern analysis on each type of sub-basin based on the R-O function. Take the centroid point of the sub-basin as a point event, select several distance scales according to the basin area to calculate the corresponding R-O function values, draw the R-O function curves of random distribution and the actual R-O function curves, thereby judging the spatial distribution types of each type of sub-basin and summarizing the position law of spatial distribution.
[0125] It should be noted that in the above embodiments, the precipitation characteristics include extremity, intensity, rainfall amount, time, and space. For the precipitation process, the requirements are relatively strict. Actually, the precipitation process is restricted within a relatively small spatial range. It is found in the research that for the precipitation process in a certain area, using the above embodiments may result in trend analysis only for the data of the rainfall center station or the stations outside the rainfall center, and the results may deviate greatly. For the regional precipitation trend composed of multiple precipitation processes, the precipitation process of the whole region can be decomposed into several independent precipitation processes. Analyze each precipitation process and then superimpose them to obtain the overall trend analysis structure. For each precipitation process, the method provided in the following embodiments is used.
[0126] According to one aspect of the present application, the step S22 further includes:
[0127] Step S22a: Read the precipitation data of each station in each sub-basin of the study area, and analyze to obtain the precipitation center and precipitation boundary of each precipitation process;
[0128] Step S22b: Obtain the distances of each station within the precipitation boundary from the precipitation center and the distances from the precipitation boundary;
[0129] Step S22c: Take the precipitation data of the station closest to the precipitation center as the standard data, calculate the ratio of the precipitation data of each station within the precipitation boundary to the standard data, and draw a distribution change diagram of similar precipitation processes.
[0130] In the above embodiments, actually, through the existing precipitation data, the precipitation center and precipitation range data are obtained, so as to extract the areas with similar precipitation processes. In other words, in a certain precipitation process, there may be differences in intensity between the precipitation trend change process at the precipitation center and the precipitation change process in the peripheral area, but the evolution of the precipitation process is similar. Assuming that in some cases, it can be analogized to a three-dimensional Gaussian distribution. Therefore, the range covered by this precipitation process can be incorporated into the similar sub-areas to make up for the problem of inaccurate sub-area clustering. Of course, the actual precipitation distribution is quite complex.
[0131] If there are multiple precipitation processes, they can be decomposed into independent precipitation processes and then superimposed in time and space. For example, if several sub-basins experience two precipitation processes simultaneously within a certain period of time, the proportion in each precipitation process can be obtained through precipitation data analysis and then superimposed. The total superimposed data should be the same as the precipitation trend of the sub-basin. Therefore, by collecting data from the precipitation center and surrounding similar sub-regions and clustering them, more accurate results can be obtained. In other words, if the result after superposition is different from the result calculated by the overall trend, it indicates that there are some unexpected problems in data collection or other influencing factors. Specifically, thresholds can be set according to the distance from the center point and the similarity of precipitation data, so that sub-basins with similar precipitation processes are grouped into one category. In the subsequent clustering process, the accuracy can be improved, or by comparing with the clustering results, it can be judged whether there are sub-basins with inconsistent trend analysis, and the trend analysis results of the sub-basins can be verified again.
[0132] According to one aspect of the present application, the step S3 is further as follows:
[0133] Step S31: Adopt the meta-analysis method to collect the driving factors of precipitation change in the study area and construct a set of driving factors for precipitation change;
[0134] Adopt meta-analysis to preliminarily screen the driving factors of precipitation change from existing literature, reports, historical records and other materials from aspects such as generation mechanism, statistical law, and indirect influence to form a set of driving factors for precipitation change.
[0135] Meta-analysis refers to using statistical methods to analyze and summarize multiple research data collected to provide a quantitative average effect to answer research questions. Its advantage is to increase the credibility of the conclusion by increasing the sample size and solve the inconsistency of research results. Meta-analysis is a systematic and quantitative comprehensive analysis of the results of multiple independent studies on the same topic. The main steps for mining and preliminary screening of indicators are as follows:
[0136] (1) Search for relevant literature: Search for articles on the impact and adaptation of orderly flow on flood control safety, water supply safety, and water environment ecological safety in retrieval databases such as CNKI and Web of Science; screen the literature, and select articles with a quantitative response relationship and from which the mean, standard deviation, and number of samples can be obtained for subsequent analysis.
[0137] (2) Calculate the effect value: Extract the indicators that can be used for meta-analysis from the collected data, select the logresponse ratio to calculate the effect value, use the maximum likelihood model to analyze the variation between different articles, and use the maximum likelihood model and the random model to calculate the total variation value.
[0138] (3)Statistical analysis: The reliability of the data was detected by calculating publication bias and performing sensitivity analysis. All calculations were completed using the Metafor package in R language.
[0139] Step S32: For each research period in the research area, a GBDT model optimized by GA hyperparameters was constructed, with the precipitation change driving factors as the input and the precipitation parameters as the output, to simulate the precipitation in the research area, obtain the sensitive factors of the precipitation change in the research area, and use the sensitive factors as the key driving factors.
[0140] Step S33: A comprehensive discrimination criterion was constructed to quantify and rank the key driving factors in descending order.
[0141] The precipitation change driving factors include the Northern Hemisphere Polar Vortex Area Index (NHPVA), the Northern Hemisphere Polar Vortex Intensity Index (NHPVI), the Northern Hemisphere Polar Vortex Central Meridional Position Index (NHPVCLON), the Northern Hemisphere Polar Vortex Central Zonal Position Index (NHPVCLAT), the Western Pacific Subtropical High Area Index (WPSHA), the Western Pacific Subtropical High Intensity Index (WPSHI), the Western Pacific Subtropical High Ridge Line Position Index (WPSHRP), the Western Pacific Subtropical High Westward Extension Ridge Point Index (WPSHWRP), the Western Pacific Subtropical High North Boundary Position Index (WPSHNBP), the Eurasian Zonal Circulation Index (EZC), the Eurasian Meridional Circulation Index (EMC), the East Asian Trough Position Index (EATP), the East Asian Trough Intensity Index (EATI), the Nino1+2 Sea Surface Temperature Index, the Nino3 Sea Surface Temperature Index, the Nino4 Sea Surface Temperature Index, the Nino3.4 Sea Surface Temperature Index, trough (T), ridge (R), upper-level jet (HJ), shear line (SL), low vortex (VO), cyclone (CL), front (FS), typhoon (TY), cold air (CA), low-level jet (LJ).
[0142] In this embodiment, since there are many driving factors dug out, some driving factors have high correlations, and the correlations of the indicators will affect the prediction results: Since there are inevitably different degrees of correlations among the driving factors dug out in the preliminary screening step and they cannot meet the requirement of mutual independence of the driving factors, and there is a large amount of redundancy and interference in the information reflected by the highly correlated driving factors, which linearly enhances or weakens the trend of the same precipitation element, easily leading to distortion of the prediction results. On the other hand, the scale of the driving factors also grows exponentially with the increase in the number of precipitation elements, increasing the complexity of the prediction and even possibly resulting in the "curse of dimensionality" making the prediction difficult to implement. Therefore, considering that the key driving factors of precipitation change in different research areas are not the same, in this step, based on the historical precipitation data of the research area, a GBDT model optimized by GA hyperparameters was used to perform a secondary screening on the precipitation change driving factors dug out in the previous step to further identify the key driving factors of precipitation change in the research area.
[0143] The deviation between the output calculated by the model and the true result of the sample is called the training error, which affects the accuracy of the model. How to improve the model accuracy so that the model prediction result is as close as possible to the true value is an important part of building the model. Without changing the model structure, parameter optimization is the main means to reduce the deviation and improve the accuracy of the model.
[0144] To obtain a good simulation effect, it is often necessary to adjust the parameters to make the model learn the general laws hidden in the samples as much as possible. However, the model cannot judge which laws are common to all data and which laws are unique to the samples. If the learning is too poor, the model will not be able to grasp the overall law and cause the underfitting problem; if the learning is too good, the model will regard the characteristics of the samples themselves as general laws and cause the overfitting problem.
[0145] The generalization ability is an index to evaluate the goodness of model simulation. The most ideal way is to evaluate the generalization errors of different models (referring to GBDT models with different parameters), and then select the parameters with the smallest generalization error. In machine learning, in order to obtain the generalization error before testing, the sample set is usually divided into two parts, the training set and the validation set. The training set is used to train the model. After the model is built, the independent variables of the validation set data are used as inputs to obtain the model prediction results, and then the validation error between the prediction results and the actual outputs of the validation set is calculated. If the validation samples are randomly and independently sampled from the overall sample, then the validation set and the overall sample follow the independent and identical distribution, and the validation error of the validation set can be approximately regarded as the generalization error. In order to make the validation error of the model on the validation set as representative of the generalization error as possible, it is particularly important to select a suitable validation set.
[0146] Cross-validation divides the sample set D into k mutually exclusive subsets D1, D2, …, D k , where D1 ∪ D2 ∪ … ∪ D k = D, and for any D i ∩ D j = empty set (i ≠ j). Each time, k - 1 of them are selected as the training set, and the remaining one is used as the validation set. After k cycles, k validation errors are obtained, and the mean of all validation errors is taken as the representation of the generalization ability. If the number of subsets divided is too small, the training set will be very small and the model cannot represent the overall sample; if the number of subsets divided is too large, the computational amount will increase, and at the same time the validation set will become smaller, resulting in a worse approximation of the generalization ability. Usually k is taken as 10, and at this time it is also called 10-fold cross-validation.
[0147] GBDT mainly includes two parts of parameters: Boosting framework parameters and weak learner parameters. The former includes the maximum number of iterations of the weak learner, the weight shrinkage coefficient of each weak learner, and the loss function, etc.; the latter includes the maximum depth, the minimum number of samples in the leaf nodes, etc. This technology mainly focuses on two sensitive parameters: the maximum number of iterations of the weak learner and the weight shrinkage coefficient of the weak learner (also known as the learning rate). Too many weak learners may lead to overfitting, while too few will result in underfitting. Generally, a moderate number is selected. Regularization can effectively reduce model overfitting, and the learning rate is the result of considering regularization. Selecting an appropriate learning rate can reduce model overfitting and improve accuracy. According to the GBDT algorithm process, it is easy to see that a smaller learning rate means more weak learners, and the two restrict each other. Therefore, usually, these two parameters, the maximum number of weak learners and the learning rate, are optimized and adjusted together.
[0148] According to the GA algorithm principle, the maximum number of weak learners Estimators and the learning rate Rate are used as target variables and input, discretized within a certain range to form a solution space; the fitness function selects the generalization error of the model, and the average validation error obtained by cross-validation is used for approximation. After looping a certain number of times, Estimators and Rate tend to stable values. At this time, the model parameters reach the optimal state, and the model has the strongest generalization ability.
[0149] According to one aspect of the present application, step S4 is further as follows:
[0150] Step S41: Based on the values of k key driving factors within each sub-basin as the coordinate values of the sub-basin in the k-dimensional space, calculate the distances between each sub-basin through Euclidean distance, and cluster the sub-basins; k is a natural number.
[0151] To analyze the correlation between precipitation and influencing factors, it is first necessary to determine the spatial distribution pattern of each influencing factor, specifically including the characteristics of the influencing factors within each time period and the spatial distribution of the change characteristics of the influencing factors between each time period. Taking the sub-basin as the unit, cluster the sub-basins according to the characteristics of the influencing factors, and then describe the spatial distribution characteristics of each influencing factor in combination with the spatial location of the sub-basins. To reduce the workload of the spatial distribution analysis of the influencing factors, a generalization method is adopted to cluster all the sub-basins in the basin according to the characteristics of the influencing factors, and describe the spatial distribution characteristics of the influencing factors in the form of classes.
[0152] When clustering based on the characteristics of the influencing factors within a time period, abstract the values of each factor as points in the Euclidean space, and use this as the characteristic index. Assume that a total of n types of factors are selected for analysis. According to the i th sub-basin n values of x i1, xi2, …, x in , consider this sub - watershed as n a point in X i - dimensional space x i1 , x i2 ,…, x in ). Using the position of the point as the characteristic index of the sub - watershed and the Euclidean distance between points as the clustering criterion, perform sub - watershed clustering. The processing method for underlying surface factors is similar. Among them, the position of the point in Euclidean space is determined by the area proportion of each land - use type within the sub - watershed.
[0153] When clustering based on the change characteristics of influencing factors between time periods, abstract the change values of each factor as vectors in Euclidean space and use them as characteristic indices. Taking a factor as an example for illustration: Suppose a total of n types of factors are selected for analysis. According to the method described above, determine the positions of the i th sub - watershed in the n - dimensional space before and after the change, denoted as point X i ( x i1 , x i2 ,…, x in )and point Y i ( y i1 , y i2 ,…, y in ). Construct a X i - dimensional vector Y i pointing from point n - Z - t = ( y i1 - x i1 , y i2 - x i2 ,…, y in - x in ); To uniformly compare the vectors corresponding to each sub - watershed, vector Z - tTranslate it to the position starting from the origin of coordinates, so the vector Z - t The end point of can be denoted as point Z i = ( y i1 - x i1 , y i2 - x i2 ,…, y in - x in ); For each sub - basin, uniformly take the origin of coordinates as the starting point of the vector, take the position of the end point of the vector as the characteristic index of the sub - basin, and take the Euclidean distance between each end point as the clustering criterion to conduct sub - basin clustering.
[0154] When conducting the attribution analysis of the evolution of basin precipitation, it focuses on the exploration of spatial differences. First, test the correlation between the spatial distribution of precipitation and the spatial distribution of influencing factors. If the degree of correlation is relatively high, then combined with the spatial distribution pattern of influencing factors, aiming at the spatial differences existing in the precipitation evolution process, conduct attribution analysis on the precipitation evolution of different sizes and different degrees of change. When studying the correlation between two variables, it is divided into two cases: for quantitative variables, the method of regression analysis can be used, such as scatter plot drawing, correlation coefficient calculation, and residual analysis; for categorical variables, the method of independence test can be used. This technology conducts clustering processing on precipitation and its influencing factors respectively, and uniformly analyzes them in the form of classes. So precipitation and its influencing factors are both regarded as categorical variables, so the independence test method is used to analyze their correlation.
[0155] When analyzing the correlation between the spatial distribution of precipitation and influencing factors, conduct independence tests based on the static distribution pattern within each time period and the dynamic change pattern between each time period respectively. Take all the sub - basins in the basin as samples, take precipitation and influencing factors as categorical variables, and assume that the research period contains m time periods. The objects of the independence test include:
[0156] (1) The independence test of the precipitation characteristics and influencing factor characteristics within the i - th time period (i = 1, 2,..., m);
[0157] (3) The independence test of the precipitation change characteristics and influencing factor change characteristics between the i - th time period and the (i + 1) - th time period (i = 1, 2,..., m - 1);
[0158] According to the results of the above 2m - 1 groups of independence tests, a comprehensive evaluation can be made on the degree of correlation between the spatial distribution of precipitation and influencing factors.
[0159] Step S42: Examine the correlation between the spatial distribution of precipitation and that of key driving factors. Use the driving factor - precipitation mapping relationship model to simulate the degree of precipitation change under the individual or combined action of each key driving factor. Cluster each sub - basin according to the degree of precipitation change, and calculate the proportion of each type affected by the key driving factors.
[0160] Step S43: If the degree of correlation is higher than the threshold, then, in combination with the spatial distribution of key influencing factors, conduct an attribution analysis of the precipitation evolution with spatial differences during the precipitation evolution process for different sizes and different degrees of change.
[0161] Referring to the results of the independence test, determine the degree of correlation between the spatial distribution of precipitation and influencing factors. Take the factors with high correlation as influencing factors and conduct an attribution analysis of precipitation evolution. First, analyze the causes of the spatial distribution pattern of precipitation evolution characteristics, determine the dominant factors of precipitation evolution characteristics at different spatial positions, and then conduct an attribution analysis of precipitation evolution for each size and each degree of change.
[0162] To analyze the influence degree of influencing factors on the spatial distribution of precipitation evolution characteristics, it is first necessary to determine the precipitation evolution characteristics of each sub - basin under the individual influence of each factor. Consistent with the above, assume that the entire basin is divided into n sub - basins in total, and the research period contains m time periods in total. Denote the precipitation of the j - th sub - basin (j = 1, 2, …, n) in the i - th time period (i = 1, 2, …, m - 1) as R j i , and the precipitation in the (i + 1) - th time period (i = 1, 2, …, m - 1) as R j i+1 . Keep the conditions of other driving factors in the i - th time period unchanged, replace the X - th factor with the corresponding value in the (i + 1) - th time period, substitute it into the driving factor - precipitation mapping relationship model to simulate precipitation, and denote the precipitation of the j - th sub - basin as R C,j i+1 ; keep the X - th factor of the basin unchanged in the i - th time period, replace the underlying surface conditions with the corresponding values in the (i + 1) - th time period, substitute it into the driving factor - precipitation mapping relationship model to simulate precipitation, and denote the precipitation of the j - th sub - basin as R L,j i+1 . Calculate the degree of precipitation change ΔR j i of the j - th sub - basin from the i - th time period to the (i + 1) - th time period, and the degree of precipitation change ΔR C,j i caused by the change of the X - th factor.
[0163] ΔR j i =(R j i+1 -Rj i ) / (R j i ) (i = 1, 2, …, m - 1; j = 1, 2, …, n),
[0164] ΔR C,j i =(R C,j i+1 -R j i ) / (R j i ) (i = 1, 2, …, m - 1; j = 1, 2, …, n),
[0165] Other influencing factors can be analogized.
[0166] According to ΔR j i 、ΔR C,j i and the precipitation change degree values caused by other influencing factors, all sub - basins are clustered respectively, and the same number of clusters is used for the three - time clustering. Let the number of clusters be k. Then, according to the ΔR j i values, ΔR C,j i values and the precipitation change degree values caused by other influencing factors of each sub - basin, each sub - basin is classified into k classes with different precipitation evolution degrees.
[0167] Introduce the concept of precipitation evolution type number. Define the precipitation evolution type numbers of the k classes with different precipitation evolution degrees as positive integers with values from 1 to k. The precipitation evolution type numbers of each class are recorded as 1 to k in ascending order of the ΔR value. The precipitation evolution type number of the class with the largest precipitation reduction degree is 1, and the precipitation evolution type number of the class with the largest precipitation increase degree is k.
[0168] Record the precipitation evolution type number of each type of sub - basin clustered according to the ΔR j i value as Num, and record the precipitation evolution type number of each type of sub - basin clustered according to the ΔR C,j i value as NumC. For the precipitation change of the j - th sub - basin from the i - th time period to the i + 1 - th time period: NumC j i is closer to the value of Num j i , the more similar the precipitation evolution degree of the sub - basin is to the precipitation evolution degree caused by the change of factor X, that is, the greater the influence of factor X on the sub - basin in the spatial distribution pattern of precipitation evolution. The judgment of the dominant influencing factor is carried out according to the following formula:
[0169] θ j i =|NumC j i -Num j i | (i = 1, 2, …, m - 1; j = 1, 2, …, n)
[0170] In the formula: θ j i is the dominant factor indicator variable of precipitation change in the j-th sub-basin from the i-th period to the (i + 1)-th period; the meanings of the remaining variables are the same as above.
[0171] According to the calculation results of θ j i (i = 1, 2, …, m - 1; j = 1, 2, …, n), the spatial distribution characteristics of the evolution characteristics of basin precipitation between periods can be attributed, taking the sub-basin as the research unit to clarify the dominant influencing factors of the characteristics presented in the basin precipitation evolution pattern.
[0172] According to one aspect of the present application, the process of attributing the evolution of precipitation of different magnitudes specifically includes:
[0173] In the i th period, the precipitation characteristics of each sub-basin are divided into l categories, and the difference between each category lies in the different precipitation amounts. The attribution analysis of the characteristics presented in the evolution pattern is carried out for the l categories of precipitation respectively. Taking the area affected by the dominant influence as the measurement index, the influence degree of the influencing factor on it is determined.
[0174] ,
[0175] δ j,p i = 1, the rainfall magnitude in the j-th sub-basin in the i-th period does not belong to the p-th category;
[0176] δ j,p i = 0, the rainfall magnitude in the j-th sub-basin in the i-th period does not belong to the p-th category.
[0177] η j,C i = 1, the rainfall evolution in the j-th sub-basin from the i-th period to the (i + 1)-th period is dominated by factor X.
[0178] η j,C i= 0, the rainfall evolution of the j-th sub-basin from the i-th period to the i+1-th period is not dominated by factor X.
[0179] Where: α i C,p is the proportion of precipitation in the i period within the p category size level that is dominated by the key influencing factor among meteorological factors in the subsequent precipitation evolution process; A j is the j area of the j-th sub-basin; n is the number of sub-basins in the basin; m is the number of periods included in the study period; l is the i number of sub-basin clusters based on precipitation size levels within the i-th period, i = 1, 2, …, m - 1, p = 1, 2, …, l.
[0180] In this embodiment, there are spatial differences in the precipitation characteristics of the basin within each period, and the precipitation amounts of each sub-basin belong to different size levels. The dominant influencing factors for the evolution of precipitation at each level are not the same during the period-to-period evolution. Therefore, the attribution analysis of precipitation evolution for different size levels is carried out separately.
[0181] According to one aspect of the present application, the process of attribution analysis for precipitation evolution with different degrees of change specifically includes:
[0182] From the i i-th period to the i i + 1-th period, the precipitation change characteristics of each sub-basin are divided into k q categories. The difference between each category lies in the different degrees of precipitation change. Attribution analysis is carried out for the characteristics presented in the evolution pattern of the k q categories of precipitation. Using the area affected by the dominant influence as a measurement index, the influence degree of the influencing factor on it is determined.
[0183] ,
[0184] δ j,q i = 1, the degree of rainfall change of the j-th sub-basin from the i-th period to the i+1-th period belongs to the q-th category;
[0185] δ j,q i = 0, the degree of rainfall change of the j-th sub-basin from the i-th period to the i+1-th period does not belong to the q-th category;
[0186] η j,C i= 1, the rainfall evolution of the j-th sub-basin from the i-th time period to the i+1-th time period is dominated by factor X;
[0187] η j,C i = 0, the rainfall evolution of the j-th sub-basin from the i-th time period to the i+1-th time period is not dominated by factor X.
[0188] β i C,q For the rainfall from the i time period to the i +1 time period, the proportion of the precipitation with the q -th degree of change being dominated by a certain key influencing factor during its evolution; A j is the j area of the j-th sub-basin; n is the number of sub-basins in the basin; m is the number of time periods included in the research period; k For the rainfall from the i time period to the i +1 time period, the number of sub-basin clusters based on the degree of precipitation change, i = 1, 2, …, m - 1, q = 1, 2, …, k.
[0189] In this embodiment, there are also spatial differences in the characteristics of basin rainfall evolution between time periods, and the degrees of precipitation change in each sub-basin belong to different levels. The main influencing factors for the precipitation evolution with different degrees of change are not the same, so the attribution analysis of precipitation evolution with different degrees of change is carried out separately.
[0190] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.
Claims
1. A method for analyzing the spatio-temporal variation of multiple precipitation elements and identifying their attribution under climate change, characterized in that, Including the following steps: Step S1: Define the scope of the research area and the scope of the research period, collect research data, and extract precipitation information and the mapping relationship between clustering and sub-basins. Each research area includes at least two sub-basins, and each research period includes at least two research time periods. Step S2: Cluster based on the precipitation characteristics of each sub-basin, read each sub-basin in the research area and generalize it into a point, construct a point pattern analysis method, and conduct a spatial distribution pattern analysis. Step S3: Construct a precipitation change driving factor set and a GBDT model optimized by GA hyperparameters. For each research period in the research area, screen sensitive precipitation change driving factors as key driving factors. Step S4: Conduct a cluster analysis on the sub-basins in the research area based on the precipitation change driving factors as key driving factors, obtain the spatial distribution characteristics of each precipitation change driving factor, and conduct an attribution analysis on the evolution of the precipitation spatial distribution pattern. The further steps of Step S4 are as follows: Step S41: Based on the values of k key driving factors within each sub-basin as the coordinate values of the sub-basin in the k-dimensional space, calculate the distances between each sub-basin through the Euclidean distance, and cluster the sub-basins. k is a natural number. Step S42: Examine the correlation between the precipitation spatial distribution and the spatial distribution of the key driving factors, use the driving factor and precipitation mapping relationship model to simulate the precipitation change degree under the individual or combined action of each key driving factor, cluster each sub-basin according to the precipitation change degree, and calculate the proportion of each category affected by the key driving factors. Step S43: If the correlation degree is higher than the threshold, then in combination with the spatial distribution of the key influencing factors, conduct an attribution analysis on the precipitation evolution with spatial differences during the precipitation evolution process for different sizes and different change degrees of precipitation evolution. The process of conducting an attribution analysis on the precipitation evolution of different sizes specifically includes: In the i precipitation characteristics of each sub-basin during the l time periods are divided into l categories. The difference between each category lies in the amount of precipitation. Attribution analysis is carried out on the characteristics presented in the evolution pattern of each l category of precipitation. Taking the area affected by the dominant influence as a measurement index, the influence degree of the influencing factors on it is determined. , Wherein: α i C,p is the proportion of precipitation in the i time period of the p th class size level that is dominated by the key influencing factors among meteorological factors in the subsequent precipitation evolution process; A j is the j area of the n th sub - basin; m is the number of sub - basins in the basin; l is the i number of sub - basin clusters based on precipitation size levels in the th time period, i = 1, 2, …, m - 1, p = 1, 2, …, l; The process of conducting an attribution analysis on the precipitation evolution of different change degrees specifically includes: From the i th time period to the i +(1)th time period, the precipitation change characteristics of each sub-basin are divided into k categories. The difference between categories lies in the different degrees of precipitation change. Attribution analysis of the characteristics presented in the evolution pattern is carried out for k categories of precipitation respectively. Taking the area affected by the dominant influence as a metric, the influence degree of the influencing factors on it is determined. , β i C,q For the proportion of precipitation with a certain degree of change in class during its evolution process from the i time period to the i +1 time period that is dominated by a certain key influencing factor; A q class j For the j area of the n sub-watershed; m For the number of time periods included in the study period; k For the number of sub-watershed clusters based on the degree of precipitation change from the i time period to the i +1 time period, where i = 1, 2, …, m - 1 and q = 1, 2, …, k.
2. The spatio-temporal variation analysis and attribution identification method of precipitation multi-elements under climate change according to claim 1, characterized in that The further steps of Step S1 are as follows: Step S11: Read the geographical information of the research area, extract the watershed data, and divide the research area into a sub-basins. Collect the research data of each sub-basin in the research area and preprocess it to form time series data of precipitation information. a is a natural number greater than 2. Step S12: Construct a set of trend analysis methods. For each sub-basin, use the trend analysis method one by one to conduct trend tests and mutation tests on the time series data, obtain trend data and mutation point data, and divide the research period into b research time periods according to the mutation point data. b is a natural number greater than 2.
3. The method for analyzing the spatio-temporal variation and attribution identification of multiple precipitation elements under climate change according to claim 2, characterized in that, The further steps of Step S11 are as follows: Step S11a: Obtain the digital elevation model data of the research area, and use the ArcGIS analysis tool to preprocess the digital elevation model until the quality of the digital elevation model data meets the standard. Step S11b: Call the pre-configured hydrological analysis tool to extract the watershed data from the digital elevation model data, determine the boundary and outlet points of the research area according to the terrain characteristics and water flow direction, generate a watershed layer. Conduct sub-basin division on the watershed layer to generate a sub-basin layer. Step S11c: Construct a precipitation feature set for the study area, where the precipitation features include extremity, intensity, rainfall amount, time, and space; collect precipitation information of each station in each sub-basin of the study area, and perform preprocessing using the GIS module or a pre-configured statistical analysis module; Step S11d: Perform spatial interpolation on the precipitation information using the GIS module or an interpolation module to generate a precipitation raster layer. Based on the precipitation raster layer, calculate the precipitation amount of each sub-basin according to different time scales using the GIS module or a time series analysis module to form time series data of the precipitation information.
4. The method for spatio-temporal variation analysis and attribution identification of multiple precipitation elements under climate change according to claim 3, characterized in that The step S11b is further as follows: Read the digital elevation model data to generate a study area map, aggregate adjacent and similar pixels in the study area map into one area, convert it into superpixels, and use it as the nodes of the neural network analysis module; Determine the connection relationship between each node according to the spatial distance, color similarity, and gradient direction between the superpixels, and construct the edges of the neural network analysis module; Calculate the new feature vectors according to the adjacency matrix and the feature vectors of the previous layer, and thus determine the updated feature vectors of each node and edge; Construct the loss function and optimization method of the neural network analysis module to form the neural network analysis module, analyze and obtain the output result, and obtain the watershed layer.
5. The method for analyzing the spatio-temporal variation and attribution identification of multiple precipitation elements under climate change according to claim 3, characterized in that, The step S12 is further as follows: Step S12a: Construct a set of trend analysis methods, where the trend analysis methods include MK, pre-whitened MK, detrended pre-whitened MK, variance-corrected MK, bootstrap MK, TS analysis method, and innovative trend analysis method; Step S12b: For each sub-basin, use each trend analysis method to perform trend test and mutation test on each precipitation feature in the time series data one by one to form a trend data set and a mutation point data set of the precipitation features; Step S12c: Synthesize the trend data of each precipitation feature obtained by each method to generate the final trend data and mutation point data of each precipitation feature; Step S12d: Segment the study period according to the mutation point data, and judge whether the sequence length of adjacent study periods is less than the threshold. If it is less than, perform cluster analysis on the precipitation features of adjacent study periods, and decide whether to merge them into one study period according to whether the similarity of the precipitation features of adjacent study periods is greater than the threshold; Perform trend analysis on each precipitation feature of each clustered study period to determine the trend direction and amplitude.
6. The spatio-temporal variation analysis and attribution identification method for multiple precipitation elements under climate change according to claim 5, characterized in that The step S2 is further as follows: Step S21: Read the boundary data of each sub-basin in the study area, calculate the centroid coordinates of each sub-basin, and generalize each sub-basin into a point; Step S22: Select precipitation features as clustering indicators, perform standardization processing on each sub-basin to eliminate the dimensional difference between the indicators; use the hierarchical clustering method to cluster each sub-basin, select the Euclidean distance as the clustering criterion, select the complete linkage method as the method for calculating the distance between clusters, draw the hierarchical clustering diagram, determine the number of categories, and obtain the sub-basins included in each category; Step S23: Use the R-O function method to analyze the spatial distribution pattern of sub-watersheds in each category. Select the distance scale range and interval, calculate the R-O function values of each category at different distance scales, plot the R-O function curves of random distribution and the actual R-O function curves, judge the spatial distribution type of each category, and obtain the position law of spatial distribution.
7. The method for analyzing the spatio-temporal variation and attribution identification of multiple precipitation elements under climate change according to claim 6, characterized in that The said step S3 is further as follows: Step S31: Use the meta-analysis method to collect the driving factors of precipitation change in the study area and construct a set of driving factors for precipitation change. Step S32: For each study period in the study area, construct a GBDT model optimized by GA hyperparameters. Take the driving factors of precipitation change as input and precipitation parameters as output to simulate the precipitation in the study area, obtain the sensitive factors of precipitation change in the study area, and use the sensitive factors as key driving factors. Step S33: Construct a comprehensive discrimination criterion to quantify and rank the key driving factors in descending order.
8. The method for analyzing the spatio-temporal variation and attribution identification of multiple precipitation elements under climate change according to claim 5, characterized in that The said step S12a is further as follows: Obtain the trend analysis method and the applicability information of each trend analysis method. Analyze the characteristics of the research data of each sub-watershed in each study period one by one. Based on the applicability information, establish the mapping relationship among the sub-watershed, the study period and the trend analysis method, and assign weights to the trend analysis methods in each period of each sub-watershed. Read the test data set to verify the mapping relationship, obtain the trend analysis methods for each sub-watershed in each study period, and the overall trend analysis method for the study period. Based on the trend analysis methods for each sub-watershed in each study period and the overall trend analysis method for the study period, construct a set of trend analysis methods.
9. The method for analyzing the spatio-temporal variation and attribution identification of multiple precipitation elements under climate change according to claim 6, characterized in that, The said step S22 also includes: Step S22a: Read the precipitation data of each station in each sub-watershed of the study area and analyze to obtain the precipitation center and precipitation boundary of each precipitation process. Step S22b: Obtain the distances of each station located within the precipitation boundary from the precipitation center and from the precipitation boundary. Step S22c: Take the precipitation data of the station closest to the precipitation center as the standard data, calculate the ratio of the precipitation data of each station within the precipitation boundary to the standard data, and draw a distribution change map of similar precipitation processes.
Citation Information
Patent Citations
Regional evapotranspiration change attribution analysis method considering multi-element spatial dependency
CN114462518A
Rainfall intensity estimation method integrating multi-temporal-spatial-scale Doppler radar data
CN114742206A