Rainfall station optimization layout method based on precipitation time sequence and geospatial feature aggregation

Through the optimization method based on precipitation timing and geospatial feature aggregation, the problem that rainfall station network layout in the prior art is difficult to fully consider the similarity of precipitation process and geographically spatial similarity, and more efficient and high-quality rainfall information acquisition and early warning of precipitation events in the river basin are achieved.

CN120069160APending Publication Date: 2025-05-30CHINA YANGTZE POWER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510030237.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing rainfall station optimization layout method is difficult to fully consider the similarity of the precipitation process and the spatial similarity of geographical elements, resulting in insufficient monitoring coverage and data accuracy.

Method used

The optimization layout method based on precipitation timing and geospatial feature aggregation is adopted, and the rainfall site layout is scientifically optimized through steps such as data preparation and processing, feature distance matrix calculation, feature aggregation analysis and optimization layout.

Benefits of technology

The efficiency and quality of basin rainfall information acquisition has been improved, the ability to obtain precipitation information in different regions and time periods has been enhanced, and the basin’s early warning and response capabilities for precipitation events have been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069160A_ABST
    Figure CN120069160A_ABST
Patent Text Reader

Abstract

The invention discloses a precipitation station optimization layout method based on precipitation time sequence and geospatial feature aggregation. The method comprises the following steps: S1, data preparation and processing; s2, calculating a feature distance matrix; s3, feature aggregation analysis; s4, optimizing the layout; according to the method, the similarity of the rainfall time sequence data in rainfall and process is fully considered on the basis of considering the rainfall cause and the spatial similarity relationship of geographic elements aiming at the optimized layout requirements of the rainfall sites, so that the optimized layout of the rainfall sites of the drainage basin is realized, and the efficiency and the quality of obtaining the rainfall information of the drainage basin are improved to the maximum extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hydrology and water resources, in particular to a method for optimizing the layout of rain gauges based on the aggregation of precipitation time series and geospatial characteristics. Background Art

[0002] Accurately obtaining the water regime of a basin from the monitoring data of rain gauge stations is crucial for improving the comprehensive utilization efficiency of basin water resources and the hydrological forecasting efficiency. However, with the increasing social demands and the continuous changes in natural conditions, the layout requirements of rain gauge stations have also changed. In some basins, due to monitoring blind spots or the improvement of accuracy requirements, it is urgent to add rain gauge stations to enhance monitoring coverage and data accuracy; in some basins, it is necessary to optimize the existing rain gauge stations to improve monitoring efficiency and resource utilization rate.

[0003] The optimized layout of a rain gauge network aims to maximize the acquisition efficiency and quality of rain gauge information by streamlining the number of stations while ensuring the accuracy of regional rain gauge data collection. When optimizing the layout of rain gauge stations currently, considerations are mainly made from the perspective of network density and station representativeness. Network density can provide data guarantee for the accurate presentation of areal rainfall, but it is not that the greater the density, the better, and cost-benefit analysis needs to be considered. The representativeness of a station involves the selection of the station layout location. The current representativeness is mainly characterized by comparing the rainfall differences, correlations, or mutual information entropy between adjacent stations. However, only focusing on rainfall differences will ignore the comparison in the change process. In addition, the time series precipitation data of different stations may only have a front-back displacement on the time axis. After restoring the displacement, the two time series are consistent. However, correlations or information entropy can only be analyzed with the precipitation data time series strictly aligned, which will lose some similar information in the change process.

[0004] The optimized layout of a rain gauge network is a complex multi-objective decision-making process. How to fully consider the similarity of precipitation processes, simultaneously consider precipitation causes, take into account the spatial similarity relationship of geographical elements, and scientifically and reasonably optimize the layout design of the rain gauge network is an important link for making full use of the information of hydrometeorological station networks and deepening the hydrometeorological coupled forecasting of downstream basins. Summary of the Invention

[0005] The purpose of the present invention is to overcome the above deficiencies and provide a method for optimizing the layout of rain gauges based on the aggregation of precipitation time series and geospatial characteristics. Aiming at the optimization layout requirements of precipitation stations, on the basis of considering precipitation causes and taking into account the spatial similarity relationship of geographical elements, the similarity of precipitation time series data in terms of rainfall amount and process is fully considered to realize the optimized layout of basin rain gauge stations, and maximize the acquisition efficiency and quality of basin rain gauge information.

[0006] To solve the above technical problems, the technical solution adopted by the present invention is: an optimized layout method for rain gauges based on precipitation time series and geographical space feature aggregation, which includes the following steps:

[0007] S1. Data preparation and processing;

[0008] S2. Calculation of feature distance matrix;

[0009] S3. Feature aggregation analysis;

[0010] S4. Optimized layout.

[0011] Further, the specific steps of S1 include:

[0012] S1-1. Data preparation;

[0013] S1-2. Generation of sample point grid;

[0014] S1-3. Smoothing processing of digital elevation data corresponding to the basin and precipitation fusion product data.

[0015] Furthermore, the specific steps of S1-1 are as follows: The data required for the optimized layout of precipitation stations includes basic basin geographical information data and precipitation fusion product data. The basic basin geographical information data includes basin boundary vector files and digital elevation data corresponding to the basin; the precipitation fusion product data is a high-precision and high-resolution precipitation information product generated by fusing historical precipitation data from multiple sources using algorithms and models.

[0016] The specific steps of S1-2 are as follows: The sample point grid is the basic sample points for feature aggregation analysis and is a grid of points evenly distributed at a certain distance for the optimized layout of rain gauge stations; according to the preset station density of the basin, a set of sample point grids A is generated; the sample point grid resolution (°) is:

[0017]

[0018] Where: F d is the preset station density of the basin (km 2 / station); L latlon is the approximate length corresponding to a unit of longitude and latitude (km); a is the sampling multiple (minimum number of aggregation classes) of sample points for feature aggregation analysis.

[0019] Further, the specific steps of S2 include:

[0020] S2-1. Calculation of geographical element feature distance matrix;

[0021] S2-2. Calculation of precipitation time series data feature distance matrix;

[0022] S2-3. Standardization processing of feature distance matrix;

[0023] S2-4. Calculate the weighted distance matrix.

[0024] Furthermore, the specific steps of S2-1 are as follows: Considering that precipitation is affected by terrain, topography, etc., and there are large regional differences in rainfall, longitude and latitude coordinates, altitude, and slope aspect are selected as characteristic indicators to quantify the spatial similarity degree of geographical elements; Calculate the horizontal distance matrix from the longitude and latitude coordinates:

[0025]

[0026] In the formula: S is the number of sample points in the sample point grid, dist i,j is the horizontal distance between sample point A i and sample point A j ;

[0027] dist i,j = R × arccos(sin(rlat i ) × sin(rlat j ) + cos(rlat i ) × cos(rlat i ) × cos(rlon j - rlon i ));

[0028] In the formula: rlon i , rlat i , rlon j , rlat j are the radian values corresponding to the longitude and latitude of sample point A i and sample point A j respectively; R is the average radius of the earth;

[0029] Calculate the altitude distance matrix from the grid sample point coordinates and digital elevation data:

[0030]

[0031] In the formula: S is the number of sample points in the sample point grid, elev i,j is the altitude distance between sample point A i and sample point A j :

[0032] elev i,j = |DEM(lon i , lat i ) - DEM(lon j , lat j )|;

[0033] In the formula: lon i , lat i, lon j , lat j are the longitude and latitude of sampling point A i and sampling point A j respectively, and DEM(lon i , lat i ) is the altitude corresponding to sampling point A i ;

[0034] Calculate the aspect value corresponding to each grid sampling point from the grid sampling point coordinates and digital elevation data:

[0035]

[0036] In the formula: are the gradients of sampling point A i in the longitude and latitude directions respectively;

[0037] Then calculate the aspect distance matrix from the aspect corresponding to the sampling point:

[0038]

[0039] In the formula: S is the number of sampling points in the sampling point grid, aspt i,j is the aspect difference between sampling point A i and sampling point A j :

[0040] aspt i,j = |Aspect i - Aspect j |.

[0041] Furthermore, the specific steps of S2-2 are as follows: In order to fully measure the similarity of precipitation time series data in terms of rainfall amount and process, the dynamic time distance is used to calculate the similarity of precipitation between two sampling points in terms of amount and process; the dynamic time distance minimizes the distance between two precipitation time series data by finding the best non-linear correspondence between them; this correspondence allows some parts of the precipitation time series to be stretched or compressed to make the two series match as much as possible;

[0042] First, according to the precipitation time series data P i , A j = p i , p i,1 , p i,2 ... p i,T , P j = p j,1 , p j,2 ... p j,T , where T is the length of the precipitation time series data, calculate the distance matrix between all possible time point pairs in the precipitation time series data of the two sampling points:

[0043]

[0044] Where: dp m,n is the sample point A i at time m and sample point A j the rainfall difference between time n:

[0045] dp m,n = p i,m - p j,n ;

[0046] Then, dynamic programming is used to calculate the dynamic time distance dtw i between the precipitation time series data of sample points A j and A i,j , which represents the minimum cumulative distance on the path path from the starting point (0, 0) to the ending point (T, T) of the distance matrix DP i,j and can be used to measure the similarity of the precipitation time series data of sample points A i and A j in terms of rainfall and process:

[0047]

[0048] From this, the dynamic time distance between the precipitation time series data of all grid sample points is calculated to construct a dynamic time distance matrix:

[0049]

[0050] Where: S is the number of sample points in the sample point grid.

[0051] Furthermore, the specific steps of the S2-3 are as follows: To avoid the results being dominated by features with larger scales during the feature clustering analysis based on the feature distance matrix, the feature distance matrix is subjected to min-max normalization:

[0052]

[0053] Where: D i ' ,j is the element in the standardized distance matrix D'; D i,j is the element in the feature distance matrix, and D is the feature distance matrix (D dist / D elev / D aspt / D dtw );

[0054] The specific steps of the S2-4 are as follows: To comprehensively consider different feature distance information, the calculated distance matrix is weighted, and information from different sources is comprehensively considered to improve the accuracy and reliability of the feature clustering results; the weighted mixed matrix is:

[0055] D ω = ω dist × D dist + ω elev × D elev + ω aspt × D aspt + ω dtw × D dtw ;

[0056] Where: ω dist , ω elev , ω aspt , ω dtw are the weights of the horizontal distance matrix, altitude distance matrix, slope aspect distance matrix, and dynamic time distance matrix, respectively.

[0057] Furthermore, the specific steps of S3 are as follows:

[0058] S3-1. Initialize the feature cluster center;

[0059] S3-2. Assign grid sample points;

[0060] S3-3. Update the cluster center;

[0061] S3-4. Iteratively output.

[0062] Even further, the specific steps of S3-1 are as follows: Determine the number of feature cluster classes K, that is, the preset number of stations, according to the preset station density and the basin area in the basin, and randomly select K points from the sample point grid set A as the initial cluster center point set E = e 1 , e 2 ... e K ;

[0063] The specific steps of S3-2 are as follows: For each grid sample point A in the non-cluster center point set E i , calculate its distance to all cluster center points, and select the cluster center point with the smallest distance as the cluster center of this point; The assignment process is shown as follows:

[0064]

[0065] Where: C k is the set of points assigned to the cluster center point e k , D ω (A i , e k ) and D ω (A i , e) respectively represent the similarity weighted distances between point A i and point e k , point e;

[0066] The specific steps of S3-3 are as follows: For each clustering set C k , try to use each non-clustering center point in the set as a potential clustering center e' k , and calculate the total distance within the set after replacement. The total distance within the set is calculated by the following formula:

[0067]

[0068] where: D ω (a, e' k ) is the similarity weighted distance between point a and the potential clustering center e' within the set C k ; k For each clustering center, select the non-clustering center point that minimizes the total distance of the set as the new clustering center point. If the current clustering center point is already optimal, it remains unchanged;

[0069] The specific steps of S3-4 are as follows: Repeat S3-2 to assign grid sample points and S3-3 to update the clustering center until the clustering center point no longer changes; finally, obtain K clustering sets, and each set is represented by a clustering center point e

[0070] k k .

[0071] Furthermore, S4 specifically includes: Optimize the layout of rain gauges in the basin according to the results of feature clustering analysis and clustering centers. When there is no rain gauge station in the set area, it is recommended to add a new one, and the addition location is near the clustering center point; when there are multiple rain gauge stations in the set area, it is recommended to optimize the existing stations and retain the existing stations that are closer to the clustering center.

[0072] Advantages of the present invention:

[0073] 1. In response to the need for optimizing the layout of precipitation stations, the present invention proposes an innovative method for optimizing the layout of rain gauge stations. This method deeply integrates multiple feature dimensions such as precipitation cause analysis, spatial similarity assessment of geographical elements, and comparison of precipitation time series data processes, aiming to scientifically optimize the layout of rain gauge stations in the basin; By using the method proposed by the present invention to optimize the layout of rain gauge stations in the basin, not only can the resource allocation be optimized, the costs of station construction and maintenance be reduced, but also the ability of rain gauge stations to obtain precipitation information in different regions and different time periods within the basin can be effectively enhanced, significantly improving the collection efficiency and accuracy of rain gauge data, thereby enhancing the basin's early warning and response capabilities to various precipitation events.

[0074] 2. In view of the requirements for optimizing the layout of precipitation stations, the present invention fully considers the similarity of precipitation time-series data in terms of rainfall amount and process on the basis of taking into account the causes of precipitation and taking into account the spatial similarity relationship of geographical elements, so as to realize the optimized layout of rainfall stations in the basin and maximize the efficiency and quality of obtaining rainfall information in the basin. Description of the Drawings

[0075] Figure 1 It is a schematic flow chart of a method for optimizing the layout of rainfall stations based on the aggregation of precipitation time series and geographical spatial characteristics;

[0076] Figure 2 It is a schematic diagram of the basin boundary and DEM;

[0077] Figure 3 It is a schematic diagram of the basin sample point grid;

[0078] Figure 4 It is a schematic diagram of the feature aggregation set and the aggregation center. Detailed Embodiment

[0079] The present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0080] Embodiment 1: As Figure 1 shown, a method for optimizing the layout of rainfall stations based on the aggregation of precipitation time series and geographical spatial characteristics includes the following steps:

[0081] S1. Data preparation and processing;

[0082] S1-1. Data preparation;

[0083] The data required for optimizing the layout of precipitation stations includes the basic data of basin geographical information and precipitation fusion product data. The basic data of geographical information includes the basin boundary vector file and the corresponding digital elevation data (DEM) of the basin. In this embodiment, the spatial resolution of the DEM is 0.00028°×0.00028°, and the basin boundary and the corresponding DEM are shown in Figure 2 . The precipitation fusion product data is a high-precision and high-resolution precipitation information product generated by fusing historical precipitation data from multiple sources using algorithms and models. In this embodiment, the time resolution of the precipitation fusion product data is hourly, and the spatial resolution is 0.01°×0.01°.

[0084] S1-2. Generation of sample point grid;

[0085] The sample point grid is the basic sample point for feature aggregation analysis and is a grid point evenly distributed at a certain distance for optimizing the layout of rainfall stations. According to the preset station density in the basin, a set A of sample point grids is generated. The spatial resolution (°) of the sample point grid is:

[0086]

[0087] Where: F d is the preset station density of the basin (km 2 / station); L latlon is the approximate length corresponding to the unit longitude and latitude (km); a is the sampling multiple (minimum number of clustering classes) of the sample points for feature clustering analysis.

[0088] In this embodiment, the preset station density of the basin is 400 km 2 / station, L latlon is taken as 100 km, a is taken as 4. Therefore, the spatial resolution of the sample point grid is 0.05°×0.05°, and the generated grid sample points are as Figure 3 shown.

[0089] S1-3. Smoothing processing of the digital elevation data and precipitation fusion product data corresponding to the basin;

[0090] Resample and smooth the digital elevation data and precipitation fusion product data according to the determined sample point grid resolution to ensure the representativeness of the data taken by the sample point grid.

[0091] In this embodiment, to ensure the representativeness of calculations such as elevation and slope direction, the DEM is resampled, and the spatial resolution of the resampled DEM is 0.01°×0.01°. The precipitation fusion product data meets the requirements and does not need to be processed.

[0092] S2. Calculation of the feature distance matrix;

[0093] S2-1. Calculation of the geographical element feature distance matrix;

[0094] Considering that precipitation is affected by terrain, topography, etc., and there are large regional differences in rainfall, longitude and latitude coordinates, altitude, and slope direction are selected as the feature indicators for quantifying the spatial similarity degree of geographical elements. Calculate the horizontal distance matrix from the longitude and latitude coordinates:

[0095]

[0096] In the formula: S is the number of sample points in the sample point grid, dist i,j is the sample point A i and the sample point A j horizontal distance between:

[0097] dist i,j = R×arccos(sin(rlat i )×sin(rlat j ) + cos(rlat i )×cos(rlat i )×cos(rlon j - rloni ))

[0098] Where: rlon i , rlat i , rlon j , rlat j are the radian values corresponding to the longitude and latitude of sample point A i and sample point A j respectively; R is the average radius of the earth.

[0099] Calculate the elevation distance matrix from the grid sample point coordinates and digital elevation data:

[0100]

[0101] In the formula: S is the number of sample points in the sample point grid, elev i,j is the elevation distance between sample point A i and sample point A j :

[0102] elev i,j = |DEM(lon i , lat i ) - DEM(lon j , lat j )|

[0103] In the formula: lon i , lat i , lon j , lat j are the longitudes and latitudes of sample point A i and sample point A j respectively, and DEM(lon i , lat i ) is the elevation corresponding to sample point A i .

[0104] First calculate the aspect values corresponding to each grid sample point from the grid sample point coordinates and digital elevation data:

[0105]

[0106] In the formula: are the gradients of point A i in the longitude and latitude directions respectively.

[0107] Then calculate the aspect distance matrix from the aspect corresponding to the sample point:

[0108]

[0109] In the formula: S is the number of sample points in the sample point grid, aspt i,j is sample point Ai and sample point A j Aspect difference between slopes:

[0110] aspt i,j = |Aspect i - Aspect j |

[0111] S2-2. Calculation of precipitation time series data feature distance matrix;

[0112] To fully measure the similarity of precipitation time series data in terms of rainfall amount and process, the dynamic time distance is used to calculate the similarity of precipitation between two sample points in terms of amount and process. The dynamic time distance minimizes the distance between two precipitation time series data by finding the best non-linear correspondence between them. This correspondence allows some parts of the precipitation time series to be stretched or compressed to make the two series match as closely as possible.

[0113] First, based on the precipitation time series data P i 、A j = p i ,p i,1 ...p i,2 ,P i,T = p j ,p j,1 ...p j,2 ...p j,T (T is the length of the precipitation time series data), calculate the distance matrix between all possible time point pairs in the two series:

[0114]

[0115] In the formula: dp m,m is the rainfall difference between sample point A i at time m and sample point A j at time n:

[0116] dp m,n = p i,m - p j,n

[0117] Then, use dynamic programming to calculate the dynamic time distance dtw i 、A j between the precipitation time series data, which represents the minimum cumulative distance on the path path from the starting point (0, 0) to the ending point (T, T) of the distance matrix DP i,j , and can be used to measure the similarity of the precipitation time series data of sample point A i,j 、A i 、A j in terms of rainfall amount and process:

[0118]

[0119] Calculate the dynamic time distance of precipitation time series data between all grid sample points, and construct a dynamic time distance matrix:

[0120]

[0121] In the formula: S is the number of sample points in the sample point grid.

[0122] S2-3. Standardize the feature distance matrix;

[0123] To avoid the result being dominated by features with larger scales when performing feature clustering analysis based on the feature distance matrix, perform min-max standardization on the feature distance matrix:

[0124]

[0125] Where: D i ′ ,j is the element in the standardized distance matrix D′; D i,j is the element in the feature distance matrix, D is the feature distance matrix (D dist / D elev / D aspt / D dtw ).

[0126] S2-4. Calculate the weighted distance matrix;

[0127] To comprehensively consider different feature distance information, perform weighted calculation on the calculated distance matrix, comprehensively consider information from different sources, and improve the accuracy and reliability of the feature clustering analysis results. The weighted mixed matrix is:

[0128] D ω = ω dist × D dist + ω elev × D elev + ω aspt × D aspt + ω dtw × D dtw

[0129] Where: ω dist , ω elev , ω aspt , ω dtw are the weights of the horizontal distance matrix, altitude distance matrix, slope aspect distance matrix, and dynamic time distance matrix, respectively.

[0130] In this embodiment, ω dist , ω elev , ω aspt , ω dtwTake values 0.4, 0.2, 0.1, and 0.3 respectively.

[0131] S3. Feature clustering analysis;

[0132] S3-1. Initialize the feature clustering center;

[0133] Determine the number of clusters K (preset number of stations) according to the preset station density and catchment area of the catchment, and randomly select K points from the sample grid set A as the initial clustering center point set E = e 1 ,e 2 ...e K . In this embodiment, the number of clusters K is set to 20.

[0134] S3-2. Assign grid samples;

[0135] For each grid sample A in the non-clustering center point set E i , calculate its distances to all clustering center points, and select the clustering center point with the minimum distance as the clustering center of this point. The assignment process is shown as follows:

[0136]

[0137] Where: C k is the point set assigned to the clustering center point e k , D ω (A i ,e k ) and D ω (A i ,e) respectively represent the similarity weighted distances between point A i and point e k , point e;

[0138] S3-3. Update the clustering center;

[0139] For each clustering set C k , try to take each non-clustering center point in the set as the potential clustering center e′ k , and calculate the total distance within the set after replacement. The total distance within the set is calculated by the following formula:

[0140]

[0141] Where: D ω (a,e′ k ) is the similarity weighted distance between point a and the potential clustering center e′ k within the set C k .

[0142] For each clustering center, select the non-clustering center point that minimizes the total intra-cluster distance as the new clustering center point (if the current clustering center point is already optimal, it remains unchanged).

[0143] S3-4. Iterative output;

[0144] Repeat S3-2 to assign grid sample points and S3-3 to update the clustering center until the clustering center point no longer changes. Finally, obtain K clustering sets, each clustering set represented by a clustering center point e k denotes.

[0145] The feature clustering set and the clustering center are as Figure 4 shown.

[0146] S4. Optimized layout;

[0147] According to the feature clustering analysis results and the clustering center, optimize the layout of the rainfall stations in the basin. When there is no rainfall station in the set area, it is recommended to add a new one, and the added position is near the clustering center point; when there are multiple rainfall stations in the set area, it is recommended to optimize the existing stations and retain the existing stations that are closer to the clustering center.

[0148] The above embodiments are only the preferred technical solutions of the present invention and should not be regarded as limitations on the present invention. The protection scope of the present invention should be the technical solutions recorded in the claims, including the equivalent replacement solutions of the technical features in the technical solutions recorded in the claims. That is, the equivalent replacement improvements within this scope are also within the protection scope of the present invention.

Claims

1. A method for optimizing the layout of rain gauges based on precipitation time series and geographic spatial feature aggregation, characterized in that: It includes the following steps: S1. Data preparation and processing; S2, feature distance matrix calculation; S3, feature cluster analysis; S4. Optimize layout.

2. The method for optimizing the layout of rain gauges based on precipitation time series and geographic spatial feature aggregation according to claim 1 is characterized in that: The S1 specifically includes: S1-1. Data preparation; S1-2, sample point grid generation; S1-3. Smoothing of digital elevation data and precipitation fusion product data corresponding to the watershed.

3. The method for optimizing the layout of rain gauges based on precipitation time series and geographic spatial feature aggregation according to claim 2 is characterized in that: The specific steps of S1-1 are as follows: the data required for the optimized layout of precipitation stations include basin geographic information basic data and precipitation fusion product data, the geographic information basic data includes basin boundary vector files and basin corresponding digital elevation data; precipitation fusion product data is a high-precision, high-resolution precipitation information product generated by fusing historical precipitation data from multiple sources and using algorithms and models; The specific steps of S1-2 are as follows: the sample point grid is the basic sample point for feature cluster analysis, which is a grid point evenly distributed at a certain distance for optimizing the layout of rainfall stations; the sample point grid set A is generated according to the preset station density of the basin; the sample point grid resolution (°) is: Among them: F d Preset station density for the basin (km 2 / station); L latlon is the approximate length corresponding to unit longitude and latitude (km); a is the sampling multiple (minimum number of clusters) used for feature cluster analysis.

4. The method for optimizing the layout of rain gauges based on precipitation time series and geographic spatial feature aggregation according to claim 1 is characterized in that: The S2 specifically includes: S2-1, calculation of geographical element characteristic distance matrix; S2-2, calculation of characteristic distance matrix of precipitation time series data; S2-3, feature distance matrix standardization; S2-4. Calculate the weighted distance matrix.

5. The method for optimizing the layout of rain gauges based on precipitation time series and geographic spatial feature aggregation according to claim 4 is characterized in that: The specific steps of S2-1 are: considering that precipitation is affected by topography and terrain, and the regional differences in rainfall are large, longitude and latitude coordinates, altitude, and slope aspect are selected as characteristic indicators for quantifying the spatial similarity of geographical elements; the horizontal distance matrix is ​​calculated from the longitude and latitude coordinates: Where: S is the number of sample points in the sample grid, dist i,j Sample point A i Sample A j Horizontal distance between dist i,j =R×arccos(sin(rlat i )×sin(rlat j )+cos(rlat i )×cos(rlat i )×cos(rlon j -rlon i )); Where: rlon i , rlat i , rlon j , rlat j Sample point A i Sample A j The radian values ​​corresponding to longitude and latitude; R is the average radius of the earth; Calculate the elevation distance matrix from the grid sample point coordinates and digital elevation data: Where: S is the number of sample points in the sample grid, elev i,j Sample point A i Sample A j Altitude distance: elev i,j =|DEM(lon i ,lat i )-DEM(lon j ,lat j )|; Where: lon i ,lat i ,lon j ,lat j Sample point A i Sample A j Longitude and latitude, DEM (lon i ,lat i ) is sample point A i The corresponding altitude; First calculate the corresponding slope value of each grid sample point based on the grid sample point coordinates and digital elevation data: Where: Sample point A i gradients in longitude and latitude; Then calculate the aspect distance matrix from the corresponding aspects of the sample points: Where: S is the number of sample points in the sample grid, aspt i,j Sample point A i Sample A j Slope difference between aspt i,j =|Aspect i -Aspect j |。 6. The method for optimizing the layout of rain gauges based on precipitation time series and geographic spatial feature aggregation according to claim 4 is characterized in that: The specific steps of S2-2 are: in order to fully measure the similarity of precipitation time series data in terms of rainfall and process, the dynamic time distance is used to calculate the similarity of precipitation in terms of amount and process between two sample points; the dynamic time distance minimizes the distance between the two precipitation time series data by finding the best nonlinear correspondence between them; this correspondence allows some parts of the precipitation time series to be stretched or compressed so that the two series match as much as possible; First, according to the two sample points A i , A j Precipitation time series data i =p i,1 ,p i,2 ...p i,T , P j =p j,1 ,p j,2 ...p j,T , T is the length of the precipitation time series data, and the distance matrix between all possible time point pairs in the precipitation time series data of two sample points is calculated: Where: dp m,n Sample point A i m time and sample point A j Rainfall difference at time n: dp m,n =p i,m -p j,n ; Then, dynamic programming is used to calculate the sample point A i , A j Dynamic time distance dtw between precipitation time series data i,j , which represents the distance matrix DP i,j The minimum cumulative distance on the path from the starting point (0, 0) to the end point (T, T) can be used to measure the sample point A. i , A j Similarities of precipitation time series data in terms of rainfall and process: The dynamic time distance of precipitation time series data between all grid sample points is calculated and the dynamic time distance matrix is ​​constructed: Where: S is the number of sample points in the sample grid.

7. The method for optimizing the layout of rain gauges based on precipitation time series and geographic spatial feature aggregation according to claim 4 is characterized in that: The specific steps of S2-3 are: to avoid the results of feature clustering analysis based on the feature distance matrix being dominated by features with larger scales, the feature distance matrix is ​​subjected to minimum-maximum normalization processing: Where: D i ' ,j is the element in the standardized distance matrix D'; D i,j is an element in the characteristic distance matrix, D is the characteristic distance matrix (D dist / D elev / D aspt / D dtw ); The specific steps of S2-4 are: in order to comprehensively consider different feature distance information, weighted calculation is performed on the calculated distance matrix, and information from different sources is comprehensively considered to improve the accuracy and reliability of the feature clustering results; the weighted mixed matrix is: D ω =ω dist ×D dist +oh elev ×D elev +oh aspt ×D aspt +oh dtw ×D dtw ; Where: dist ,ω elev ,ω aspt ,ω dtw They are the weights of the horizontal distance matrix, altitude distance matrix, aspect distance matrix, and dynamic time distance matrix respectively.

8. The method for optimizing the layout of rain gauges based on precipitation time series and geographic spatial feature aggregation according to claim 1 is characterized in that: The S3 specifically includes: S3-1, initializing the feature clustering center; S3-2, assigning grid sample points; S3-3, update the agglomeration center; S3-4, iterative output.

9. The method for optimizing the layout of rain gauges based on precipitation time series and geographic spatial feature aggregation according to claim 8, characterized in that: The specific steps of S3-1 are: determine the number of feature set clusters K according to the preset site density and the basin area, that is, the preset number of sites, and randomly select K points from the sample point grid set A as the initial cluster center point set E = e1, e2...e K ; The specific steps of S3-2 are: for each grid sample point A in the non-clustered center point set E i , calculate its distance to all cluster center points, and select the cluster center point with the smallest distance as the cluster center of the point; the allocation process is expressed as follows: Where: C k is assigned to the cluster center e k The point set, D ω (A i ,e k ) and D ω (A i ,e) respectively represent point A i and point e k , similarity weighted distance between point e; The specific steps of S3-3 are: for each cluster C k , try to use each non-cluster center point in the set as a potential cluster center e′ k , and calculate the total distance within the set after replacement. The total distance within the set is calculated by the following formula: Where: D ω (a,e′ k ) is the set C k Interior point a and potential agglomeration center e′ k Similarity weighted distance between them; For each cluster center, select the non-cluster center point that minimizes the total distance of the set as the new cluster center point. If the current cluster center point is already the best, keep it unchanged. The specific steps of S3-4 are: repeat S3-2 to allocate grid samples and S3-3 to update the cluster center until the cluster center point no longer changes; finally, K clusters are obtained, each of which consists of a cluster center point e k express.

10. The method for optimizing the layout of rain gauges based on precipitation time series and geographic spatial feature aggregation according to claim 1, characterized in that: The S4 specifically includes: optimizing the layout of rain gauges in the basin according to the characteristic clustering analysis results and the clustering center; when there is no rain gauge station in the clustering area, it is recommended to add one near the clustering center; when there are multiple rain gauge stations in the clustering area, it is recommended to optimize the existing stations and retain the existing stations that are closer to the clustering center.