Basin similarity analysis method and system based on cumulative probability distribution
Patent Information
- Application Number
- CN202311208001.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-19
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2043-09-19
AI Technical Summary
然而空间邻近并不一定代表物理特征会相似,聚类和主成分分析也存在一定的随意性,此外,目前流域相似性分析多采用流域面平均特征进行相似性分析,缺少流域特征局部分布的考量,且不同流域之间的面积大小不一致,导致流域特征值长度无法保证一致,难以直接进行聚类和相似性比较,迫切需要提出科学合理的流域相似性分析方法
[0025]综合考虑有资料流域和无资料流域特征值长度的一致性问题,基于累积概率分布,进行流域同频率采样得到相同长度的流域特征值,并进行流域逐网格特征值提取,充分考虑流域全局特征和局部细节特征,针对流域大小不一致条件下也能适用,提高了流域相似性分析的准确性,为无水文观测资料流域的相似性分析提供技术指导,进而对于无水文观测资料流域水文预报具有重要作用。
Smart Images

Figure CN117273215B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of hydrological resource data processing technology applicable to management or forecasting purposes, and in particular to a watershed similarity analysis method and system based on cumulative probability distribution. Background Technology
[0002] Watershed similarity analysis is a commonly used method for hydrological forecasting in data-scarce areas. Due to the lack of monitoring stations in many small watersheds, their hydrological characteristics cannot be accurately understood, limiting the rational development and utilization of water resources in these areas. Watershed similarity analysis allows the application of model parameters and hydrological information from watersheds with existing observational data to watersheds without such data, enabling runoff forecasting in data-scarce areas, which is of great significance for water resource development. Current research on watershed similarity analysis in data-scarce areas mainly considers spatial proximity, physical attributes, and hydrological sequence characteristics. For example, the geographically closest watershed to the data-scarce area is selected as a substitute watershed; or watershed physical characteristics such as topography, soil texture, and land use are used as watershed similarity indicators; or principal component analysis and clustering are performed based on watershed hydrological characteristics to group watersheds with similar attributes. These methods can all help researchers conduct watershed similarity analysis in data-scarce areas to a certain extent for hydrological forecasting. However, spatial proximity does not necessarily imply similar physical characteristics. Clustering and principal component analysis also have a certain degree of arbitrariness. Furthermore, current watershed similarity analysis often uses the average feature of the watershed surface for similarity analysis, lacking consideration of the local distribution of watershed features. Moreover, the inconsistent area sizes between different watersheds make it difficult to guarantee the consistency of watershed feature value lengths, making it difficult to directly perform clustering and similarity comparisons. Therefore, there is an urgent need to propose a scientific and reasonable watershed similarity analysis method. Summary of the Invention
[0003] The purpose of this invention is to disclose a watershed similarity analysis method and system based on cumulative probability distribution, which can compare watershed similarity under different watershed feature value lengths, and fully consider both global and local detailed features of the watershed to improve the accuracy of watershed similarity analysis.
[0004] To achieve the above objectives, the watershed similarity analysis method based on cumulative probability distribution disclosed in this invention includes:
[0005] Step S0: Obtain the distribution information of the first catchment area corresponding to the outlet point of the first watershed without hydrological observation data, and obtain the distribution information of the second catchment area corresponding to the outlet point of the second watershed with hydrological observation data.
[0006] Step S1: Determine a common hydrological feature shared by the first and second catchment areas. Extract the feature value of the first catchment area grid by grid to form a column vector V1; and extract the feature value of the first catchment area grid by grid to form a column vector V2; sort V1 and V2 in ascending order to obtain VS1 and VS2:
[0007]
[0008]
[0009] In the formula, v ij For the j-th feature value of the i-th watershed, vs ij Let be the j-th feature value after ascending sorting of the i-th watershed, m be the number of feature values in the first catchment area, and n be the number of feature values in the second catchment area;
[0010] Step S2: Calculate the cumulative probability distribution functions F1(x) and F2(x) for VS1 and VS2 respectively; where F1(x) = P(VS1≤x) and F2(x) = P(VS2≤x);
[0011] Step S3: Sample F1(x) and F2(x) at the same frequency to obtain two column vector data VS1' and VS'2 of the same length;
[0012]
[0013]
[0014] Where T is the length of the two column vectors VS1' and VS'2; vs1' t Let t be the t-th data point in the column vector data VS1'; vs' 2t Let t be the t-th data in the column vector data VS'2;
[0015] Step S4: Calculate the KGE index to obtain the similarity between the first watershed and the second watershed. The calculation formula is as follows:
[0016]
[0017]
[0018]
[0019]
[0020] Where r is the correlation coefficient between the characteristic values of the first watershed and the second watershed, α is the ratio of the standard deviations of the characteristic values of the first watershed and the second watershed, and β is the ratio of the means of the characteristic values of the first watershed and the second watershed. The mean of VS1'; This is the mean of VS'2.
[0021] Preferably, during the sampling of F1(x) and F2(x) at the same frequency, two column vector data VS1' and VS'2 of the same length are obtained by linear interpolation or spline interpolation.
[0022] Preferably, the method of the present invention further includes: obtaining the KGE index between the first catchment area and the catchment areas corresponding to the outlet points of the second watershed and other watersheds outside the second watershed with hydrological observation data, based on the same feature, to obtain the target watershed most similar to the first watershed based on the feature.
[0023] To achieve the above objectives, the present invention also discloses a watershed similarity analysis system based on cumulative probability distribution, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method.
[0024] Therefore, this invention has a clear concept, is easy to operate, and is highly practical. It also has at least the following beneficial effects:
[0025] Taking into account the consistency of feature value lengths between watersheds with and without data, this study uses cumulative probability distribution to sample watersheds at the same frequency to obtain watershed feature values of the same length. It then performs grid-by-grid feature value extraction, fully considering both global and local detailed features of the watershed. This approach is applicable even under conditions of inconsistent watershed sizes, improving the accuracy of watershed similarity analysis. It provides technical guidance for similarity analysis of watersheds without hydrological data, and thus plays a crucial role in hydrological forecasting for watersheds without hydrological data.
[0026] The present invention will now be described in further detail with reference to the accompanying drawings. Attached Figure Description
[0027] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:
[0028] Figure 1 This is the NDVI mean distribution map of watershed 1 (ID: 01667500) disclosed in the embodiments of the present invention.
[0029] Figure 2This is the NDVI mean distribution map of watershed 2 (ID: 03164000) disclosed in the embodiments of the present invention.
[0030] Figure 3 This is a distribution map of NDVI values after sampling at the same frequency in watershed 1 and watershed 2, as disclosed in an embodiment of the present invention. Detailed Implementation
[0031] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings, but the present invention can be implemented in many different ways as defined and covered by the claims.
[0032] Example 1
[0033] This embodiment discloses a watershed similarity analysis method based on cumulative probability distribution, including:
[0034] Step S0: Obtain the distribution information of the first catchment area corresponding to the outlet point of the first watershed without hydrological observation data, and obtain the distribution information of the second catchment area corresponding to the outlet point of the second watershed with hydrological observation data.
[0035] In this step, the so-called "catchment area" refers to a region within a certain range where rainwater and groundwater converge. Its distribution information mainly includes the area used for grid division and the feature values of each grid used for subsequent feature calculation.
[0036] Step S1: Determine a hydrological-related feature shared by the first and second catchment areas, extract the feature value of the first catchment area grid by grid to form a column vector V1; and extract the feature value of the first catchment area grid by grid to form a column vector V2; sort V1 and V2 in ascending order to obtain VS1 and VS2.
[0037]
[0038]
[0039] In the formula, v ij For the j-th feature value of the i-th watershed, vs ij Let be the j-th feature value after ascending sorting of the i-th watershed, m be the number of feature values in the first catchment area, and n be the number of feature values in the second catchment area.
[0040] In this step, the selected features include, but are not limited to, Normalized Difference Vegetation Index (NDVI), altitude or topography, soil texture, land use, and other watershed physical characteristics. Preferably, the grid size used to extract feature values from the two regions in this step is the same. As a degraded implementation, when the actual geographical area covered by the grids of the two regions differs within a certain threshold range, certain technical effects can still be achieved. Such variations are all within the protection scope of this invention.
[0041] Step S2: Calculate the cumulative probability distribution functions F1(x) and F2(x) for VS1 and VS2 respectively; where F1(x) = P(VS1≤x) and F2(x) = P(VS2≤x).
[0042] In this step, the method for finding the cumulative probability distribution function is as follows: First, sort all M data points from smallest to largest. The probability value of the first value is 1 / M, the probability value of the second value is 2 / M, and so on. Then, with the sorted data as the x-axis and the probability value as the y-axis, the cumulative probability distribution can be obtained.
[0043] Step S3: Sample F1(x) and F2(x) at the same frequency to obtain two column vector data VS1' and VS'2 of the same length.
[0044]
[0045]
[0046] Where T is the length of the two column vectors VS1' and VS'2; vs1' t Let t be the t-th data point in the column vector data VS1'; vs' 2t Let t be the t-th data point in the column vector data VS'2.
[0047] In this step, preferably, linear interpolation or spline interpolation can be used to obtain two column vector data VS1' and VS'2 of the same length. For example, if F1(x) has 100 numbers and F2(x) has 50 numbers, to obtain two data sequences of the same length, F1(x) can be divided into groups of two numbers and averaged, resulting in a new sequence of 50 numbers, which has the same length as F2(x). Alternatively, one number can be interpolated between every two numbers in F2(x), resulting in another new sequence of 100 numbers, which also has the same length as F1(x).
[0048] Therefore, through the combined effect of steps S1 to S3, the following unexpected technical effects can also be achieved:
[0049] The watershed similarity can also be compared between grids of different shapes (such as triangular grids and rectangular grids); moreover, the angular relationship between the grid and the catchment area has little impact on the overall accuracy, thereby greatly improving the universality and reliability of this embodiment.
[0050] like Figure 1 and Figure 2 The image shows the grid-by-grid NDVI distribution map for watershed 1 (ID: 01667500) and watershed 2 (ID: 03164000); the sampled data of consistent length obtained after linear interpolation sampling at the same frequency based on the cumulative probability distribution function are shown below. Figure 3 As shown.
[0051] Step S4: Calculate the KGE index to obtain the similarity between the first watershed and the second watershed. The calculation formula is as follows:
[0052]
[0053]
[0054]
[0055]
[0056] Where r is the correlation coefficient between the characteristic values of the first watershed and the second watershed, α is the ratio of the standard deviations of the characteristic values of the first watershed and the second watershed, and β is the ratio of the means of the characteristic values of the first watershed and the second watershed. The mean of VS1'; This is the mean of VS'2.
[0057] In a specific application, based on Figure 3 According to the KGE index, the similarity between the two watersheds is 0.92, indicating that the two watersheds are very similar and the parameters related to NDVI in watershed 1 can be transferred to watershed 2.
[0058] Preferably, the method of this embodiment further includes: obtaining the KGE index between the first catchment area and the catchment areas corresponding to the outlet points of the second watershed and other watersheds outside the second watershed with hydrological observation data, based on the same feature, to obtain the target watershed most similar to the first watershed based on the feature.
[0059] Example 2
[0060] The present invention also discloses a watershed similarity analysis system based on cumulative probability distribution, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method disclosed in Embodiment 1 above.
[0061] In summary, the methods and systems disclosed in the embodiments of this invention are clear in concept, convenient to operate, and highly practical. Furthermore, they possess at least the following beneficial effects:
[0062] Taking into account the consistency of feature value lengths between watersheds with and without data, this study uses cumulative probability distribution to sample watersheds at the same frequency to obtain watershed feature values of the same length. It then performs grid-by-grid feature value extraction, fully considering both global and local detailed features of the watershed. This approach is applicable even under conditions of inconsistent watershed sizes, improving the accuracy of watershed similarity analysis. It provides technical guidance for similarity analysis of watersheds without hydrological data, and thus plays a crucial role in hydrological forecasting for watersheds without hydrological data.
[0063] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A watershed similarity analysis method based on cumulative probability distribution, characterized in that, include: Step S0: Obtain the distribution information of the first catchment area corresponding to the outlet point of the first watershed without hydrological observation data, and obtain the distribution information of the second catchment area corresponding to the outlet point of the second watershed with hydrological observation data. Step S1: Determine a common hydrological feature shared by the first and second catchment areas. Extract the feature value of the first catchment area grid by grid to form a column vector V1; and extract the feature value of the second catchment area grid by grid to form a column vector V2; sort V1 and V2 in ascending order to obtain VS1 and VS2: In the formula, v ij For the j-th feature value of the i-th watershed, vs ij Let be the j-th feature value after ascending sorting of the i-th watershed, m be the number of feature values in the first catchment area, and n be the number of feature values in the second catchment area; Step S2: Calculate the cumulative probability distribution functions F1(x) and F2(x) for VS1 and VS2 respectively; where F1(x) = P(VS1≤x) and F2(x) = P(VS2≤x); Step S3: Sample F1(x) and F2(x) at the same frequency to obtain two column vector data VS1' and VS'2 of the same length; Where T is the length of the two column vectors VS1' and VS'2; vs1' t Let t be the t-th data point in the column vector data VS1'; vs' 2t Let t be the t-th data in the column vector data VS'2; Step S4: Calculate the KGE index to obtain the similarity between the first watershed and the second watershed. The calculation formula is as follows: Where r is the correlation coefficient between the characteristic values of the first watershed and the second watershed, α is the ratio of the standard deviations of the characteristic values of the first watershed and the second watershed, and β is the ratio of the means of the characteristic values of the first watershed and the second watershed. The mean of VS1'; This is the mean of VS'2.
2. The method according to claim 1, characterized in that, During the sampling of F1(x) and F2(x) at the same frequency, two column vector data VS1' and VS'2 of the same length are obtained by linear interpolation or spline interpolation.
3. The method according to claim 1 or 2, characterized in that, Also includes: Based on the same feature, the KGE index between the outflow points of the first catchment area and the second watershed, as well as the catchment areas of other watersheds outside the second watershed with hydrological observation data, is used to obtain the target watershed that is most similar to the first watershed based on this feature.
4. A watershed similarity analysis system based on cumulative probability distribution, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method described in any one of claims 1 to 3.
Citation Information
Patent Citations
Quantitative basin similarity comprehensive evaluation index computing method
CN107391939A
Basin similarity classification method and device
CN113887635A