Similarity discrimination and clustering method, system and device for middle-swimming small watershed of Yellow River watershed and storage medium

By using weighted Euclidean distance and K-means clustering algorithms, the problem of quantifying the similarity between watersheds in the high sediment content area of ​​the middle reaches of the Yellow River was solved, enabling accurate hydrological model transplantation and classification management of watersheds without data, and providing scientific data support.

CN122020196APending Publication Date: 2026-05-12INST OF EARTH ENVIRONMENT CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF EARTH ENVIRONMENT CHINESE ACAD OF SCI
Filing Date
2026-01-23
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies are insufficient to comprehensively and accurately quantify the similarity between river basins in the high sediment content areas of the middle reaches of the Yellow River, and cannot effectively solve the problems of hydrological model transplantation and accurate classification and management of river basins without data.

Method used

We adopted a weighted Euclidean distance algorithm combined with the K-means clustering algorithm to construct a multi-factor fusion feature discrimination system by acquiring static topographic and dynamic hydrological indicators. We used the coefficient of variation method to objectively determine the indicator weights, quantify the similarity between watersheds, and perform cluster analysis.

Benefits of technology

It has achieved accurate similarity identification and clustering of small watersheds in the middle reaches of the Yellow River, providing scientific data support for hydrological model transplantation and classification management, which conforms to the natural geographical differentiation law.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122020196A_ABST
    Figure CN122020196A_ABST
Patent Text Reader

Abstract

The invention discloses a Yellow River basin midstream small watershed similarity discrimination and clustering method, system and device, and a storage medium, and the method comprises the steps: obtaining the static topographic index data and dynamic hydrological index data of a target small watershed, and carrying out the standardization processing of the static topographic index data and dynamic hydrological index data, and obtaining the standardized data; calculating variable coefficients of the static terrain indexes and the dynamic hydrological indexes, and determining weights of the static terrain indexes and the dynamic hydrological indexes according to the variable coefficients; on the basis of the weight and the standardized data, a weighted Euclidean distance algorithm is adopted to calculate the feature space distance between the small watersheds, and the similarity coefficient of the small watersheds is determined according to the feature space distance; and grouping the small watershed by using a K-means clustering algorithm, calculating a contour coefficient, determining an optimal clustering number, and obtaining a small watershed similarity clustering result. And the similarity degree between the drainage basins can be quantified comprehensively and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of hydrological prediction and relates to a method, system, device and storage medium for similarity discrimination and clustering of small watersheds in the middle reaches of the Yellow River. Background Technology

[0002] The Loess Plateau region in the middle reaches of the Yellow River is the world's largest continuous loess deposition area. Soil erosion is a prominent problem in this region, making it one of the areas with the most severe soil erosion globally. The Yellow River produces a massive amount of sediment in this section, with complex runoff mechanisms. In recent years, vegetation cover in most parts of the Loess Plateau has improved significantly (Zhang et al., 2000). However, this vegetation restoration has also triggered a negative soil moisture balance, with soil drying and the thickening of the dry soil layer becoming increasingly prominent, posing challenges to sustainable vegetation restoration and ecosystem stability (Shao et al., 2016). A complex coupling relationship exists between water and sediment processes and vegetation dynamics in the middle reaches of the Yellow River (Tang et al., 2023). The complexity of its runoff mechanisms directly affects the sediment transport patterns of the Yellow River and has significant impacts on basin hydrological forecasting, flood control and drought relief, and water resource regulation. Therefore, conducting research on water and sediment prediction in the Loess Plateau region is of great significance for optimizing water resource allocation in the Yellow River Basin, implementing water and sediment regulation, improving the flood control and disaster reduction system, and promoting comprehensive river management (Li Linqi et al., 2025). Reliable data support is the foundation of hydrological forecasting (Vasiletal. 2021). However, the region suffers from strong spatial heterogeneity of underlying surface conditions, significant impacts from human activities (such as urbanization, ecological restoration, and silt-retention dam construction), and sparse and unevenly distributed hydrological stations, leading to a lack of data in many small watersheds, which severely restricts the in-depth development of water and sediment prediction work.

[0003] The watershed characteristic similarity method based on hydrological similarity theory (Liu Jintao et al., 2014) has become an effective approach to solving hydrological prediction problems in areas with no or insufficient data. This method quantifies the degree of similarity in attribute characteristics between watersheds and transplants existing mature hydrological models to watersheds with similar water and sediment characteristics that lack data, thereby achieving scientific prediction of hydrological processes.

[0004] Although the watershed feature similarity method based on hydrological similarity theory is an effective way to solve the prediction problem in areas without data, the existing technology is difficult to fully and accurately quantify the degree of similarity between watersheds when dealing with high sediment content areas such as the middle reaches of the Yellow River. It cannot effectively solve the problems of hydrological model transplantation and accurate classification and management of watersheds without data. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method, system, device and storage medium for similarity discrimination and clustering of small watersheds in the middle reaches of the Yellow River Basin. This method can comprehensively and accurately quantify the degree of similarity between watersheds and effectively solve the problems of hydrological model transplantation and accurate classification and management of watersheds without data.

[0006] To achieve the above objectives, the present invention employs the following technical solution: A method for similarity discrimination and clustering of small watersheds in the middle reaches of the Yellow River Basin, comprising: Obtain static topographic data and dynamic hydrological data of the target small watershed, and standardize the static topographic data and dynamic hydrological data to obtain standardized data; Calculate the coefficient of variation for each static topographic index and dynamic hydrological index, and determine the weight of each static topographic index and dynamic hydrological index based on the coefficient of variation. Based on weighted and standardized data, a weighted Euclidean distance algorithm is used to calculate the feature space distance between small watersheds, and the similarity coefficient of small watersheds is determined based on the feature space distance. The K-means clustering algorithm was used to group the small watersheds, and the silhouette coefficient was calculated to determine the optimal number of clusters, thus obtaining the small watershed similarity clustering results.

[0007] Optional static topographic indicators include watershed area, watershed shape coefficient, average slope, topographic humidity index, river length, river gradient, cultivated land ratio, forest ratio, grassland ratio, terrace ratio, and number of silt-retaining dams; dynamic hydrological indicators include rainfall, river runoff, and river sediment content.

[0008] Optionally, the standardization process employs the Z-Score standardization method to eliminate the dimensional differences between static topographic indicators and dynamic hydrological indicators.

[0009] Optionally, the process of determining the weights of each static topographic index and dynamic hydrological index includes:

[0010]

[0011] In the formula, E i As an indicator i coefficient of variation, σ i and μ i These are the standard deviation and mean of the indicator, respectively. ω i As an indicator i The weight, n This represents the total number of indicators.

[0012] Optionally, the process of calculating the characteristic spatial distance between small watersheds includes:

[0013] In the formula, D ij For the basin i With the basin j The weighted Euclidean distance between them; ω k For the first k The weight of each indicator; x ik and x jk respectively watershed i and the basin j In the k Standardized values ​​for each indicator.

[0014] Optionally, the process of determining the small watershed similarity coefficient includes:

[0015] in: The similarity coefficient, For Euclidean distance, This represents the maximum distance.

[0016] Optionally, the process of calculating the silhouette coefficient to determine the optimal number of clusters includes:

[0017] In the formula, a i For the sample i The average distance to all other samples within the same cluster reflects the density of the cluster. b i For the sample i The average distance to all samples in its nearest neighbor cluster reflects the degree of separation between clusters; S i The value of is in the range of [-1, 1], and the larger the value, the more reasonable the clustering of the sample. For a complete clustering result, its overall quality is determined by the mean of the silhouette coefficients of all samples, i.e., the average silhouette coefficient. S k To measure:

[0018] In the formula, n This represents the total number of samples. By calculating different presets... k value corresponding S k And select to make S k Maximize kThe value is used as the final cluster number.

[0019] A similarity discrimination and clustering system for small watersheds in the middle reaches of the Yellow River Basin includes: The indicator data acquisition module is used to acquire static topographic indicator data and dynamic hydrological indicator data of the target small watershed, and to standardize the static topographic indicator data and dynamic hydrological indicator data to obtain standardized data. The indicator weight determination module is used to calculate the coefficient of variation of each static topographic indicator and dynamic hydrological indicator, and determine the weight of each static topographic indicator and dynamic hydrological indicator based on the coefficient of variation. The similarity calculation module is used to calculate the feature space distance between small watersheds based on weighted and standardized data using a weighted Euclidean distance algorithm, and to determine the similarity coefficient of the small watersheds based on the feature space distance. The clustering module is used to group small watersheds using the K-means clustering algorithm, calculate the silhouette coefficient to determine the optimal number of clusters, and obtain the small watershed similarity clustering results.

[0020] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the similarity discrimination and clustering method for small watersheds in the middle reaches of the Yellow River Basin.

[0021] A computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the similarity discrimination and clustering method for small watersheds in the middle reaches of the Yellow River Basin.

[0022] Compared with the prior art, the present invention has the following beneficial effects: This invention constructs a multi-factor fusion feature discrimination system by acquiring static topographic and dynamic hydrological indicators of target small watersheds. It systematically characterizes the interaction between meteorological and hydrological elements and underlying surface conditions from three dimensions: climate characteristics, runoff characteristics, and underlying surface characteristics. The coefficient of variation method is used to objectively determine indicator weights based on the degree of data dispersion, avoiding the bias of subjective weighting. Furthermore, a weighted Euclidean distance algorithm is employed to accurately calculate the differences between watersheds in the feature space, achieving a quantitative assessment of similarity. Further combining K-means clustering and silhouette coefficient optimization reveals the natural grouping patterns of small watersheds (such as large-basin type, medium-basin type, and small-basin type), ensuring that the classification results conform to natural geographical differentiation patterns. This similarity-based clustering method can find the most similar reference watersheds for watersheds without data, thereby enabling the scientific transfer of mature hydrological models and providing scientific data support for regional soil and water conservation planning and classified management. Attached Figure Description

[0023] Figure 1This is a map showing the location distribution of the study area in this invention; Figure 2 This is a schematic diagram of the watershed hydrological index data of the present invention; Figure 3 This is a pie chart showing the weights of the watershed similarity discrimination index of the present invention. Figure 4 This is a heatmap showing the similarity of small watersheds in this invention; Figure 5 is a schematic diagram of the main similarity indicators for the most similar watershed pair according to the present invention, wherein... Figure 5a For the Anse-Wuqi section, Figure 5b For the Wuqi-Zhidan section; Figure 6 is a schematic diagram of the watershed clustering situation according to the present invention, wherein... Figure 6a The cluster distribution after dimensionality reduction of the feature data is shown in the figure. Figure 6b Create radar charts for key indicators; Figure 7 This is a schematic diagram illustrating the clustering effect evaluation of the present invention. Detailed Implementation

[0024] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0025] The following disclosure provides many different embodiments or examples for implementing various structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the invention. Furthermore, reference numerals and / or letters may be repeated in different examples; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed. In addition, examples of various specific processes and materials are provided in this invention, but those skilled in the art will recognize the application of other processes and / or the use of other materials.

[0026] This embodiment uses hydrological stations as control sections for watershed division. Nine representative and widely covered sub-watersheds were selected from the tributaries of the middle reaches of the Yellow River flowing through the Loess Plateau as the research objects. Taking into account geographical distribution, underlying surface characteristics, and data availability, the catchment areas controlled by nine hydrological stations—Ansai, Gaojiabao, Hengshan, Liulin, Qingyangcha, Wuqi, Zaoyuan, Zhidan, and Zichang—were ultimately determined as the research units. Each sub-watershed was named after its corresponding hydrological station. The location of the study area in the middle reaches of the Yellow River is shown in the figure below. Figure 1 As shown.

[0027] The study area in this embodiment covers the northern, central, and southern parts of the Loess Plateau, and runs through multiple river basins, including the Weihe River, Yanhe River, and Beiluo River. This region is located in the transition zone between temperate monsoon climate and temperate continental climate, exhibiting significant monsoon characteristics, with simultaneous rainfall and heat, concentrated summer precipitation, frequent rainstorms, and obvious seasonal climate differences.

[0028] The runoff generation and runoff characteristics of a watershed are a comprehensive reflection of the interaction between meteorological and hydrological factors and underlying surface conditions. Therefore, the determination of watershed hydrological similarity should be systematically considered from three aspects: climate characteristics, runoff characteristics, and underlying surface characteristics. Regarding climate characteristics, rainfall is the main characteristic factor; runoff characteristics include indicators such as flow rate and sediment concentration; human activities affect the runoff and sediment generation processes of the watershed through direct water intake and alteration of the underlying surface. Especially in the middle reaches of the Yellow River, the construction of large-scale silt-retention dams has played an important role in flood control and disaster reduction, sediment retention and water storage, and consolidating the achievements of returning farmland to forest. In addition, the watershed area, as a basic geographical parameter, not only determines the total river runoff but also directly affects the runoff process. The watershed shape coefficient is a hydrological parameter reflecting the geometric morphology of the watershed and can be used to analyze flood evolution patterns. The topographic humidity index, through the functional relationship between slope and upstream catchment area, characterizes the control effect of topography on hydrological processes. Factors such as topographic slope and river gradient are related to hydrodynamic conditions.

[0029] In summary, this embodiment selects the watershed area and the watershed shape coefficient (…). Ke ), average slope, topographic humidity index ( TWI Eleven static topographic indicators, including river length, river gradient, cultivated land rate, forest land rate, grassland rate, terraced field ratio, and number of silt-retaining dams, were selected, along with three dynamic hydrological indicators, including rainfall, river runoff, and river sediment content, to jointly construct a multi-factor integrated watershed characteristic discrimination system.

[0030] The following is the method for similarity discrimination and clustering of small watersheds in the middle reaches of the Yellow River Basin as described in this embodiment, including the following process: S1, acquisition of static topographic indicators and dynamic hydrological indicators.

[0031] Using hydrological stations as outlets, the DEM data of the catchment area is extracted using ArcGIS software. The catchment area and catchment shape coefficient can then be calculated. Ke ), average slope, topographic humidity index ( TWIBased on a 1:250,000 Yellow River Basin water system dataset, river sections from the source of the main stream to hydrological stations were extracted using ArcGIS software, and their river lengths were calculated. River gradient was calculated using the ratio of the elevation difference to the horizontal distance from the source of the main stream to the hydrological station, with elevation data sourced from the basin's DEM. Based on 30-meter resolution land use raster data from the 1980s in China, ArcGIS software was used for mask extraction and reclassification of the study basin, statistically analyzing the areas of cultivated land, forest land, grassland, and other land types, and calculating their proportions. The proportion of the total watershed area; terrace data are derived from the aforementioned land use data, and in this embodiment, they are defined as the sum of two types of patches: "terraced land" and "dry slope land" with slopes between 3° and 25°; the number of silt-retaining dams is derived from the "Shaanxi Province Soil and Water Conservation Data Set (1949-1979)," and the total number of silt-retaining dams in each small watershed is obtained through spatial registration and statistics; hydrological sequence data (rainfall, runoff, sediment concentration) are all derived from the daily dataset of hydrological stations in the upper and middle reaches of the Yellow River released by the National Earth System Science Data Center. The data periods for each station are as follows: Anse Station (1981-1985), Gaojiabao, Hengshan, Liulin, Qingyangcha, Zaoyuan, Zhidan, Zichang, etc. (1975-1985), Wuqi Station (1980-1985). To eliminate the influence of differences in the dimensions and value ranges of each indicator on the analysis, the Z-Score standardization method is used to process the raw data.

[0032] S2, index weighting based on the coefficient of variation method.

[0033] To objectively determine the relative importance of each characteristic indicator in watershed similarity discrimination, this embodiment uses the coefficient of variation method for weighting. This method assigns weights based on the degree of variation of the indicator data itself: the greater the dispersion of the indicator data, the stronger its ability to distinguish samples, and therefore it should be given a higher weight. The coefficient of variation and weight of each indicator are calculated using equations (1) and (2), respectively: (1) (2) In the formula, E i As an indicator i coefficient of variation, σ i and μ i These are the standard deviation and mean of the indicator, respectively. ω i As an indicator i The weight, n This represents the total number of indicators.

[0034] S3, Watershed Similarity Calculation To quantify the overall similarity between watersheds, this embodiment uses weighted Euclidean distance to calculate the differences in the watershed feature space. Weighted Euclidean distance not only considers the differences in the values ​​of each feature index but also incorporates their importance (i.e., weight) in similarity judgment. Its calculation method is shown in equation (3): (3) In the formula, D ij For the basin i With the basin j The weighted Euclidean distance between them; ω k For the first k The weight of each indicator; x ik and x jk respectively watershed i and the basin j In the k Standardized values ​​for each indicator.

[0035] The similarity coefficient is calculated as shown in equation (4): (4) in: The similarity coefficient, For Euclidean distance, This represents the maximum distance.

[0036] After obtaining the similarity coefficients between each pair of watersheds, the degree of similarity between watersheds is qualitatively classified according to the hydrological similarity evaluation criteria (Table 1).

[0037] Table 1. Evaluation Criteria for Hydrological Similarity in Different Value Ranges

[0038] S4, Cluster Analysis.

[0039] To reveal the natural grouping patterns of small watersheds within the study area, this embodiment employs the K-means clustering algorithm, a classic unsupervised machine learning method that iteratively optimizes the data to divide samples into several clusters with high internal similarity and significant inter-group differences. A silhouette coefficient is then introduced. S i Evaluation to determine the optimal number of clusters k And evaluate the clustering quality. For a single sample i ,That S i The calculation formula is: (5) In the formula, a i For the sample iThe average distance to all other samples within the same cluster reflects the density of the cluster. b i For the sample i The average distance to all samples in its nearest neighbor cluster reflects the degree of separation between clusters. S i The value of is in the range of [-1, 1], and a larger value indicates a more reasonable clustering of the samples. For a complete clustering result, its overall quality is determined by the mean of the silhouette coefficients of all samples, i.e., the average silhouette coefficient. S k To measure: (6) In the formula, n This represents the total number of samples. By calculating different presets... k value corresponding S k And select to make S k Maximize k The value is used as the final cluster number to ensure the most reasonable watershed classification result is obtained.

[0040] The following analysis uses actual data to illustrate the results of the above method: S1, Analysis of static topographic indicators and dynamic hydrological indicators.

[0041] The raw data was processed using ArcGIS software to obtain land use information for each small watershed. Based on the aforementioned methods, 11 static topographic indicators and 3 dynamic hydrological indicators were obtained, as detailed in Table 2. Figure 2 .

[0042] Table 2. Watershed Topographic Indicators Data

[0043] Analysis shows that the topographic and hydrological characteristics of the various watersheds exhibit both similarities and significant differences. Regarding topographic indicators, factors such as the watershed shape coefficient and topographic humidity index show relatively small differences, while factors such as watershed area, cultivated land ratio, and forest land ratio show significant differences and exhibit obvious spatial clustering characteristics. In terms of hydrological indicators, rainfall varies little among the watersheds, which is related to their location within the same climatic unit; however, river runoff and sediment load differ significantly, with some extreme values ​​observed.

[0044] S2, Watershed similarity discriminant analysis.

[0045] Calculate the coefficient of variation and weight of each indicator, and draw a weight distribution diagram. Figure 3 The results showed that indicators such as river sediment content, runoff, forest coverage, and watershed area had relatively high weights in the discrimination system, indicating that these factors had a more significant impact on watershed similarity.

[0046] The similarity between each pair of watersheds was further calculated, and a similarity heatmap was drawn. Figure 4 Based on the hydrological similarity evaluation criteria (Qi Xiaoming et al., 2007), five pairs of watersheds reached the level of "basic similarity" or above. Among them, Ansai-Wuqi and Zhidan-Wuqi were rated as "basic similarity"; Liulin-Zaoyuan, Zhidan-Qingyangcha, and Zhidan-Ansai were rated as "relatively similar" (Table 3). Two pairs of watersheds with "basic similarity" (Ansai-Wuqi and Wuqi-Zhidan) were selected, and their nine main similarity indicators were visually compared (Figure 5). The results show that the similarity between the two pairs of watersheds in terms of topographic indicators is higher than that in terms of hydrological indicators. Among them, the slope, land use structure, and topographic humidity index are the most similar, and the difference of all indicators is within 100%, indicating that their overall similarity is good.

[0047] Table 5. Characteristic Evaluation of the Most Similar Watersheds

[0048] Besides using the coefficient of variation (CVM) method, subjective weighting, entropy weighting, or CRITIC methods can be combined to construct a comprehensive subjective-objective weighting system. Appropriate handling of outliers and missing values ​​before analysis can improve the robustness of weight allocation. The Euclidean distance method is widely used in hydrological similarity assessment, and the results of this embodiment verify its effectiveness. However, this method focuses on measuring absolute numerical differences. Introducing more directional (vector) features allows for a more comprehensive measurement of the similarity between watersheds by combining Euclidean distance with cosine distance.

[0049] S3, Watershed Similarity Cluster Analysis.

[0050] via K MESSING cluster analysis determined the optimal number of clusters to be 3, with a silhouette coefficient of 0.467. At this point, the similarity within each cluster is high, while the differences between clusters are significant. The cluster distribution after dimensionality reduction of the feature data (cumulative principal component contribution rate of 80.7%) is shown in the figure. Figure 6a Radar is plotted by combining key indicators with higher weightings. Figure 6b The nine small watersheds can be divided into three categories: Cluster I (large watershed group: Gaojiabao, Hengshan): large watershed area, low average slope, high runoff and sediment content, relatively humid hydrological conditions, and low proportion of cultivated land and forest land; Cluster II (medium watershed group: Ansai, Wuqi, Zhidan, Qingyangcha): medium watershed area, medium slope and runoff sediment content, and high proportion of cultivated land and grassland; Cluster III (small watershed group: Liulin, Zaoyuan, Zichang): small watershed area, large slope, high annual rainfall, high proportion of forest land, and large river gradient.

[0051] To evaluate the clustering effect, six key indicators were selected to compare the changes in the standard deviation and range of each category before and after clustering. Figure 7The results showed that the range and standard deviation of most indicators in each category decreased significantly after clustering, indicating that classification effectively reduced the dispersion within groups. The standard deviation of a few indicators decreased less or increased slightly, mainly due to the reduced sample size after clustering and the continued existence of numerical differences within groups. This also suggests that increasing the sample size helps to further improve the stability and accuracy of clustering.

[0052] The clustering results in this embodiment conform to the natural geographical differentiation patterns of small watersheds in the middle reaches of the Yellow River Basin. The characteristics within each category are similar, while the differences between categories are significant. This is in good agreement with the actual hydrological and geomorphological characteristics and can provide a categorized reference for the comprehensive management and soil and water conservation planning of small watersheds in this region.

[0053] In summary, to address the challenge of hydrological forecasting in areas with limited or no data, this embodiment selected nine small watersheds in the Loess Plateau region of the middle reaches of the Yellow River. A watershed characteristic system was constructed, comprising 11 static topographic indicators and 3 dynamic hydrological indicators. Using weighted Euclidean distance and K-means clustering methods, watershed similarity discrimination and classification were conducted. The similarity discrimination results showed that five pairs of watersheds reached a level of "basic similarity" or higher. Among them, Ansai-Wuqi and Zhidan-Wuqi showed the highest similarity (basic similarity), followed by Liulin-Zaoyuan, Zhidan-Qingyangcha, and Zhidan-Ansai (relatively similar). Cluster analysis determined the optimal number of clusters to be 3, with a silhouette coefficient of 0.467. The nine small watersheds can be divided into three categories: large watershed type (Gaojiabao, Hengshan), medium watershed type (Ansai, Wuqi, Zhidan, Qingyangcha), and small watershed type (Liulin, Zaoyuan, Zichang). Their characteristics correspond to different geomorphological-hydrological combination patterns. The classification results are in good agreement with the actual geographical conditions and can provide a classification basis for regional soil and water conservation and watershed management.

[0054] The following are embodiments of the apparatus of the present invention, which can be used to execute embodiments of the method of the present invention. For details not omitted in the apparatus embodiments, please refer to the embodiments of the method of the present invention.

[0055] In another embodiment of the present invention, a similarity discrimination and clustering system for small watersheds in the middle reaches of the Yellow River Basin is provided. This system can be used to implement the aforementioned similarity discrimination and clustering method for small watersheds in the middle reaches of the Yellow River Basin. Specifically, the system includes an indicator data acquisition module, an indicator weight determination module, a similarity calculation module, and a clustering module.

[0056] The indicator data acquisition module is used to acquire static topographic indicator data and dynamic hydrological indicator data of the target small watershed, and to standardize the static topographic indicator data and dynamic hydrological indicator data to obtain standardized data.

[0057] The indicator weight determination module is used to calculate the coefficient of variation of each static topographic indicator and dynamic hydrological indicator, and to determine the weight of each static topographic indicator and dynamic hydrological indicator based on the coefficient of variation.

[0058] The similarity calculation module is used to calculate the feature space distance between small watersheds based on weighted and standardized data using a weighted Euclidean distance algorithm, and to determine the similarity coefficient of the small watersheds based on the feature space distance.

[0059] The clustering module is used to group small watersheds using the K-means clustering algorithm, calculate the silhouette coefficient to determine the optimal number of clusters, and obtain the small watershed similarity clustering results.

[0060] In another embodiment of the present invention, a terminal device is provided, comprising a processor and a memory. The memory stores a computer program, the computer program including program instructions, and the processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to... The processor described in this embodiment of the invention can be used for the operation of a similarity discrimination and clustering method for small watersheds in the middle reaches of the Yellow River Basin, including: acquiring static topographic index data and dynamic hydrological index data of the target small watershed; standardizing the static topographic index data and dynamic hydrological index data to obtain standardized data; calculating the coefficient of variation of each static topographic index and dynamic hydrological index; determining the weight of each static topographic index and dynamic hydrological index based on the coefficient of variation; calculating the feature space distance between small watersheds using a weighted Euclidean distance algorithm based on the weight and standardized data; determining the small watershed similarity coefficient based on the feature space distance; grouping the small watersheds using a K-means clustering algorithm; calculating the silhouette coefficient to determine the optimal number of clusters; and obtaining the small watershed similarity clustering result.

[0061] In another embodiment, the present invention also provides a computer-readable storage medium (Memory), which is a memory device in a terminal device for storing programs and data. It is understood that the computer-readable storage medium here may include both the built-in storage medium in the terminal device and extended storage media supported by the terminal device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor, which may be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here may be high-speed RAM or non-volatile memory, such as at least one disk storage device.

[0062] One or more instructions stored in a computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the method for similarity discrimination and clustering of small watersheds in the middle reaches of the Yellow River Basin in the above embodiments. One or more instructions in the computer-readable storage medium are loaded and executed by the processor in the following steps: obtaining static topographic index data and dynamic hydrological index data of the target small watershed; standardizing the static topographic index data and dynamic hydrological index data to obtain standardized data; calculating the coefficient of variation of each static topographic index and dynamic hydrological index; determining the weight of each static topographic index and dynamic hydrological index based on the coefficient of variation; calculating the feature space distance between small watersheds using the weighted Euclidean distance algorithm based on the weight and standardized data; determining the small watershed similarity coefficient based on the feature space distance; grouping the small watersheds using the K-means clustering algorithm; calculating the silhouette coefficient to determine the optimal number of clusters; and obtaining the small watershed similarity clustering results.

[0063] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0064] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0065] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0066] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0067] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0068] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0069] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0070] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0071] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

[0072] It should be understood that the above description is for illustrative purposes and not for limitation. Many embodiments and applications beyond the provided examples will be apparent to those skilled in the art upon reading the above description. Therefore, the scope of this patent should not be determined by reference to the above description, but rather by reference to the foregoing claims and the full scope of their equivalents. For purposes of completeness, all articles and references, including patent applications and publications, are incorporated herein by reference. The omission of any aspect of the subject matter disclosed herein in the foregoing claims is not intended as a waiver of that subject matter, nor should it be construed as an indication that the applicant has not considered that subject matter as part of the disclosed inventive subject matter.

Claims

1. A method for similarity discrimination and clustering of small watersheds in the middle reaches of the Yellow River Basin, characterized in that, include: Obtain static topographic data and dynamic hydrological data of the target small watershed, and standardize the static topographic data and dynamic hydrological data to obtain standardized data; Calculate the coefficient of variation for each static topographic index and dynamic hydrological index, and determine the weight of each static topographic index and dynamic hydrological index based on the coefficient of variation. Based on weighted and standardized data, a weighted Euclidean distance algorithm is used to calculate the feature space distance between small watersheds, and the similarity coefficient of small watersheds is determined based on the feature space distance. The K-means clustering algorithm was used to group the small watersheds, and the silhouette coefficient was calculated to determine the optimal number of clusters, thus obtaining the small watershed similarity clustering results.

2. The method for similarity discrimination and clustering of small watersheds in the middle reaches of the Yellow River Basin according to claim 1, characterized in that, Static topographic indicators include watershed area, watershed shape coefficient, average slope, topographic humidity index, river length, river gradient, cultivated land ratio, forest land ratio, grassland ratio, terrace ratio, and number of silt-retaining dams; dynamic hydrological indicators include rainfall, river runoff, and river sediment content.

3. The method for similarity discrimination and clustering of small watersheds in the middle reaches of the Yellow River Basin according to claim 1, characterized in that, The standardization process employs the Z-Score standardization method to eliminate the dimensional differences between static topographic indicators and dynamic hydrological indicators.

4. The method for similarity discrimination and clustering of small watersheds in the middle reaches of the Yellow River Basin according to claim 1, characterized in that, The process of determining the weights of each static topographic index and dynamic hydrological index includes: In the formula, E i As an indicator i coefficient of variation, σ i and μ i These are the standard deviation and mean of the indicator, respectively. ω i As an indicator i The weight, n This represents the total number of indicators.

5. The method for similarity discrimination and clustering of small watersheds in the middle reaches of the Yellow River Basin according to claim 1, characterized in that, The process of calculating the characteristic spatial distance between small watersheds includes: In the formula, D ij For the basin i With the basin j The weighted Euclidean distance between them; ω k For the first k The weight of each indicator; x ik and x jk respectively watershed i and the basin j In the k Standardized values ​​for each indicator.

6. The method for similarity discrimination and clustering of small watersheds in the middle reaches of the Yellow River Basin according to claim 5, characterized in that, The process of determining the similarity coefficient of a small watershed includes: in: The similarity coefficient, For Euclidean distance, This represents the maximum distance.

7. The method for similarity discrimination and clustering of small watersheds in the middle reaches of the Yellow River Basin according to claim 1, characterized in that, The process of calculating the silhouette coefficient to determine the optimal number of clusters includes: In the formula, a i For the sample i The average distance to all other samples within the same cluster reflects the density of the cluster. b i For the sample i The average distance to all samples in its nearest neighbor cluster reflects the degree of separation between clusters; S i The value of is in the range of [-1, 1], and the larger the value, the more reasonable the clustering of the sample. For a complete clustering result, its overall quality is determined by the mean of the silhouette coefficients of all samples, i.e., the average silhouette coefficient. S k To measure: In the formula, n Given the total number of samples, different preset values ​​are calculated. k value corresponding S k And select to make S k Maximize k The value is used as the final cluster number.

8. A similarity discrimination and clustering system for small watersheds in the middle reaches of the Yellow River Basin, characterized in that, include: The indicator data acquisition module is used to acquire static topographic indicator data and dynamic hydrological indicator data of the target small watershed, and to standardize the static topographic indicator data and dynamic hydrological indicator data to obtain standardized data. The indicator weight determination module is used to calculate the coefficient of variation of each static topographic indicator and dynamic hydrological indicator, and determine the weight of each static topographic indicator and dynamic hydrological indicator based on the coefficient of variation. The similarity calculation module is used to calculate the feature space distance between small watersheds based on weighted and standardized data using a weighted Euclidean distance algorithm, and to determine the similarity coefficient of the small watersheds based on the feature space distance. The clustering module is used to group small watersheds using the K-means clustering algorithm, calculate the silhouette coefficient to determine the optimal number of clusters, and obtain the small watershed similarity clustering results.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the similarity discrimination and clustering method for small watersheds in the middle reaches of the Yellow River Basin as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the similarity discrimination and clustering method for small watersheds in the middle reaches of the Yellow River Basin as described in any one of claims 1 to 7.