Method for clustering and partitioning hydrological similarity of tributaries in reservoir area

By combining the SWAT model and SCS runoff curve method with fuzzy C-means clustering, the confluence dynamic coefficients are quantified, and the comprehensive feature vector of tributaries is generated. This solves the data sparsity problem in the hydrological similarity analysis of large reservoir areas and achieves accurate partitioning and simulation under low data density.

CN121834385APending Publication Date: 2026-04-10CHINA INST OF WATER RESOURCES & HYDROPOWER RES
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies, in the analysis of hydrological similarity in large reservoir areas, cannot comprehensively consider both runoff generation and confluence similarity, ignore factors such as confluence paths and river morphology, and are highly dependent on data, making it impossible to effectively transfer parameters under conditions of sparse hydrological stations.

Method used

The SWAT model was used to divide the watershed into sub-basins. The SCS runoff curve method and fuzzy C-means clustering algorithm were combined. The runoff dynamic coefficients were quantified using DEM, remote sensing images and soil data to generate tributary comprehensive feature vectors for similarity analysis and zoning.

Benefits of technology

Under extremely low data density conditions, it accurately reflects the similarity of hydrological functions of tributaries in the reservoir area, improves the scientificity and practical value of the zoning results, and is applicable to large reservoir areas with sparse hydrological stations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834385A_ABST
    Figure CN121834385A_ABST
Patent Text Reader

Abstract

The invention discloses a reservoir area tributary hydrological similarity clustering and partitioning method. The method comprises the steps that S1, a reservoir area drainage basin is divided into a plurality of sub-drainage basins through an SWAT model, and the sub-drainage basins are further divided into one or more hydrological response units; s2, carrying out slope runoff production calculation by adopting an SCS runoff curve method, and extracting key factors of a hydrological response unit; s3, taking the standardized key factors as attribute features, and performing spatial clustering analysis on all sub-basins of the reservoir area by adopting a fuzzy C-means clustering algorithm to complete sub-basin classification; s4, calculating a confluence feature coefficient of each sub-basin, and extracting composition factors of the confluence feature coefficients; s5, determining a comprehensive feature vector of the branch; and S6, calculating a similarity matrix among the branches, and dividing the branches of which the similarity is not lower than a set threshold value into the same hydrological similar category. According to the method, effective reservoir area branch partitioning can be achieved under the condition of extremely low data density, and a feasible analysis tool can be provided for large reservoir areas with sparse hydrologic stations such as Three Gorges.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of reservoir area hydrological simulation and water resources management technology, and in particular relates to a method for clustering and partitioning tributaries of reservoir areas based on hydrological similarity. Background Technology

[0002] Watershed hydrological similarity is a hot topic in hydrological research. If two or more watersheds are similar in key features such as topography, soil, vegetation, and river network structure, their "hydrological response" processes to inputs such as rainfall and runoff will also tend to be consistent, thus indicating watershed hydrological similarity. Therefore, watershed hydrological similarity can be used to transfer model parameters and hydrological information from watersheds with available data to watersheds without data, thereby conducting runoff prediction in data-scarce areas. With the rapid development of watershed hydrological models and remote sensing geographic information technology, parameter transfer based on hydrological similarity has become a key approach to solving the challenges of hydrological simulation and forecasting in data-scarce or data-poor areas. Its core lies in identifying watersheds or sub-regions with similar hydrological response characteristics, extending reliable model parameters calibrated at a limited number of observation stations to data-scarce areas, thereby improving the accuracy of overall simulation.

[0003] Chinese patent CN113887635A discloses a watershed similarity classification method, proposing a relatively universal watershed similarity classification framework. This method first extracts a wide range of meteorological and underlying surface factors, uses distance correlation coefficients to screen indicators, and then employs a two-stage clustering algorithm (SOM-FCM) combining self-organizing mapping and fuzzy C-means to perform two-level clustering at both the grid scale and sub-watershed scale. Finally, it uses the maximum-minimum proximity method to assess the similarity between watersheds. While this method provides a systematic analysis process, it focuses on constructing a general and universal similarity discrimination system. The two-stage process is complex and computationally expensive. Furthermore, during the clustering process, it does not explicitly consider the water flow process within the watershed, especially the differences in river confluence dynamics after runoff generation in sub-watersheds. It also does not optimize for the specific hydrological conditions of specific types of watersheds (such as artificially regulated large reservoir areas). For reservoir areas, their tributaries are affected by backwater, experience significant water level fluctuations, and have complex hydrodynamic conditions. Similarity analysis based solely on static underlying surface properties is insufficient to accurately reflect their unique runoff generation and confluence mechanisms, particularly the differences in the confluence process. Chinese Patent CN118114567A discloses a multi-site watershed hydrological model parameter integration and calibration method based on reinforcement learning. This method constructs hydrological similarity elements and calculates comprehensive similarity by dividing multiple station watersheds into clusters and classifying internal sub-watershed units based on similarity. It then uses the DDPG reinforcement learning algorithm to integrate and calibrate the model parameters within the clusters. This method effectively combines similarity analysis with parameter optimization. However, this method is mainly applicable to station watershed clusters with multiple hydrological stations and relatively abundant data, provided that sufficient data units can be found within the cluster for parameter learning and transfer. For specific scenarios such as large reservoir areas with extremely sparse hydrological stations, only a few tributaries with stations, and tributaries significantly affected by uniform backwater flow, the cluster-based parameter learning model relied upon by this method is difficult to apply directly. Furthermore, this method focuses on transferring process simulation parameters for storm runoff, requiring extremely high-quality baseline data on rainfall responses, thus lacking universality. In summary, the existing technology still has the following shortcomings:

[0004] (1) Traditional similarity methods (such as AHP-based weighted summarization of multiple independent indicators) are difficult to comprehensively reflect the entire process of runoff generation and confluence. Furthermore, the weight setting is highly subjective and does not fully consider the non-natural hydrodynamic conditions in the reservoir area under artificial scheduling, such as backwater, flow rate slowdown, and sediment deposition changes. This may cause the similarity judgment based on traditional underlying surface indicators to fail in the reservoir area.

[0005] (2) The parameter integration calibration method is highly dependent on data and requires machine learning to be carried out at a large number of data stations. This is inconsistent with the reality that there are many tributaries in large reservoir areas such as the Three Gorges and the hydrological stations are not matched. There is a lack of feasible solutions under low data density.

[0006] (3) In terms of analysis dimensions, existing methods mostly focus on the similarity of runoff generation attributes such as climate and underlying surface. The analysis dimensions are singular, and the decisive influence of runoff dynamic factors such as confluence path, river morphology and hydraulic gradient on the final hydrological response of reservoir tributaries is generally ignored.

[0007] (4) Since land use types are difficult to represent numerically, existing technologies often do not consider land use patterns when extracting feature values, resulting in imperfect feature value extraction for similarity analysis.

[0008] Therefore, given the hydrogeographic characteristics of reservoir areas, there is a need for a tributary hydrological similarity analysis and zoning method that can comprehensively consider both runoff generation and runoff confluence similarities, select appropriate parameters to represent land use type characterization, and ultimately serve spatial zoning management, under the condition of extremely limited hydrological station data. Summary of the Invention

[0009] To address the aforementioned issues, this invention aims to overcome the limitations of existing watershed hydrological similarity and parameter transfer methods in the special scenarios of large reservoir areas, and to provide a tributary hydrological functional zoning method specifically tailored to the hydrological characteristics, data constraints, and precise management needs of reservoir areas.

[0010] To achieve the objective of this invention, the following technical solution is adopted:

[0011] A method for clustering and partitioning tributaries based on hydrological similarity in reservoir areas, the steps of which are as follows:

[0012] S1. Based on the DEM data, river network vector data and watershed outlet point data of the reservoir area, the reservoir area watershed is divided into multiple sub-watersheds using the SWAT model. Furthermore, each sub-watershed is divided into one or more hydrological response units by overlaying land use type map, slope grade map and soil type map data.

[0013] S2. The SCS runoff curve method is used to calculate slope runoff and determine the key factors reflecting the runoff capacity of the sub-watershed. Taking the sub-watershed as the unit, the key factors of all hydrological response units in the sub-watershed are extracted. The key factors include soil particle size composition, CN value, NDVI index, and area of ​​hydrological response unit. The weighted average of soil particle size composition, CN value, and NDVI index of the sub-watershed is calculated based on the area of ​​hydrological response unit and used as the key runoff factor of the sub-watershed. The factor is then standardized.

[0014] S3. Using standardized key factors as attribute features, spatial clustering analysis was performed on all sub-basins of the reservoir area using the fuzzy C-means clustering algorithm. The effectiveness of the Xie-Beni clustering function was then calculated. Parameter values, drawing The evaluation curve of parameter change with the number of clusters is used to determine the optimal number of clusters based on the inflection point of the curve or the number of clusters that are flat and have the best clustering effect, and the sub-basin classification is completed.

[0015] S4. Considering the runoff characteristics of the reservoir area affected by backwater, the runoff of the sub-basin is regarded as a water ball about to enter the river channel, and the river channel is generalized as a slope. Ignoring the influence of the difference in river runoff roughness, the runoff characteristic coefficient of each sub-basin is calculated. Taking the sub-basin as the unit, the component factors of the runoff characteristic coefficient are extracted. The component factors include the sub-basin area, average slope, and slope length from the sub-basin to the tributary outlet, which are used to quantify the transport capacity of the runoff of the sub-basin to the tributary outlet.

[0016] S5. Multiply the fuzzy membership matrix of each sub-basin by the corresponding confluence feature coefficient to obtain the weighted feature vector of the sub-basin under the classification of its tributary; sum the feature vectors of all sub-basins of the same tributary to form the comprehensive feature vector of the tributary.

[0017] S6. Based on the comprehensive feature vector of the tributaries, calculate the similarity matrix between each tributary, set a similarity threshold, classify tributaries with similarity not lower than the threshold into the same hydrological similarity category, and divide the tributaries flowing into the reservoir into several hydrological functional similarity zones, wherein the tributaries in the same zone have similar runoff generation attributes and confluence dynamics characteristics.

[0018] Furthermore, the confluence characteristic coefficient mentioned in step S5 The calculation formula is as follows:

[0019]

[0020] in, Indicates the first The catchment area of ​​each sub-basin For the first Length of the river channel from the sub-basin to the tributary outlet Indicates the first Average riverbed slope in each sub-basin.

[0021] Furthermore, the specific method for determining the optimal number of clusters in step S3 is as follows: Calculate the number of clusters sequentially from 2 to 20. Parameter values, drawing The curve of parameters changing with the number of clusters; when When the downward trend of the value slows down significantly and the curve shows a stable segment, the number of clusters corresponding to the starting point of the stable segment is selected as the optimal number of clusters.

[0022] Furthermore, the similarity threshold mentioned in step S6 can be set to 0.8.

[0023] Furthermore, the CN value mentioned in S2 is a dimensionless empirical parameter used to comprehensively reflect the watershed's underlying surface's "interception-runoff generation" capacity in the rainfall-runoff relationship. It is an integer between 0 and 100 that comprehensively reflects three major factors: soil type, land use / vegetation, and previous wet conditions (AMC).

[0024] Furthermore, when the similarity between two tributaries is greater than or equal to the similarity threshold of 0.8, they are judged to be highly similar and classified into the same type of hydrological similarity zone.

[0025] Furthermore, in step S6, the similarity of hydrological functions is measured by calculating the similarity between the comprehensive feature vectors of each tributary.

[0026] Furthermore, after step S3, the silhouette coefficient function is used to calculate the clustering quality and determine the quality of the clustering results; for each data point, its distance to other points in the same cluster and its distance to other clusters are calculated, and the basic formula for its silhouette coefficient is obtained:

[0027]

[0028] in, Represents sample points The average distance to all other points within its cluster. Represents sample points The minimum average distance to all sample points in any other cluster is the minimum value. When the silhouette coefficient is close to 1, the clustering result is better, and when the silhouette coefficient is closer to -1, the clustering result is worse.

[0029] Furthermore, the method also includes S7. Within the same hydrologically similar area, at least one tributary with a hydrological monitoring station is selected as a reference watershed, and its calibrated hydrological model parameters are transferred to the tributary in the area without data, thereby completing the construction of hydrological model parameters for the entire target watershed.

[0030] Furthermore, the several hydrologically similar zones obtained by the method can also be used for constructing hydrological models of lake and reservoir watersheds, calibrating and verifying parameters, simulating inflow hydrological processes, and providing a basis for subsequent watershed non-point source pollution simulation.

[0031] The beneficial effects of this invention are:

[0032] By introducing the confluence dynamics coefficient, the differences in the confluence process affected by backwater in the reservoir area were quantified for the first time in similarity analysis. This achieved a dual consideration of runoff generation attributes and confluence dynamics, and more accurately reflected the hydrological structure within the basin. Under the special hydrodynamic environment of the reservoir area, the similarity judgment of tributary hydrological functions is more scientific and accurate, overcoming the defects of distortion when general methods are applied in the reservoir area.

[0033] Instead of relying on a large amount of data on tributaries for machine learning or parameter calibration, this method is based on readily available DEM, remote sensing images, and soil data. It achieves effective tributary zoning of reservoir areas under extremely low data density conditions through a path of sub-basin clustering → tributary vector construction → similarity comparison. This provides a practical analytical tool for large reservoir areas with sparse hydrological stations, such as the Three Gorges Dam.

[0034] This method creatively couples (multiplies) the fuzzy clustering membership degree of sub-basins with the confluence dynamics coefficient to generate a comprehensive tributary feature vector that simultaneously reflects the similarity of static attributes and dynamic processes. This expands the dimension and depth of existing similarity analysis methods, and the zoning results more realistically reflect the complex hydrological functions of reservoir tributaries, significantly enhancing the practical value of the zoning results in hydrological simulation.

[0035] This method creatively uses CN value and NDVI index to represent land use type, a key factor influencing runoff generation, thus solving the problem of existing methods that do not include land use type as a feature value for cluster analysis due to the difficulty in quantifying it. It employs cosine similarity between the comprehensive feature vectors of tributaries for measurement, a method that is objective and computationally efficient. This overcomes the shortcomings of traditional similarity analysis, which relies on subjective weighting (such as AHP) or complex regression, making the zoning results more stable and reliable. Attached Figure Description

[0036] Figure 1 This is a flowchart of the method steps of the present invention;

[0037] Figure 2 This is an embodiment of the invention, specifically the Three Gorges Reservoir area. Graph showing how parameters change with increasing number of cluster groups;

[0038] Figure 3 This is a diagram showing the clustering quality evaluation results of the Three Gorges Reservoir area according to an embodiment of the present invention;

[0039] Figure 4 This is a diagram showing the similarity analysis results among 30 tributaries in the Three Gorges Reservoir area according to an embodiment of the present invention;

[0040] Figure 5 This is a classification result diagram of 30 tributaries in the Three Gorges Reservoir area according to an embodiment of the present invention;

[0041] Figure 6 This is a zoning result diagram of 30 tributaries in the Three Gorges Reservoir area according to an embodiment of the present invention. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0043] The SCS (Soil Conservation Service Curve Number Method) involved in this invention refers to an empirical method proposed by the U.S. Department of Agriculture's Soil and Water Conservation Service to estimate surface runoff based on land use, soil type, and antecedent moisture conditions. Its core parameter is the runoff curve number (CN).

[0044] Example 1

[0045] This embodiment describes a method for clustering and partitioning tributaries based on hydrological similarity in a reservoir area. Figure 1 As shown, the method includes the following steps:

[0046] S1. Based on the DEM data, river network vector data and watershed outlet point data of the reservoir area, the reservoir area watershed is divided into multiple sub-watersheds using the SWAT model. Furthermore, by overlaying land use type map, slope grade map and soil type map data, each sub-watershed is further divided into one or more hydrological response units.

[0047] To achieve refined hydrological similarity analysis of the reservoir basin, the continuous geographic space must first be discretized into a series of hydrologically independent basic computational units. This invention preferably uses the SWAT (Soil and Water Assessment Tool) model for subbasin and hydrological response unit (HRU) division. The SWAT model is a long-term, basin-scale distributed hydrological model. It assumes that the hydrological responses of each subbasin or subunit within the basin are independent. The hydrological response of the entire basin is obtained by overlaying the hydrological responses of each subbasin or subunit. By inputting basin elevation information, the actual internal river network, and the basin outlet point, the basin is divided into multiple subbasins. To reflect the spatial heterogeneity of the underlying surface within each subbasin, land use type maps, slope grade maps, and soil data are spatially overlaid within each subbasin. Then, based on user-defined thresholds (such as the percentage of minimum area for each category), each subbasin is further subdivided into one or more HRUs. Simultaneously, an attribute database is constructed based on long-term meteorological data of the watershed to simulate and calculate processes such as water flow, sediment transport, and nutrient migration within the watershed. Sub-watersheds serve as the basic units for subsequent cluster analysis, while the HRU partitioning information provides the foundation for extracting spatial statistics of key factors such as soil particle size composition within sub-watersheds in step S2.

[0048] S2. The SCS runoff curve method is used to calculate the slope runoff generation, and the key factors reflecting the runoff generation capacity of the sub-basin are determined. The key factors are extracted on a sub-basin basis. The key factors include at least the sub-basin area, average slope, soil particle size composition and CN value. The extracted key factors are then standardized.

[0049] The SWAT simulation of the watershed hydrological cycle is mainly divided into the land surface hydrological component of the water cycle, namely watershed runoff and slope runoff, and the water surface component of the water cycle, namely river runoff. The former controls the input to the main channel within each sub-watershed; the latter determines the transport and collection motion from the river network to the watershed outlet. When precipitation occurs in the watershed, some precipitation is intercepted by the vegetation canopy, and some falls to the ground as runoff or infiltrates into the soil. Some of the infiltrated water evaporates or is absorbed and utilized by soil roots, or continues to infiltrate and remain in the soil or flows laterally into the river channel. The water balance equation used in the SWAT model is:

[0050]

[0051] In the formula, , Soil moisture content at the beginning and end of the time period, respectively, in mm; For the first Rainfall for the day, in mm; For the first The depth of surface runoff on the day, in mm; For the first Evapotranspiration per day, mm; For the first The amount of water entering the vadose zone from the soil profile, in mm; For the first The amount of water flowing back to the sluice gate, in mm. Similar to traditional models, the SWAT model is a distributed application of lumped models, and the data is spatially dispersed. However, its parameters (especially runoff parameters) still reflect the overall underlying surface characteristics of a specific sub-basin, and cannot reflect the interception and infiltration processes of precipitation on the underlying surface under different vegetation and soil combinations, or the effects of various vegetation and soil on evapotranspiration.

[0052] This invention uses SCS to calculate slope runoff, and the calculation formula is as follows:

[0053]

[0054]

[0055] in, Surface runoff, in mm; Rainfall amount, in mm; This represents the initial loss, in mm. The infiltration rate is expressed in mm.

[0056] Early losses Influenced by factors such as land use, interception of branches and leaves, infiltration, and filling of depressions, it is directly proportional to the soil's saturated water storage (potential infiltration), that is... The U.S. Soil and Water Conservation Service proposed an appropriate ratio. ,Right now: Infiltration rate Using empirical parameters This is reflected in the equation: 𝑆=25400 / 𝐶𝑁−254.

[0057] The core parameter of the SCS model is a dimensionless empirical parameter used to comprehensively reflect the watershed's underlying surface's "interception-runoff generation" capacity in the rainfall-runoff relationship. It integrates three major factors—soil type, land use / vegetation, and anterior moisture condition (AMC)—into an integer between 0 and 100.

[0058] CN values ​​are closely related to underlying surface factors such as land use type, soil type, and antecedent soil moisture. Since land use type is difficult to represent numerically, CN values ​​and vegetation cover are used directly as substitutes. The particle size distribution (average proportions of CLAY, SAND, and SILT), CN value, and NDVI index are extracted for each sub-watershed. After slope runoff, the amount of water flowing into the river channel is directly related to the sub-watershed area and slope; therefore, sub-watershed area and slope are extracted as key factors. Thus, the key factors determining the runoff of a sub-watershed include sub-watershed area, particle size distribution, CN value, and slope. Fuzzy clustering analysis is used to cluster these factors. Data standardization should be performed before clustering analysis to avoid clustering bias due to differences in the order of magnitude of the parameters.

[0059] S3. Using standardized key factors as attribute features, spatial clustering analysis was performed on all sub-basins of the reservoir area using the fuzzy C-means clustering algorithm. The effectiveness of the Xie-Beni clustering function was then calculated. Parameter values, drawing The evaluation curve of parameter variation with the number of clusters is used to determine the optimal number of clusters based on the inflection point of the curve or the number of clusters that flattens out and achieves the best clustering effect, thus completing the sub-basin classification.

[0060] As a type of multivariate statistical analysis, cluster analysis has been widely applied in fields such as pattern recognition and data mining. Given a dataset with a known clustering trend, clustering can be performed using appropriate algorithms. However, the validity of the clustering results needs to be analyzed. Typically, cluster validity analysis can be transformed into the automatic determination of the number of clusters *c* and the fuzzy weight coefficients *m*.

[0061] This embodiment preferably adopts the approach proposed by Xie and Beni. The parameter characterizes the clustering effect. This parameter is the first clustering effectiveness function that takes the structure of the dataset into account. It is the ratio of in-class compactness to between-class separation. The larger the between-class distance, the more dispersed the clusters are; the smaller the within-class distance, the more compact the clusters are. The minimum value indicates the most efficient clustering result. The specific method for determining the optimal number of clusters is to calculate the number of clusters sequentially from 2 to 20. Parameter values, drawing The curve of parameters changing with the number of clusters; when When the downward trend of the value slows down significantly and the curve shows a stable segment, the number of clusters corresponding to the starting point of the stable segment is selected as the optimal number of clusters.

[0062] Based on the above clustering results, this invention further employs the silhouette coefficient function to calculate the clustering quality and determine the quality of the clustering results. The method involves calculating the distance to other points in the same cluster and the distance to other clusters for each data point to obtain its silhouette coefficient. The basic formula used is:

[0063]

[0064] in, Represents sample points The average distance to all other points within its cluster (ambiguity). Represents sample points The minimum average distance to all sample points in any other cluster (density). When the silhouette coefficient is close to 1, the clustering result is better; when the silhouette coefficient is closer to -1, the clustering result is worse.

[0065] It should be noted that when verifying the quality of the clustering results based on fuzzy C-means, the fuzzy clustering results need to be 'hardened' first. That is, for each sub-basin, the category with the highest membership degree is determined as the final category of that sub-basin. On this basis, the silhouette coefficient is then used to evaluate the overall clustering quality of all sub-basins.

[0066] S4. Considering the runoff characteristics of the reservoir area affected by backwater, the runoff of the sub-basin is regarded as a water ball about to enter the river channel, and the river channel is generalized as a slope. Ignoring the influence of the difference in river runoff roughness, the runoff characteristic coefficient of each sub-basin is calculated. Taking the sub-basin as the unit, the component factors of the runoff characteristic coefficient are extracted. The component factors include the sub-basin area, average slope, and slope length from the sub-basin to the tributary outlet, which are used to quantify the transport capacity of the runoff of the sub-basin to the tributary outlet.

[0067] In addition to considering runoff similarity, the hydrological similarity of this invention also considers confluence similarity. The amount and time of runoff reaching the basin outlet from different sub-basins determine the similarity of confluence. Considering the confluence characteristics of the reservoir area affected by backwater, the runoff from each sub-basin is regarded as a water sphere about to enter the river channel, and the river channel is generalized as a slope. The influence of differences in river confluence roughness is ignored, and the confluence characteristic coefficient of each sub-basin is calculated. The confluence characteristic coefficient... The calculation formula is as follows:

[0068]

[0069] in, Indicates the first The catchment area of ​​each sub-basin For the first Length of the river channel from the sub-basin to the tributary outlet Indicates the first Average riverbed slope in each sub-basin.

[0070] S5. Multiply the fuzzy membership matrix of each sub-basin by the corresponding confluence characteristic coefficient to obtain the weighted feature vector of the sub-basin under the classification of its tributary; sum the feature vectors of all sub-basins of the same tributary to form the comprehensive feature vector of the tributary.

[0071] S6. Based on the comprehensive feature vectors of the tributaries, calculate the similarity matrix between each tributary, set a similarity threshold, and classify tributaries with a similarity not lower than the threshold into the same hydrological similarity category. Divide the tributaries flowing into the reservoir into several hydrological functional similarity zones, wherein tributaries within the same zone have similar runoff generation attributes and confluence dynamics characteristics. Preferably, the similarity threshold can be set to 0.8. When the similarity between two tributaries is greater than or equal to the similarity threshold of 0.8, they are judged to be highly similar and classified into the same hydrological similarity zone.

[0072] S7. Within the same hydrologically similar area, select at least one tributary with a hydrological monitoring station as a reference watershed, and transfer the calibrated hydrological model parameters of the tributary to the tributary without data in the same area to complete the construction of the hydrological model parameters for the entire target watershed.

[0073] In step S6 above, the calculation of cosine similarity between the comprehensive feature vectors of the tributaries and the classification based on a threshold of 0.8 can be implemented by programming, such as using tools like MATLAB. The core algorithm logic is as described in this embodiment.

[0074] Example 2

[0075] Based on Example 1, the following uses the Three Gorges Reservoir area as an example to further elaborate on the method of the present invention.

[0076] The Three Gorges Reservoir area is a typical ecologically fragile and environmentally sensitive area in my country, with severe water and soil pollution in Chongqing. The main sources of pollution are agricultural non-point source pollution caused by soil erosion, excessive use of chemical fertilizers and pesticides, and the improper disposal of livestock manure from large-scale livestock farming. Its unique ecological environment promotes non-point source pollution. Surveys have found that agricultural non-point source pollutants account for a large proportion of water pollution sources in the Three Gorges Reservoir area. Statistics show that non-point source pollutants account for 70.8%, 60.6%, and 74.9% of the total pollution load in the Three Gorges Reservoir area, respectively. With the completion and operation of the Three Gorges Project, soil erosion, algal blooms in tributaries, and especially eutrophication have become increasingly prominent problems in the reservoir area. Natural factors such as topography, resources, and climate, as well as unreasonable human activities, have exacerbated agricultural environmental pollution in the reservoir area. The extensive use of pesticides, plastic film, chemical fertilizers, and livestock manure in the Three Gorges Reservoir area has caused significant damage to agricultural development and the ecological environment.

[0077] However, to date, there is still no dynamic assessment model for non-point source pollution in the Three Gorges Reservoir area, which poses a certain threat to the water quality safety management of the reservoir area. Therefore, it is necessary to construct an assessment and management model for pollution entering the Three Gorges Reservoir area, dynamically simulate pollution entering the reservoir area, and propose corresponding pollution reduction measures for key source areas to reduce water quality risks in the Three Gorges Reservoir area.

[0078] Since non-point source pollution input is closely related to water and sediment transport in the basin, the accuracy of hydrological and sediment simulation is fundamental to the accuracy of land-based non-point source pollution simulation. The Three Gorges Reservoir area exhibits diverse geographical spatial attributes and numerous tributaries, but only a few tributaries have hydrological stations for model calibration of water and sediment. Therefore, the key to the accuracy of hydrological simulation in the reservoir area lies in how to reasonably extend the parameters calibrated from a few tributaries to all tributaries.

[0079] The pollution assessment and management model for the Three Gorges Reservoir covers the watershed area from Cuntan and Wulong to the dam, with a total area of ​​60,000 square kilometers. The area includes 30 major first-level tributaries flowing into the reservoir, such as Shennongxi, Daninghe, Sanxihe, Meixihe, Liangdouhe, and Guanduhe. Currently, only 9 tributaries within the area have hydrological stations.

[0080] 1. Extraction of key factors in sub-basins of the Three Gorges Reservoir area

[0081] Based on the SWAT model, the entire study basin is divided into a series of independent sub-basins. The runoff generation process of each sub-basin is superimposed to obtain the hydrological response process of the entire basin.

[0082] (1) Data input and preprocessing

[0083] Input the DEM data of the reservoir area and surrounding areas. This data is the basis for SWAT to automatically generate river networks and divide sub-basins.

[0084] (2) Divide into sub-basins

[0085] The basin is divided into 980 sub-basins by setting a minimum upstream catchment area threshold to determine the boundaries and outlets of each sub-basin.

[0086] (3) Divide the hydrological response units

[0087] By overlaying land use type, slope grade, and soil data, the sub-basin is divided into several hydrological response units. The number of hydrological response units varies from sub-basin to sub-basin, ranging from approximately 1 to 20.

[0088] 2. Slope runoff generation was calculated using the SCS runoff curve method, and key factors were extracted.

[0089] The particle size distribution (average proportions of clay, sand, and silt), CN value, NDVI index, and area of ​​all hydrological response units within each of 980 sub-basins were extracted. CN is a core parameter of the SCS model, reflecting the runoff generation capacity of the underlying surface units. The CN value is closely related to underlying surface factors such as land use type, soil type, and antecedent soil moisture. Since land use type is difficult to represent numerically, CN value and vegetation cover were used directly as substitutes. The weighted average of soil particle size distribution, CN value, and NDVI index of the sub-basin was calculated based on the area of ​​the hydrological response units, serving as the key runoff generation factor for the sub-basin. Fuzzy clustering analysis was used to cluster these factors. Before clustering analysis, data should be normalized to avoid clustering bias caused by differences in the order of magnitude of the parameters.

[0090] 3. Sub-basin cluster analysis

[0091] Using standardized key factors as attribute features, a fuzzy C-means clustering algorithm was employed to perform spatial clustering analysis on all sub-basins of the reservoir area. The effectiveness of the Xie-Beni clustering function was then calculated. Parameter values, drawing The evaluation curve of parameter variation with the number of clusters is used to determine the optimal number of clusters based on the inflection point of the curve or the number of clusters that shows the best clustering effect, thus completing the sub-basin classification. Figure 2 As shown. With the increase of cluster groups, The parameters are getting smaller, but too many clustering groups will affect the next step of the analysis, especially after 10 groups. The parameter values ​​tend to level off. Therefore, 10 was chosen as the number of clusters to perform cluster analysis on 980 sub-basins in the Three Gorges Reservoir area. The clusters are as follows: Cluster 1 contains 95 sub-basins, Cluster 2 contains 84 sub-basins, Cluster 3 contains 76 sub-basins, Cluster 4 contains 104 sub-basins, Cluster 5 contains 90 sub-basins, Cluster 6 contains 176 sub-basins, Cluster 7 contains 95 sub-basins, Cluster 8 contains 122 sub-basins, Cluster 9 contains 54 sub-basins, and Cluster 10 contains 84 sub-basins.

[0092] For the above classification, the silhouette function is used to calculate the clustering quality, which is used to judge the quality of the clustering results. The optimal number of clusters is determined by the value calculated by the silhouette function. The distances to other points in the same cluster and the distances to other clusters are calculated to obtain the silhouette coefficient. When the silhouette coefficient is close to 1, the clustering result is better; when the silhouette coefficient is closer to -1, the clustering result is worse. The clustering quality calculated by the silhouette function is as follows: Figure 3 As shown in the figure. According to statistics, among the 980 sub-basins in the Three Gorges Reservoir area, 731 have a silhouette value greater than 0.5, accounting for 74.5% of the total, indicating a good overall clustering effect.

[0093] 4. Calculate the flow characteristic coefficients of sub-basins

[0094] Hydrological similarity considers not only runoff generation similarity but also confluence similarity. The amount and time of runoff reaching the tributary outlet from different sub-basins determine the similarity of confluence. The amount of water flowing into the river channel is directly related to the sub-basin area, slope, and length. These factors are extracted as key components. Considering the confluence characteristics of the reservoir area affected by backwater, the runoff from each sub-basin is considered as a water sphere about to enter the river channel. The river channel is generalized as a slope, ignoring the influence of differences in river confluence roughness. The confluence characteristic coefficients of each sub-basin are calculated to quantify the transport capacity of the sub-basin runoff to the tributary outlet. The calculation formula is as follows:

[0095]

[0096] in, Indicates the first The catchment area of ​​each sub-basin For the first Length of the river channel from the sub-basin to the tributary outlet Indicates the first Average riverbed slope in each sub-basin.

[0097] 5. Form the comprehensive characteristic vector of this tributary.

[0098] The sub-basin clustering results are represented by a fuzzy membership matrix, with a value of 1 assigned to belonging to the category and a value of 0 assigned to not belonging to the category. The confluence coefficient of each sub-basin is multiplied by the fuzzy membership matrix to serve as the weight of that sub-basin under a certain category. The weighted values ​​of each category for the 30 tributaries are summed to form the sum feature vector of the 30 tributaries.

[0099] 6. Calculate tributary similarity and partition the data.

[0100] Based on the comprehensive feature vectors of the 30 tributaries, a similarity matrix is ​​calculated between each tributary. This similarity matrix is ​​obtained by calculating the cosine similarity between the comprehensive feature vectors of each tributary, using the following formula:

[0101]

[0102] in These are the combined feature vectors of branches A and B, respectively.

[0103] Comparing the similarity among 30 vectors, those with high similarity were grouped into one category. The similarity analysis results among the branches are as follows: Figure 4 As shown in the figure. Similarity scores higher than 0.8 are grouped into one category and assigned a value of 1, while the rest are assigned 0. The results are as follows. Figure 5 As shown. According to Figure 5 According to the classification results, the 30 first-level tributaries flowing into the Three Gorges Reservoir area are divided into four categories, such as... Figure 6 As shown.

[0104] In summary, this invention, based on the DEM, land use types, and soil types of the Three Gorges Reservoir area, and according to the runoff generation and runoff calculation mechanism of the SWAT model, divides the area into sub-basins and hydrological response units, extracting key factors determining the hydrological similarity of the Three Gorges Reservoir area. These key factors include sub-basin area, average slope, soil particle size composition, and CN value. Using 980 sub-basins as calculation units, fuzzy C-means clustering and determining the optimal number of clusters were employed to plot... The curve showing the change with the number of clusters indicates that the curve tends to flatten out after 10 clusters. Therefore, 10 clusters for the sub-basin was determined to be the optimal number of clusters. The clustering quality was evaluated, and the results showed that 74.5% of the sub-basins had a profile coefficient >0.5, indicating good performance. The runoff generation of the sub-basin was generalized as a water sphere, and the river channel as a slope. The roughness difference was ignored, and the confluence coefficient was calculated. The sub-basin clustering results were represented by a fuzzy membership matrix. The confluence coefficient of each sub-basin was multiplied by the fuzzy membership matrix and used as a weight. The weighted values ​​of all sub-basins of the same tributary were summed to form the vector of that tributary. The similarity between the 30 tributary vectors was calculated using cosine similarity and other methods. A similarity threshold of 0.8 was set, and tributaries with a similarity higher than 0.8 were grouped into one category. Finally, the 30 tributaries were divided into four categories: A, B, C, and F, forming the hydrological similarity zoning scheme for the Three Gorges Reservoir area.

[0105] The four zoning categories are as follows: Category A tributaries are characterized by calcareous leached soil as the main soil type, slopes greater than 50 degrees, and a mix of forest and farmland; Category B tributaries are characterized by gleyed leached soil as the main soil type, slopes greater than 50 degrees, and a mix of forest and farmland; Category C tributaries are characterized by saturated gleyed soil as the main soil type, slopes greater than 25 degrees and less than 50 degrees, and a mix of farmland and some forest and farmland; and Category F tributaries are characterized by saturated gleyed soil as the main soil type, slopes less than 25 degrees, and a mix of farmland and farmland. Under these zoning results, within the same hydrologically similar zone, at least one tributary with a hydrological monitoring station is selected as a reference watershed. The calibrated hydrological model parameters of this tributary are then transferred to tributaries in the same zone without data, completing the construction of hydrological model parameters for the entire target watershed. This can be used to guide the construction of hydrological models for lake and reservoir-type watersheds, parameter calibration and verification, and simulation of inflow hydrological processes. It can also provide a foundation for subsequent watershed non-point source pollution simulation.

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features, and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for clustering and partitioning tributaries based on hydrological similarity in a reservoir area, characterized in that, The steps are as follows: S1. Based on the DEM data, river network vector data and watershed outlet point data of the reservoir area, the reservoir area watershed is divided into multiple sub-watersheds using the SWAT model. Furthermore, each sub-watershed is divided into one or more hydrological response units by overlaying land use type map, slope grade map and soil type map data. S2. The SCS runoff curve method is used to calculate slope runoff and determine the key factors reflecting the runoff capacity of the sub-basin. Taking the sub-basin as the unit, the key factors of all hydrological response units in the sub-basin are extracted. The key factors include soil particle size composition, CN value, NDVI index and hydrological response unit area. The weighted average of soil particle size composition, CN value and NDVI index of the sub-basin is calculated based on the area of ​​hydrological response units and used as the key runoff factor of the sub-basin. The factor is then standardized. S3. Using standardized key factors as attribute features, spatial clustering analysis was performed on all sub-basins of the reservoir area using the fuzzy C-means clustering algorithm. The effectiveness of the Xie-Beni clustering function was then calculated. Parameter values, drawing The evaluation curve of parameter change with the number of clusters is used to determine the optimal number of clusters based on the inflection point of the curve or the number of clusters that are flat and have the best clustering effect, and the sub-basin classification is completed. S4. Considering the runoff characteristics of the reservoir area affected by backwater, the runoff of the sub-basin is regarded as a water ball about to enter the river channel, and the river channel is generalized as a slope. Ignoring the influence of the difference in river runoff roughness, the runoff characteristic coefficient of each sub-basin is calculated. Taking the sub-basin as the unit, the component factors of the runoff characteristic coefficient are extracted. The component factors include the sub-basin area, average slope, and slope length from the sub-basin to the tributary outlet, which are used to quantify the transport capacity of the runoff of the sub-basin to the tributary outlet. S5. Multiply the fuzzy membership matrix of each sub-basin by the corresponding confluence feature coefficient to obtain the weighted feature vector of the sub-basin under the classification of its tributary; sum the feature vectors of all sub-basins of the same tributary to form the comprehensive feature vector of the tributary. S6. Based on the comprehensive feature vector of the tributaries, calculate the similarity matrix between each tributary, set a similarity threshold, classify tributaries with similarity not lower than the threshold into the same hydrological similarity category, and divide the tributaries flowing into the reservoir into several hydrological functional similarity zones, wherein the tributaries in the same zone have similar runoff generation attributes and confluence dynamics characteristics.

2. The method for clustering and partitioning tributaries based on hydrological similarity in a reservoir area according to claim 1, characterized in that, The confluence characteristic coefficients mentioned in step S5 The calculation formula is as follows: in, Indicates the first The catchment area of ​​each sub-basin For the first Length of the river channel from the sub-basin to the tributary outlet Indicates the first Average riverbed slope in each sub-basin.

3. The method for clustering and partitioning tributaries based on hydrological similarity in a reservoir area according to claim 1, characterized in that, The specific method for determining the optimal number of clusters in step S3 is as follows: calculate the number of clusters from 2 to 20 sequentially. Parameter values, drawing The curve of parameters changing with the number of clusters; when When the downward trend of the value slows down significantly and the curve shows a stable segment, the number of clusters corresponding to the starting point of the stable segment is selected as the optimal number of clusters.

4. The method for clustering and partitioning tributaries based on hydrological similarity in a reservoir area according to claim 1, characterized in that, The similarity threshold mentioned in step S6 can be set to 0.

8.

5. The method for clustering and partitioning tributaries of reservoir areas based on hydrological similarity according to claim 4, characterized in that, When the similarity between two tributaries is greater than or equal to the similarity threshold of 0.8, they are judged to be highly similar and classified into the same type of hydrological similarity zone.

6. The method for clustering and partitioning tributaries based on hydrological similarity in a reservoir area according to claim 1, characterized in that, In step S6, the similarity of hydrological functions is measured by calculating the similarity between the comprehensive feature vectors of each tributary.

7. The method for clustering and partitioning tributaries based on hydrological similarity in a reservoir area according to claim 1, characterized in that, After determining the optimal number of clusters and completing the sub-basin classification in step S3, the process also includes a step to verify the clustering results: The silhouette coefficient function is used to calculate the clustering quality and determine the quality of the clustering results; for each data point, its distance to other points in the same cluster and its distance to other clusters are calculated, and the basic formula for its silhouette coefficient is: in, Represents sample points The average distance to all other points within its cluster. Represents sample points The minimum average distance to all sample points in any other cluster is the silhouette coefficient. The closer the silhouette coefficient is to 1, the better the clustering result; the closer it is to -1, the worse the clustering result.

8. The method for clustering and partitioning tributaries based on hydrological similarity in a reservoir area according to claim 1, characterized in that, The CN value mentioned in S2 is a dimensionless empirical parameter used to comprehensively reflect the watershed's underlying surface's "interception-runoff generation" capacity in the rainfall-runoff relationship. It is an integer between 0 and 100 that comprehensively reflects three major factors: soil type, land use / vegetation, and early moisture condition (AMC).

9. The method for clustering and partitioning tributaries based on hydrological similarity in a reservoir area according to claim 1, characterized in that, The method further includes S7. Within the same hydrologically similar area, at least one tributary with a hydrological monitoring station is selected as a reference watershed, and its calibrated hydrological model parameters are transferred to the tributary in the area without data, thereby completing the construction of hydrological model parameters for the entire target watershed.

10. The method for clustering and partitioning tributaries of reservoir areas based on hydrological similarity according to claim 1, characterized in that, The several hydrologically similar zones obtained by the method can also be used for the construction of hydrological models for lake and reservoir type watersheds, parameter calibration and verification, and simulation of inflow hydrological processes. They can also provide a basis for subsequent watershed non-point source pollution simulation.

Citation Information

Patent Citations

  • Basin similarity classification method and device

    CN113887635A

  • Multi-site watershed hydrological model parameter integration calibration method based on reinforcement learning

    CN118114567A