A small watershed classification method based on louvain community detection algorithm
Patent Information
- Application Number
- CN202311487535.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-09
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2043-11-09
AI Technical Summary
[0002]小流域水文相似性研究是水文学科领域的基础课题和前沿问题,基于聚类算法的小流域分类方法和水文相似性评价是小流域水文相似性研究的主要方法,但传统的聚类算法无法体现每个类别中流域的相似性程度且没有量化每一类中各小流域的角色,而水文相似性评价能够体现各流域间的相似程度,但无法反映所选的相似流域间的类别同质性
1、根据流域具体情况,结合已有数据,确定进行相似度计算的相似性指标;将选取相似度的计算方法,将各小流域的相似性指标数据输入到相似性评价指标的公式中计算各流域间的相似度;将小流域作为节点,确定最佳的连接阈值,在相似度大于该阈值的两个小流域间添加连接;将相似度的值作为连接边的权重,构建出复杂网络;使用Louvain算法将所有小流域划分成不同社区;选取网络平均度、中心度、K核心三种网络指标进一步解析各类中小流域之间的水文联系;确定各流域的洪峰模数,对分类结果进行合理性验证。
Smart Images

Figure CN117851876B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of hydrological forecasting research technology for small watersheds, specifically a small watershed classification method based on the Louvain community detection algorithm. Background Technology
[0002] The study of hydrological similarity in small watersheds is a fundamental and cutting-edge topic in the field of hydrology. Small watershed classification methods and hydrological similarity evaluation based on clustering algorithms are the main methods for studying hydrological similarity in small watersheds. However, traditional clustering algorithms cannot reflect the degree of similarity between watersheds in each category and do not quantify the role of each small watershed in each category. Hydrological similarity evaluation can reflect the degree of similarity between watersheds, but it cannot reflect the homogeneity of categories among the selected similar watersheds. Summary of the Invention
[0003] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a small watershed classification method based on the Louvain community detection algorithm, which improves the accuracy of small watershed classification and further analyzes the hydrological relationships between various types of small watersheds.
[0004] (II) Technical Solution To achieve the above objectives, the present invention provides the following technical solution: a small watershed classification method based on the Louvain community detection algorithm, comprising determining similarity indicators: determining similarity indicators for similarity calculation based on the specific conditions of the watershed and in combination with existing data; Determine similarity: Select appropriate similarity evaluation indicators, input the values of the similarity indicators of each sub-basin into the formula of the similarity evaluation indicators, and calculate the similarity between each sub-basin.
[0005] Small watershed community partitioning: Treating small watersheds as nodes, determining the optimal connection threshold, and adding connections between two small watersheds with similarity greater than the threshold; using the similarity value as the weight of the connection edge to construct a complex network; using the Louvain algorithm to divide all small watersheds into different communities (i.e., classifying small watersheds).
[0006] Network indicator evaluation: Three network indicators, namely network mean degree, centrality, and K-core, were selected to further analyze the hydrological connections between small watersheds in each category.
[0007] Reasonableness verification: Determine the flood peak modulus for each watershed and verify the reasonableness of the classification results.
[0008] (III) Beneficial Effects Compared with the prior art, the present invention has the following beneficial effects: 1. Based on the specific conditions of the watershed and combined with existing data, determine the similarity index for similarity calculation; select the similarity calculation method, input the similarity index data of each sub-watershed into the formula of the similarity evaluation index to calculate the similarity between watersheds; treat the sub-watersheds as nodes, determine the optimal connection threshold, and add connections between two sub-watersheds with similarity greater than the threshold; use the similarity value as the weight of the connection edge to construct a complex network; use the Louvain algorithm to divide all sub-watersheds into different communities; select three network indicators—network mean degree, centrality, and K-core—to further analyze the hydrological connections between various types of small and medium-sized watersheds; determine the flood peak modulus of each watershed, and verify the rationality of the classification results.
[0009] 2. This method selects similarity indicators, calculates the similarity between watersheds through similarity evaluation indicators, sets the similarity as the weight of the connecting edges, uses the Louvain community detection algorithm for classification, and selects three network indicators, namely network mean degree, centrality, and K-core, to further analyze the hydrological connections between various small and medium-sized watersheds, which helps to identify the similarity of watershed runoff generation and confluence conditions and implement flood risk management strategies.
[0010] 3. This method uses the inference formula method to calculate the flood peak modulus of each watershed to further verify the rationality of the classification of small watershed results. The overall distribution of the flood peak modulus is similar to the classification results of small watersheds, indicating that this method can obtain relatively reliable classification results of small watersheds and better identify the similarity of watershed runoff conditions, providing a new approach to solving the flood forecasting problem in areas without data. Attached Figure Description
[0011] Figure 1 Flowchart of a small watershed classification method based on the Louvain community detection algorithm; Figure 2 Complex network structure diagram; Figure 3 Example graph of network centrality; Figure 4 A map showing the centrality results of a typical small watershed in a certain province; Figure 5 K-core result map of a typical small watershed in a certain province; Figure 6 Distribution of characteristic values of flood peak modulus in a typical small watershed of a certain province. Detailed Implementation
[0012] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0013] Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] The study of hydrological similarity in small watersheds is a fundamental and cutting-edge topic in the field of hydrology. Small watershed classification methods and hydrological similarity evaluation based on clustering algorithms are the main methods for studying hydrological similarity in small watersheds. However, traditional clustering algorithms cannot reflect the degree of similarity between watersheds in each category and do not quantify the role of each small watershed in each category. Hydrological similarity evaluation can reflect the degree of similarity between watersheds, but it cannot reflect the homogeneity of categories among the selected similar watersheds. Based on the specific conditions of the watershed and combined with existing data, similarity indicators for similarity calculation were determined. Appropriate similarity evaluation indicators were selected, and the values of the similarity indicators for each sub-watershed were input into the formula of the similarity evaluation indicators to calculate the similarity between sub-watersheds. Using sub-watersheds as nodes, an optimal connection threshold was determined, and connections were added between two sub-watersheds with similarity greater than this threshold. The similarity values were used as the weights of the connection edges. The Louvain algorithm was used to divide all sub-watersheds into different communities. Network indicators were determined, and three network indicators—network mean degree, centrality, and K-core—were selected to further analyze the hydrological connections between various types of small and medium-sized watersheds. The flood peak modulus of each watershed was determined, and the rationality of the community division results was verified.
[0015] This embodiment provides a small watershed classification method based on the Louvain community detection algorithm; Determine similarity indicators: Based on the specific conditions of the watershed and in conjunction with existing data, determine the similarity indicators for similarity calculation; Determine similarity: Select appropriate similarity evaluation indicators, input the values of the similarity indicators of each sub-basin into the formula of the similarity evaluation indicators, and calculate the similarity between each sub-basin.
[0016] Small watershed community partitioning: Treating small watersheds as nodes, determining the optimal connection threshold, and adding connections between two small watersheds with a similarity greater than the threshold; using the similarity value as the weight of the connection edge to construct a complex network; and using the Louvain algorithm to partition all small watersheds into different communities.
[0017] Network indicator evaluation: Three network indicators, namely network mean degree, centrality, and K-core, were selected to further analyze the hydrological connections between small watersheds in each category.
[0018] Reasonableness verification: Determine the flood peak modulus for each watershed and verify the reasonableness of the classification results.
[0019] As one or more embodiments, the specific steps of the small watershed community delineation include: (1) Louvain community detection algorithm The Louvain community detection algorithm is a network partitioning method, and broadly speaking, a graph clustering algorithm. Its main function is to divide graph data into different communities, where nodes within a community are closely connected or similar, while the connections between nodes in different communities are sparse or dissimilar. Community detection algorithms are now widely used in various fields. The Louvain algorithm is based on modularity, and its optimization goal is to maximize the modularity of the entire community network.
[0020] The formula for calculating modularity is: (Formula 1) In the formula: It refers to the community Internal weights It refers to the community The weights of edges connecting internal points, including edges within the community and edges outside the community. This represents the sum of the weights of all edges. The modularity ranges from -1 to 1, and a value between 0.3 and 0.7 is generally considered a suitable partitioning result.
[0021] The formula for calculating modularity gain is: (Formula 2) In the formula: K i,j This represents a node. (or community) The sum of the weights of the edges connecting to the community B to which the object is to be moved; This indicates the relationship with the node. (or community) The sum of the weights of the connecting edges; This indicates the relationship with the node. (or community) The community to be moved to B The sum of the weights of the connecting edges.
[0022] Based on the specific conditions of the watershed and the existing data, a similarity index for similarity calculation is determined. A suitable similarity evaluation index is selected, and the values of the similarity index for each sub-watershed are input into the formula of the similarity evaluation index to calculate the similarity between sub-watersheds. Using sub-watersheds as nodes, an optimal connection threshold is determined, and connections are added between two sub-watersheds with similarity greater than this threshold. The similarity values are used as the weights of the connection edges to construct a complex network. The Louvain algorithm is used to perform community partitioning on the constructed complex network, ultimately obtaining the community partitioning results for all sub-watersheds.
[0023] (2) Similarity evaluation indicators To determine a similarity metric to characterize the degree of similarity between watersheds, let a certain watershed feature be... Then the watershed and the basin In watershed characteristics The similarity on is: (Formula 3) In the formula, and Watershed characteristics Two values; For the basin and the basin In features The absolute distance on the surface.
[0024] The above formula mainly analyzes a single characteristic of a watershed. If multiple indicators are selected to measure similar watersheds, then the watershed... and the basin The formula for calculating the similarity between them is: (Formula 4) In the formula, For the first In this study, each watershed characteristic is assigned the same weight. ; and respectively watershed and the basin In the Features The value on; For the basin and the basin In features The similarity is calculated on [the surface]. The higher the similarity between them, the more similar the watersheds they reflect.
[0025] As one or more embodiments, the network metric evaluation specifically includes the following steps: (1) Average network degree The average degree of a network refers to the average number of connections per node. In an undirected subnetwork, the degree of each node is the number of edges connected to it. The average degree is the sum of the degrees of all nodes divided by the number of nodes. Therefore, the average degree of a community... The average degree is: (Formula 5) In the formula: Indicates community The sum of the degrees of all nodes, Indicates community The number of nodes.
[0026] (2) Centrality Centrality describes the role of a node in a community network. Centrality includes degree centrality and betweenness centrality. Degree centrality indicates that a given node in a community can represent the characteristics of all nodes in that community. Node betweenness refers to the number of shortest paths passing through a node in a network. The degree centrality of a node is calculated as follows: (Formula 6) In the formula: Represents a node Connect to the community The number of connections to other nodes in the middle, Indicates community The number of nodes.
[0027] The formula for calculating the betweenness centrality of a node is: (Formula 7) In the formula: Indicates community Middle node To the node The number of shortest paths, express The path through nodes The quantity.
[0028] (3) K core K cores are essentially a densely connected set of nodes. Determining K cores can be divided into the following three steps: First, remove all nodes with only one connection edge and put these nodes together, denoted as subset 1; Second, remove all nodes with only i connection edges, denoted as subset i; Third, continue until all nodes have been assigned, and all nodes in subset k (where k is the maximum value of i) are included; here k is also called core degree.
[0029] After obtaining the community division results of the small watersheds, three network evaluation indicators, namely network mean degree, centrality, and K-core, were selected to further analyze the hydrological connections between small watersheds in each community.
[0030] As one or more embodiments, the rationality verification specifically includes the following steps: The rationality of the watershed classification results was further verified by calculating the flood peak modulus of each small watershed using the inference formula method.
[0031] The formula for calculating the peak flood modulus is: (Formula 8) (Formula 9) (Formula 10) In the formula: The peak value modulus; For time period The maximum net rainfall within the area; The drainage area; Represents the duration of the confluence of watersheds; The longest confluence path distance; For along the process The average drop; For the bus parameters; Represents the flood peak modulus.
[0032] Watershed area, slope, land use / cover, soil type, and rainfall were selected as similarity indicators. DEM topographic data, drainage data, and rainfall data for relevant sub-watersheds were collected and processed to obtain similarity index values for each sub-watershed. Based on these similarity evaluation indicators, the similarity between sub-watersheds was calculated. Sub-watersheds were used as nodes; generally, two watersheds with a similarity greater than 0.6 are considered similar watersheds. Therefore, a connection threshold of 0.6 was set. A connection was added between two nodes only when the similarity was greater than the connection threshold. Using similarity values as the weights of the connecting edges, a complex network is constructed. The constructed network is then input into the Louvain community detection algorithm model for community division, ultimately resulting in 6 large communities and 9 small communities (the number of small watersheds in the large communities accounts for more than 90% of all small watersheds). Three network evaluation indicators—network mean degree, centrality, and K-core—are selected to further analyze the hydrological connections between small watersheds in the 6 large communities. The flood peak modulus of each small watershed is calculated using the inference formula method to further verify the rationality of the community division results.
[0033] Figure 5 The community division results for the small watersheds show that all small watersheds were divided into 15 communities, including 6 large communities and 9 small communities. The community detection algorithm is not just a simple division; it can also reflect the commonalities and uniqueness among small watersheds. Large communities reflect the commonalities among small watersheds, while small watersheds with unique characteristics are reflected in isolated small communities.
[0034] Figure 6 The range and spatial distribution of the flood peak modulus of each small watershed show that the flood peak modulus generally exhibits a distribution pattern of gradually decreasing from the two centers in the middle and east to the surrounding areas, which is similar to the results of the community division of the small watershed. The flood peak modulus of the small watershed within the same community shows a certain degree of clustering, and from community 1 to community 6, the flood peak modulus of the small watershed generally shows a downward trend.
[0035] Table 1 Table 1 shows the statistical results of the similarity index values for each small watershed. Based on these index values, the next step is to calculate the similarity of the small watersheds.
[0036] Table 2 Table 2 shows the calculation results of the similarity between small watersheds, which can more specifically represent the degree of similarity between similar small watersheds, and can be used as a connection threshold to construct complex networks for classification.
[0037] Table 3 Table 3 shows the basic characteristics of small watersheds in each large community after classification. It can be seen that the differences in the characteristics of small watersheds between different communities are particularly obvious, while the differences in the characteristics of small watersheds within the same community are small or basically the same.
[0038] Table 4 (Note: "High", "Relatively High", "Medium", "Low", and "Relatively Low" in the table represent the distribution of the indicator values of this community's watershed among all its smaller watersheds.) Table 4 shows the average network degree of each large community. A higher average degree indicates more connections between the smaller watersheds within the community and a higher homogeneity of runoff generation and confluence conditions. Conversely, a lower average degree indicates fewer connections between the smaller watersheds within the community and a lower homogeneity of runoff generation and confluence conditions.
[0039] Table 5 Table 5. Average degree results for six large community sub-networks.
[0040] In summary, the small watershed classification method based on the Louvain community detection algorithm essentially determines similarity indices based on existing data and specific watershed conditions; calculates the similarity between watersheds based on these similarity evaluation indices; determines connection thresholds to construct a complex network; uses the Louvain community detection algorithm for community division; employs three network evaluation indices to analyze the hydrological connections between watersheds within a community; quantifies the role of small watersheds within the community; and finally, verifies the rationality of the classification results using the flood peak modulus. This method belongs to the attribute similarity method in small watershed similarity research, but it differs from traditional clustering algorithms. It can obtain sets of similar watersheds, demonstrating the class identity (homogeneity) among similar watersheds, and also reflect the degree of hydrological similarity between small watersheds. In practical applications, this method can effectively identify similar small watersheds and meet the application needs of flood forecasting in data-scarce areas.
[0041] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A small watershed classification method based on the Louvain community detection algorithm, characterized in that: include: Determine similarity indicators: Based on the specific conditions of the watershed and in conjunction with existing data, determine the similarity indicators for similarity calculation; Determine similarity: Select appropriate similarity evaluation indicators, input the values of the similarity indicators of each sub-basin into the formula of the similarity evaluation indicators to calculate the similarity between each sub-basin; Small watershed community partitioning: Treating small watersheds as nodes, determining the optimal connection threshold, and adding connections between two small watersheds with a similarity greater than the threshold; using the similarity value as the weight of the connection edge to construct a complex network; The Louvain algorithm is used to divide all small watersheds into different communities, thus classifying the small watersheds. Network indicator evaluation: Three network indicators, namely network mean degree, centrality, and K-core, were selected to further analyze the hydrological connections between small watersheds in each category; Reasonableness verification: Determine the flood peak modulus for each watershed and verify the reasonableness of the classification results; The specific steps for dividing the small watershed communities include: (1) Louvain community detection algorithm The Louvain community detection algorithm is a network partitioning method, and broadly speaking, it is also a graph clustering algorithm. Its main function is to divide graph data into different communities. Nodes within a community are closely connected or similar, while the connections between nodes in different communities are sparse or dissimilar. Community detection algorithms are now widely used in various fields. The Louvain algorithm is based on modularity, and its optimization goal is to maximize the modularity of the entire community network. The formula for calculating modularity is: (Official 1) In the formula: It refers to the community Internal weights It refers to the community The weights of edges connecting internal points, including edges within the community and edges outside the community. This represents the sum of the weights of all edges. The modularity ranges from -1 to 1, and a value between 0.3 and 0.7 is generally considered a suitable partitioning result. The formula for calculating modularity gain is: (Official 2) In the formula: K i,j This represents a node. or community The sum of the weights of the edges connecting to the community B to which the device is to be moved; This indicates the relationship with the node. or community The sum of the weights of the connecting edges; This indicates the relationship with the node. or community The community to be moved to B The sum of the weights of the connecting edges; Based on the specific conditions of the watershed and the existing data, a similarity index for similarity calculation is determined. A suitable similarity evaluation index is selected, and the values of the similarity index for each sub-watershed are input into the formula of the similarity evaluation index to calculate the similarity between sub-watersheds. Using sub-watersheds as nodes, an optimal connection threshold is determined, and connections are added between two sub-watersheds with similarity greater than this threshold. The similarity values are used as the weights of the connection edges to construct a complex network. The Louvain algorithm is used to perform community partitioning on the constructed complex network, ultimately obtaining the community partitioning results for all sub-watersheds.
2. The small watershed classification method based on the Louvain community detection algorithm according to claim 1, characterized in that: The similarity evaluation index: Define a similarity metric to characterize the degree of similarity between watersheds. Let a certain watershed feature be... Then the watershed and the basin In watershed characteristics The similarity on is: (Official 3) In the formula, and Watershed characteristics Two values; For the basin and the basin In features The absolute distance on; The above formula mainly analyzes a single characteristic of a watershed. If multiple indicators are selected to measure similar watersheds, then the watershed... and the basin The formula for calculating the similarity between them is: (Official 4) In the formula, For the first In this study, each watershed characteristic is assigned the same weight. ; and respectively watershed and the basin In the Features The value on; For the basin and the basin In features The similarity on the surface, the calculated similarity in The higher the similarity between them, the more similar the watersheds they reflect.
3. The small watershed classification method based on the Louvain community detection algorithm according to claim 1, characterized in that: The average degree of the network: The average degree of a network refers to the average number of connections per node. In an undirected subnetwork, the degree of each node is the number of edges connected to it. The average degree is the sum of the degrees of all nodes divided by the number of nodes. Therefore, the community... The average degree is: (Official 5) In the formula: Indicates community The sum of the degrees of all nodes, Indicates community The number of nodes.
4. The small watershed classification method based on the Louvain community detection algorithm according to claim 1, characterized in that: Centrality: Centrality describes the role of a node in a community network. Centrality includes degree centrality and betweenness centrality. Degree centrality indicates that a given node in a community can represent the characteristics of all nodes in that community. Node betweenness refers to the number of shortest paths through a node in a network. The degree centrality of a node is calculated as follows: (Official 6) In the formula: Represents a node Connect to the community The number of connections to other nodes in the middle, Indicates community The number of nodes; The formula for calculating the betweenness centrality of a node is: (Official 7) In the formula: Indicates community Middle node To the node The number of shortest paths, express The path through nodes The quantity.
5. The small watershed classification method based on the Louvain community detection algorithm according to claim 1, characterized in that: The K core: K cores are essentially a densely connected set of nodes. Determining K cores can be divided into the following three steps: First, remove all nodes with only one connection edge and put these nodes together, denoted as subset 1; Second, remove all nodes with only i connection edges, denoted as subset i; Third, continue until all nodes have been assigned, and all nodes in subset k, where k is the maximum value of i; here k is also called core degree. After obtaining the community division results of the small watersheds, three network evaluation indicators, namely network mean degree, centrality, and K-core, were selected to further analyze the hydrological connections between small watersheds in each community.
6. The small watershed classification method based on the Louvain community detection algorithm according to claim 1, characterized in that: The rationality verification process includes the following specific steps: The rationality of the watershed classification results was further verified by calculating the flood peak modulus of each small watershed using a reasoning formula method. The formula for calculating the peak flood modulus is: (Official 8) (Official 9) (Official 10) In the formula: The peak value modulus; For time period The maximum net rainfall within the area; The drainage area; Represents the duration of the confluence of watersheds; The longest confluence path distance; For along the process The average drop; For the bus parameters; Represents the flood peak modulus.