Water and rain condition factor optimization method for dynamic identification of flood season after watershed

By using methods such as systematic clustering of hydrological and rainfall factors, information entropy, and seasonality indices, key factors for the post-flood season of the watershed are identified and optimized. This solves the problems of subjectivity in factor selection and overfitting in traditional methods, and achieves efficient and accurate hydrological analysis.

CN121935873APending Publication Date: 2026-04-28BUREAU OF HYDROLOGY CHANGJIANG WATER RESOURCES COMMISSION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BUREAU OF HYDROLOGY CHANGJIANG WATER RESOURCES COMMISSION
Filing Date
2026-02-10
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively identify and optimize water and rainfall factors in the post-flood season of a watershed. Traditional methods struggle to capture the nonlinear interactions between factors, and machine learning models lack physical interpretability, resulting in a large workload and a high risk of overfitting.

Method used

We employ systematic clustering, information content analysis, seasonality discrimination, and correlation analysis of water and rainfall factors. We combine information entropy, seasonality index, and multiple correlation coefficient to perform singular value decomposition, calculate factor importance index, construct clusters through hierarchical clustering and Euclidean distance, and select the most important water and rainfall factors.

Benefits of technology

By comprehensively considering the information content, seasonality, and correlation of hydrological factors, the subjectivity of factor selection is effectively avoided, thus improving the efficiency and accuracy of hydrological analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121935873A_ABST
    Figure CN121935873A_ABST
Patent Text Reader

Abstract

The invention provides a watershed post-flood season dynamic identification-oriented water and rain condition factor optimization method. The watershed post-flood season dynamic identification-oriented water and rain condition factor optimization method comprises the steps of watershed and rain condition factor system clustering, watershed and rain condition factor information quantity analysis, watershed and rain condition factor seasonal judgment, watershed and rain condition factor correlation analysis and watershed and rain condition factor importance identification and optimization. Various characteristics such as information amount, seasonality and relevance of the water regimen factors can be comprehensively considered; and the subjectivity problem of water regimen factor selection can be effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hydrological calculation technology, and in particular to a method for optimizing hydrological and rainfall factors for dynamic identification in the post-flood season of a watershed. Background Technology

[0002] In hydrological analysis and water resources management, hydrological and rainfall factors (such as rainfall intensity, flow rate, and water level) are the core inputs driving various hydrological models. Therefore, the identification and optimization of the importance of hydrological and rainfall factors are of great significance for hydrological analysis and calculation. When a watershed involves a large number of hydrological and rainfall factors, there are usually certain correlations and associations among these factors. Analyzing all hydrological and rainfall factors one by one would be extremely labor-intensive and somewhat blind. Currently, there is no method for identifying and optimizing the importance of hydrological and rainfall factors. Traditional factor screening methods (such as principal component analysis and correlation coefficient methods) struggle to capture the nonlinear interactions between factors; machine learning models (such as random forests and XGBoost) can assess feature importance, but they lack explanation of the physical mechanisms and are prone to overfitting in high-dimensional dynamic systems. Summary of the Invention

[0003] The purpose of this invention is to address the shortcomings of the prior art by providing a method for optimizing water and rainfall factors for dynamic identification during the post-flood season of a watershed.

[0004] To achieve the above objectives, the present invention adopts the following technical solution: This invention provides a method for optimizing hydrological and rainfall factors for dynamic identification during the post-flood season of a watershed, comprising the following steps: S1, systematic clustering of water and rainfall factors; S2, Analysis of information content of water and rainfall factors; S3. Seasonal identification of water and rainfall factors; S4. Correlation analysis of water and rainfall factors; S5. Identification and optimization of the importance of water and rainfall factors.

[0005] Furthermore, S1 specifically refers to: S101, Let the basin be... The sequence of water condition factors is as follows: The time is The values ​​for each element are: ; The hydrological factor sequence is normalized to a range of 0 to 1, and the calculation formula is as follows: ; in, and These are the minimum and maximum values ​​of the hydrological factor sequence, respectively; These are the normalized eigenvalues; S102. Calculate the Euclidean distance between hydrological factor sequences for two different hydrological factor sequences. , European distance The calculation formula is: ; S103. Perform hierarchical clustering on each hydrological factor: Hierarchical clustering constructs clusters through recursive merging or splitting operations. In agglomerated hierarchical clustering, each data point is regarded as a separate cluster, and then the clusters with the closest Euclidean distance are merged step by step; the greater the distance, the greater the difference between the two clusters.

[0006] Furthermore, using aggregation strategies and linking methods, a tree diagram representing the merging order among the clusters can be constructed, specifically as follows: S1031. Initialization: In the initial state, each data point is an independent cluster; S1032. Find the nearest cluster pair: Among all cluster pairs, find the cluster pair with the smallest merge distance; S1033, Merge Clusters: Merge the nearest cluster pairs found into a new cluster; S1034. Update the distance matrix: In the new distance matrix, replace the original cluster pairs with the new clusters and recalculate. Distance between the cluster and all other clusters; S1035, Repeat: Repeat S1031 to S1034 until all hydrological factors are merged into a single cluster.

[0007] Furthermore, S2 specifically involves calculating the information entropy of the hydrological factor sequence as follows: ; in, For the first The first water condition factor The probability of each value occurring; The information entropy of a hydrological factor is denoted as . The greater the information entropy, the greater the amount of information contained in the hydrological factor.

[0008] Furthermore, S3 specifically involves: constructing the seasonal intensity of the seasonal index standard hydrological factors; and constructing a non-seasonal flood model for the hydrological factors using a circular uniform distribution: ; in, Let be the circular uniform probability density function of the hydrological factors; Let be the circular uniform probability distribution function of the hydrological factors; For a given hydrological indicator, the composite vector method treats the corresponding daily-scale component of the flood season as the indicator vector; where the indicator vector for a particular day is the indicator value corresponding to that flood season day, and the angle is determined according to the rank of the corresponding flood season day in the total number of days. Hydrological factors The seasonal index is as follows: ; ; ; in, and The composite vector of the flood season index vectors are respectively in and Components on the axis; Represents the sequence of hydrological factors The first in Each possible value Indicators The seasonality index is the highest of the seasonality indices; the higher the seasonality index, the stronger the seasonality of the indicator.

[0009] Furthermore, S4 specifically includes: S401. Determine the correlation matrix of water and rainfall factors: Based on the variance of each indicator , and the covariance between various indicators and form related arrays : ; S402. Calculate the multiple correlation coefficient for each indicator: Related array for: ; Water conditions multiple correlation coefficient The calculation formula is: .

[0010] Furthermore, the characteristic of step S5 is that it specifically comprises: S501. The information entropy, seasonality index, and multiple correlation coefficient of each hydrological factor are used to construct a set matrix of hydrological factor indicators. Singular value decomposition is then performed to calculate the influence weights of information entropy, seasonality index, and multiple correlation coefficient. ; in, The eigenvalues ​​of information entropy, seasonality index, and multiple correlation coefficient; S502. Calculate the importance index of each hydrological factor. Importance index of water-related factors The calculation is as follows: ; The most important hydrological factors in each category are obtained by ranking them according to their importance index.

[0011] The beneficial effects of this invention are: it can comprehensively consider the information content, seasonality, and correlation of hydrological factors; and it can effectively avoid the subjective problem of hydrological factor selection. Attached Figure Description

[0012] Figure 1 This is a flowchart of a method for optimizing water and rainfall factors for dynamic identification in the post-flood season of a watershed. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0014] Please see Figure 1 A method for optimizing hydrological and rainfall factors for dynamic identification during the post-flood season of a watershed includes the following steps: S1, systematic clustering of water and rainfall factors; S2, Analysis of information content of water and rainfall factors; S3. Seasonal identification of water and rainfall factors; S4. Correlation analysis of water and rainfall factors; S5. Identification and optimization of the importance of water and rainfall factors.

[0015] Specifically, S1 is: S101, Let the basin be... The sequence of water condition factors is as follows: The time is The values ​​for each element are: ; The hydrological factors selected in this embodiment are the daily average water levels of Lianhuatang Station, Shashi Station, Hankou Station, Jiujiang Station, Chenglingji Station, and Hukou Station, as well as the daily average flow rates of Chenglingji Station, Hankou Station, Yichang Station, Luoshan Station, the four rivers of Dongting Lake, Hukou Station, the five rivers of Poyang Lake, Datong Station, and Huangzhuang Station. The hydrological factor sequence is normalized to a range of 0 to 1, and the calculation formula is as follows: ; in, and These are the minimum and maximum values ​​of the hydrological factor sequence, respectively; These are the normalized eigenvalues; S102. Calculate the Euclidean distance between hydrological factor sequences for two different hydrological factor sequences. , European distance The calculation formula is: ; S103. Perform hierarchical clustering on each hydrological factor: Hierarchical clustering constructs clusters through recursive merging or splitting operations. In agglomerated hierarchical clustering, each data point is regarded as a separate cluster, and then the clusters with the closest Euclidean distance are merged step by step; the greater the distance, the greater the difference between the two clusters.

[0016] Using aggregation strategies and linking methods, a tree diagram representing the merging order among the clusters can be constructed, specifically: S1031. Initialization: In the initial state, each data point is an independent cluster; S1032. Find the nearest cluster pair: Among all cluster pairs, find the cluster pair with the smallest merge distance; S1033, Merge Clusters: Merge the nearest cluster pairs found into a new cluster; S1034. Update the distance matrix: In the new distance matrix, replace the original cluster pairs with the new clusters and recalculate. Distance between the cluster and all other clusters; S1035, Repeat: Repeat S1031 to S1034 until all hydrological factors are merged into a single cluster.

[0017] In this embodiment, the hydrological factors are divided into three categories, as shown in Table 1: Table 1 Cluster Analysis of Hydrological Factors in the Middle Reaches of the Yangtze River

[0018] Specifically, S2 involves calculating the information entropy of the hydrological factor sequence as follows: ; in, For the first The first water condition factor The probability of each value occurring; The information entropy of a hydrological factor is denoted as . The greater the information entropy, the greater the amount of information contained in the hydrological factor.

[0019] The information entropy values ​​of each hydrological factor are shown in Table 2.

[0020] Table 2 Entropy values ​​of hydrological factors

[0021] Specifically, S3 involves: constructing the seasonal intensity of the seasonal index standard hydrological factors; and constructing a non-seasonal flood model for the hydrological factors using a circular uniform distribution. ; in, Let be the circular uniform probability density function of the hydrological factors; Let be the circular uniform probability distribution function of the hydrological factors; For a given hydrological indicator, the composite vector method treats the daily-scale component of the flood season indicator as the indicator vector; where the indicator vector for a given day is the indicator value corresponding to that flood season day, and the angle is determined according to the rank of the corresponding flood season day in the total number of days. Hydrological factors... The seasonal index is as follows: ; ; ; in, and The composite vector of the flood season index vectors are respectively in and Components on the axis; Represents the sequence of hydrological factors The first in Each possible value Indicators The seasonality index is the highest of the seasonality indices; the higher the seasonality index, the stronger the seasonality of the indicator.

[0022] The seasonal indices of various hydrological factors are shown in Table 3.

[0023] Table 3 Seasonal Indices of Hydrological Factors

[0024] Specifically, S4 is: S401. Determine the correlation matrix of water and rainfall factors: Based on the variance of each indicator , and the covariance between various indicators and form related arrays : ; S402. Calculate the multiple correlation coefficient for each indicator: Related array for:: ; Water conditions multiple correlation coefficient The calculation formula is: .

[0025] The seasonal indices of various hydrological factors are shown in Table 3.

[0026] Table 3. Multiple correlation coefficients of hydrological factors

[0027] Specifically, S5 is: S501. The information entropy, seasonality index, and multiple correlation coefficient of each hydrological factor are used to construct a set matrix of hydrological factor indicators. Singular value decomposition is then performed to calculate the influence weights of information entropy, seasonality index, and multiple correlation coefficient. ; in, The eigenvalues ​​of information entropy, seasonality index, and multiple correlation coefficient; S502. Calculate the importance index of each hydrological factor. Importance index of water-related factors The calculation is as follows: ; The most important hydrological factors in each category are obtained by ranking them according to their importance index.

[0028] According to the importance index analysis, the top three water situation factors are the daily average flow of the four rivers of Dongting Lake, the daily average flow of the five rivers of Poyang Lake, and the daily average water level at Shashi Station, which are largely consistent with the actual flood control situation in this basin.

[0029] The embodiments described above are merely illustrative of implementation methods of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be defined by the appended claims.

Claims

1. A method for optimizing water and rainfall factors for dynamic identification during the post-flood season of a watershed, characterized in that, Includes the following steps: S1, systematic clustering of water and rainfall factors; S2, Analysis of information content of water and rainfall factors; S3. Seasonal identification of water and rainfall factors; S4. Correlation analysis of water and rainfall factors; S5. Identification and optimization of the importance of water and rainfall factors.

2. The method for optimizing water and rainfall factors for dynamic identification during the post-flood season of a watershed, as described in claim 1, is characterized in that... Specifically, S1 is: S101, Let the basin be... The sequence of water condition factors is as follows: The time is The values ​​for each element are: ; The hydrological factor sequence is normalized to a range of 0 to 1, and the calculation formula is as follows: ; in, and These are the minimum and maximum values ​​of the hydrological factor sequence, respectively; These are the normalized eigenvalues; S102. Calculate the Euclidean distance between hydrological factor sequences for two different hydrological factor sequences. , European distance The calculation formula is: ; S103. Perform hierarchical clustering on each hydrological factor: Hierarchical clustering constructs clusters through recursive merging or splitting operations. In agglomerated hierarchical clustering, each data point is regarded as a separate cluster, and then the clusters with the closest Euclidean distance are merged step by step; the greater the distance, the greater the difference between the two clusters.

3. The method for optimizing water and rainfall factors for dynamic identification during the post-flood season of a watershed, as described in claim 2, is characterized in that: Using aggregation strategies and linking methods, a tree diagram representing the merging order among the clusters can be constructed, specifically: S1031. Initialization: In the initial state, each data point is an independent cluster; S1032. Find the nearest cluster pair: Among all cluster pairs, find the cluster pair with the smallest merge distance; S1033, Merge Clusters: Merge the nearest cluster pairs found into a new cluster; S1034. Update the distance matrix: In the new distance matrix, replace the original cluster pairs with the new clusters and recalculate. Distance between the cluster and all other clusters; S1035, Repeat: Repeat S1031 to S1034 until all hydrological factors are merged into a single cluster.

4. The method for optimizing water and rainfall factors for dynamic identification during the post-flood season of a watershed, as described in claim 2, is characterized in that... Specifically, S2 involves calculating the information entropy of the hydrological factor sequence as follows: ; in, For the first The first water condition factor The probability of each value occurring; The information entropy of a hydrological factor is denoted as . The greater the information entropy, the greater the amount of information contained in the hydrological factor.

5. The method for optimizing water and rainfall factors for dynamic identification during the post-flood season of a watershed, as described in claim 4, is characterized in that... Specifically, S3 involves: constructing the seasonal intensity of the seasonal index standard hydrological factors; and constructing a non-seasonal flood model for the hydrological factors using a circular uniform distribution. ; in, Let be the circular uniform probability density function of the hydrological factors; Let be the circular uniform probability distribution function of the hydrological factors; For a given hydrological indicator, the composite vector method treats the daily-scale component of the flood season indicator as the indicator vector; where the indicator vector for a given day is the indicator value corresponding to that flood season day, and the angle is determined according to the rank of the corresponding flood season day in the total number of days. Hydrological factors... The seasonal index is as follows: ; ; ; in, and The composite vector of the flood season index vectors are respectively in and Components on the axis; Represents the sequence of hydrological factors The first in Each possible value Indicators The seasonality index is the highest of the seasonality indices; the higher the seasonality index, the stronger the seasonality of the indicator.

6. The method for optimizing water and rainfall factors for dynamic identification during the post-flood season of a watershed, as described in claim 5, is characterized in that... Specifically, S4 is: S401. Determine the correlation matrix of water and rainfall factors: Based on the variance of each indicator , and the covariance between various indicators and form related arrays : ; S402. Calculate the multiple correlation coefficient for each indicator: Related array for:: ; Water conditions multiple correlation coefficient The calculation formula is: 。 7. The method for optimizing water and rainfall factors for dynamic identification during the post-flood season of a watershed, as described in claim 6, is characterized in that... Specifically, S5 is: S501. The information entropy, seasonality index, and multiple correlation coefficient of each hydrological factor are used to construct a set matrix of hydrological factor indicators. Singular value decomposition is then performed to calculate the influence weights of information entropy, seasonality index, and multiple correlation coefficient. ; in, The eigenvalues ​​of information entropy, seasonality index, and multiple correlation coefficient; S502. Calculate the importance index of each hydrological factor. Importance index of water-related factors The calculation is as follows: ; The most important hydrological factors in each category are obtained by ranking them according to their importance index.