Flood classification method and system based on dynamic time warping and multi-cluster coupling

By using dynamic time warping and multi-cluster coupling methods, the similarity calculation and classification of flood flow time series data are optimized, solving the problem of inaccurate similarity judgment in traditional methods and achieving higher flood classification accuracy and management reliability.

CN120493104BActive Publication Date: 2025-10-17ZHEJIANG INST OF HYDRAULICS & ESTUARY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510873901.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-17
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

Traditional flood classification methods fail to effectively utilize flood flow time series data for direct process similarity calculation and are sensitive to the selection of flood start and end times, leading to inaccurate similarity judgments, especially in the failure to properly distinguish contribution values ​​between high and low flow intervals.

Method used

The Dynamic Time Warping (DTW) algorithm is used to optimize the similarity calculation of flood flow time series data. Combined with hierarchical clustering and K-medoids multi-clustering methods, the accuracy and precision of flood classification are improved by adjusting the distance coefficient between high and low flow intervals and optimizing the clustering process.

Benefits of technology

It improves the accuracy of flood similarity calculation and the precision of flood classification, reduces the impact of start and end time selection, enhances the focus on high-flow intervals, and improves the reliability of flood management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493104B_ABST
    Figure CN120493104B_ABST
Patent Text Reader

Abstract

The application discloses a flood classification method and system based on dynamic time warping and multi-cluster coupling. The application firstly collects and processes flood flow time series data to obtain a data set of standardized flood composition; then adopts a dynamic time warping method to calculate the similarity between each two of K fields of flood, to form a similarity matrix; and finally carries out flood classification based on DTW coupling hierarchical clustering and K-medoids multi-clustering. The system of the application comprises a data acquisition module, a data preprocessing module, a comparison module and a classification output module. The application optimizes the hierarchical clustering again, combines the multi-clustering method, improves the rationality of flood classification, and helps to improve the accuracy of flood classification. The application can be applied to the fields of flood warning and flood control scheduling, and has important practical value and social significance.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of data analysis, and relates to a flood classification system coupled with a dynamic time warping (DTW) algorithm and a multi-clustering method, which is used to improve the calculation accuracy of flood process similarity and the classification reliability. BACKGROUND

[0002] Surface hydrological process research is the theoretical basis of flood forecasting, and has complex characteristics such as randomness, fuzziness, nonlinearity, non-stationarity (time variability). The similarity research of flood aims to explore the characteristics and laws contained in the flood time series data. The traditional flood classification method has the following problems: ① Most of them are based on the classification of characteristic values of flood flow time series data, and less use flood flow time series data directly for process similarity classification. ② In the calculation of flood flow time series data directly for process similarity, the different contribution values of the high flow time series part and the low flow time series part of the same flood to the similarity calculation are not considered. ③ The flood flow time series data is usually standardized one by one, which is greatly affected by the selection of flood start and end time, and the unreasonable similarity of high rising flood peak and low rising flood peak may occur after standardization.

[0003] The dynamic time warping (DTW) algorithm can effectively handle unequal length sequences and phase shift when calculating time series similarity, so that two floods with different lengths can directly calculate the similarity of flood process shape, and the similarity of any two floods can be measured. The traditional DTW uses uniform weight for flood peak area and non-flood peak area. However, in fact, people pay more attention to flood peak area, which is also an important influencing factor of flood warning and forecasting, and is the key point of flood management in the humid areas of southern China. Therefore, it is of great significance to increase the similarity calculation weight of the high flow time series part of the flood.

[0004] In the era of big data, machine learning and data mining are two closely related important technologies for learning knowledge and discovering laws. Clustering is a very important method in data mining research and application, and clustering analysis includes data representation, feature extraction, similarity measurement and clustering algorithm of time series. Clustering algorithms often have advantages and disadvantages, K-medoids is one of the most common clustering algorithms, but it has the disadvantages of uncertain number of classifications, sensitivity to initial center and easy to fall into local optimum in clustering calculation. Therefore, the cooperation of multiple clustering methods can effectively make up for the limitations of single algorithm. SUMMARY

[0005] An object of the present application is to provide a flood classification method based on dynamic time warping and multi-cluster coupling to improve the accuracy of flood similarity calculation and the precision of flood classification. The present application improves the contribution allocation strategy of DTW, combines the multi-algorithm coupling of hierarchical clustering and K-medoids, and solves the problems of inaccurate similarity measurement and unstable clustering results in flood classification.

[0006] The method of the present application is as follows:

[0007] Step (1) collects and processes flood flow time series data, and obtains a data set of K-field standardized flood composition after processing.

[0008] Step (2) calculates the time warping distance between any two K-field standardized floods using the dynamic time warping method, and obtains the similarity matrix of K-field floods;

[0009] (2-1) Select any two flood flow time series data, and calculate the high-low flow time series data according to the maximum flood flow value and the minimum flood flow value.

[0010] (2-2) Set the high and low flow time series data distance coefficient, calculate the distance between any two points on the two flood flow time series data, and the high flow time series data distance coefficient is less than the low flow time series data distance coefficient.

[0011] (2-3) Calculate the time warping distance of the two floods using the dynamic time warping distance algorithm.

[0012] (2-4) Calculate the time warping distance between any two K-field standardized flood flow time series data according to the time series, and finally obtain the similarity matrix of K-field floods.

[0013] Step (3) Similar flood classification, flood classification based on DTW coupling hierarchical clustering and K-medoids multi-clustering;

[0014] (3-1) Using the method of DTW coupling hierarchical clustering, calculate the flood clustering condensation process, remove the isolated points, determine the flood classification number R and the corresponding classification, and finally calculate the contour coefficient S1 of the classification.

[0015] (3-2) Select the classification number R determined by hierarchical classification as the classification number of K-medoids, and select any one of the two most similar groups of standardized flood flow time series data in each classification after hierarchical classification as the initial center of K-medoids clustering; the two most similar groups of standardized flood flow time series data in a classification are the two floods with the smallest dynamic time warping distance.

[0016] (3-3) using the K-medoids clustering method to iteratively calculate the flood classification until the termination condition is met, and the profile coefficient of the classification is recorded as S2;

[0017] (3-4) selecting the classification with a high profile coefficient score between the profile coefficients S1 and S2 as the final flood classification number R', and obtaining the corresponding classification.

[0018] Another object of the present application is to provide a flood classification system based on dynamic time warping and multi-clustering coupling for executing the aforementioned flood classification method. The flood classification system comprises a data acquisition module, a data preprocessing module, a comparison module, and a classification output module.

[0019] The data acquisition module is used to acquire flow time series data in a target basin and a target time range;

[0020] The data preprocessing module is used to clean, round, interpolate, and standardize the flood flow time series data;

[0021] The comparison module is used to obtain the pairwise similarity between multiple floods;

[0022] The classification output module is used to determine the flood type and output.

[0023] The present application has the following advantages: the present application optimizes the original single standardization of each flood to annual flood flow time series data standardization, which can compensate for the possible error influence of flood start and end time selection on flood standardization and flood similarity judgment, and can compensate for the inconsistency of mean and mean difference in single flood standardization process, which can lead to similarity calculation error, such as the occurrence of high water level rising high flow and low water level rising low flow in two floods after standardization.

[0024] Dynamic time warping gives the same weight to the distance of each step of matching, and in fact we believe that the similarity of flood in high flow part should contribute more than the similarity of low flow part. The present application optimizes the DTW algorithm to give different coefficient settings to high flow time series area and low flow time series area, and improves the accuracy of flood similarity calculation.

[0025] Hierarchical clustering can demonstrate the characteristics of the clustering process and conduct preliminary classification of flood flow time series data. The objectives are: 1. Identify isolated flood points; 2. Visually determine the degree of flood similarity, thereby assisting in flood classification; and 3. Provide initial values ​​for the application of other classification methods. K-medoids classification is a further optimization of hierarchical clustering, combined with a multi-clustering algorithm, to improve the rationality of flood classification and contribute to the improvement of flood classification accuracy. The method of the present invention can be applied to fields such as flood warning and flood control scheduling, and has important practical value and social significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION

[0027] like Figure 1 As shown in the figure, the flood classification method based on dynamic time warping and multi-clustering coupling is as follows:

[0028] Step (1) Collect and process flood flow time series data:

[0029] (1-1) Based on the hydrological yearbook data, collect the flow time series data of each flood in the target basin and target time range.

[0030] (1-2) Flow data cleaning: Clean the flow time series data and remove abnormal data that exceeds historical extremes and does not conform to hydrological laws.

[0031] (1-3) Hourly processing of traffic data: If the extracted traffic data is hourly hourly data, it will be retained; if the extracted traffic data is non-hourly hourly data, it will be compared with the adjacent hourly data according to the rounding principle. If there is no data at the adjacent hourly time, the non-hourly hourly data will be stored as the adjacent hourly data; if there is data at the adjacent hourly time, the data will be compared with the adjacent hourly data. If it is greater than the adjacent hourly data, the non-hourly hourly data will replace the original adjacent hourly data and be stored. Otherwise, the non-hourly hourly data will be discarded and the original hourly data will be retained unchanged.

[0032] (1-4) Interpolation of flow time series data: If the interval between adjacent flow time series data is greater than or equal to 2 hours, the intermediate process flow time series data is linearly interpolated to supplement the hourly flow time series data during the period, and finally the hourly flow time series data for the entire period (q, t) is obtained, t = {1, 2, 3, …, T}, q is the flow, t is the time, and T represents the total time;

[0033] (1-5) Standardization of flow data time series data: The hourly flow time series data q is standardized according to the year, and q' is the standardized hourly flow time series data. μ is the mean of the annual flow time series data, and σ is the standard deviation of the annual flow time series data;

[0034] (1-6) The start and end times of each flood are determined in combination with the original data, and the data set Q composed of K standardized floods is extracted according to the start time, Q = {Q1, Q2, Q3, …, QK}.

[0035] Step (2) uses the dynamic time warping method (DTW) to calculate the similarity between K floods, forming a similarity matrix.

[0036] (2-1) Determine the high-low flow demarcation line: select the flow time series data of any two floods, the flow time series data of the flood with the earlier time sequence , and the flow time series data of the flood with the later time sequence , M and N are the lengths of and respectively; considering that the similarity of high flow time series data should be more important than that of low flow time series data in the DTW similarity calculation of two floods, the demarcation line value of the high-low flow time series data of the flow time series data QA of the flood with the earlier time sequence is calculated as , is the maximum flow value of the flood, is the minimum flow value of the flood, is the proportionality coefficient;

[0037] (2-2) Determine the DTW distance coefficient: the distance between any two points on the two flow time series data , m = 1, 2, 3, …, M, n = 1, 2, 3, …, N; when calculating the distance between any two points, the high flow time series data is given a distance coefficient α, and the low flow time series data is given a distance coefficient β, considering that the greater the distance between any two points, the lower the similarity, α < β;

[0038] (2-3) Calculate the similarity of the two floods:

[0039] The similarity of the two flood flow time series data QA and QB with lengths m and n is measured, the distance between each pair of points in the two sequences is calculated, and the distance matrix based on QA and QB is obtained;

[0040] The cumulative distance of the first i elements of the flood flow time series data QA and the first j elements of the flood flow time series data QB ; represents the distance between the i th element of the flood flow time series data QA and the j th element of the flood flow time series data QB.

[0041] According to the dynamic programming principle, the optimal path r(m, n) is found by backtracking from the end point (m, n) to the start point (1, 1). The path selection after each data point matching follows continuity and monotonicity, and the boundary is consistent.

[0042] Boundary condition: the start and end points of the twisted path must correspond to the two end elements on the diagonal of the distance matrix, i.e. (1, 1) and (i, j);

[0043] Monotonicity: ensure that the twisted path must be monotonically increasing on the time axis, i.e. the i and j indices of a given path can only increase;

[0044] Continuity: only match with adjacent points, cannot skip a point;

[0045] To eliminate the influence of different twisted path lengths, the total distance is normalized by the path length L to obtain the time warping distance , .

[0046] (2-4) Calculate all flood similarity matrices. According to the flood number order, the floods are sorted from small to large, and the time warping distance between each two K-field standardized sub-flood flow time series data is calculated in turn, finally obtaining the similarity matrix K×K of K-field floods; .

[0047] Step (3) Similar flood classification, flood classification based on DTW coupled hierarchical clustering and K-medoids multi-clustering.

[0048] (3-1) Adopt the method of DTW coupled hierarchical clustering, calculate the flood clustering condensation process, remove isolated points, determine the number of flood classifications R and the corresponding classification, and finally calculate the profile coefficient S1 of the classification.

[0049] (3-1-1) In the first classification, each flood is classified as class 1, and the classification set is denoted as Q (0) , let Q (0) =Q, given the time warping distance between each two K-field standardized sub-flood flow time series data, denoted as the initial distance matrix ;

[0050] (3-1-2) Select the two sub-floods corresponding to the smallest value in the distance matrix D (0) , combine them into a class, denoted as the new class C, , , which represents the smallest number in the K×K matrix, i.e. the minimum distance; after clustering, the number of classifications will decrease by one, and the new classification set is denoted as

[0051] (3-1-3) Calculate new class C and original Q (0) The dynamic time warping distance of each class, and the distance matrix between each class after clustering is updated as

[0052] (3-1-4) Repeat (3-1-2) and (3-1-3) until clustering is complete.

[0053] (3-1-5) Draw a clustering tree of the system clustering results, and remove isolated points; the isolated point judgment condition is that each flood event in a class must satisfy ≥Y events, and those that do not satisfy are considered isolated points, 1≤Y≤3; determine a reasonable flood classification number R, 2≤R≤5, if multiple schemes simultaneously satisfy the above conditions, calculate the silhouette coefficient of each scheme, select the classification with the largest silhouette coefficient S1 and the corresponding flood classification number R as the preferred scheme.

[0054] (3-2) Select the classification number R determined by hierarchical classification as the classification number of K-medoids, and select any one of the two most similar groups of standardized flood discharge time series data in each classification after hierarchical classification as the initial center of K-medoids clustering. The two most similar groups of standardized flood discharge time series data in a classification are the two floods with the smallest dynamic time warping distance.

[0055] (3-3) Use the K-medoids clustering method to iteratively calculate flood classification until the termination condition is met, and record the silhouette coefficient of the classification as S2.

[0056] (3-3-1) Classification identification based on DTW. Calculate the distance between all data points except the initial center medoids and each K-medoids, and assign it to the cluster Ci where the medoids is located, and select dynamic time warping algorithm for distance measurement.

[0057] (3-3-2) Calculate the new center of each of the R classifications. Traverse all classifications, and for each classification, assume all data points Qi∈Ci as new center medoids, and calculate the sum of the total distance Z from all data points to the new center medoids, and select the center medoids with the smallest distance as the new center; new center , Qi traverses all points in Ci, and DTW(Qi,Qj) is the dynamic time warping distance function. The new center is the point that minimizes the total distance within Ci.

[0058] (3-3-3) Repeat (3-3-1) and (3-3-2) until the termination condition is met; the termination condition is that the number of iterations is reached or the K-medoids in all classifications no longer changes. ​

[0059] (3-3-4) Calculate the classification silhouette coefficient S2 of the DTW coupled hierarchical clustering and K-medoids.

[0060] (3-4) Select the classification with high silhouette coefficient score between S1 and S2 as the final flood classification number R', and get the corresponding classification.

[0061] The multi-cluster method coupling is not limited to the combination of the above two clustering methods. Other clustering methods such as density clustering, graph clustering, fuzzy C clustering, etc. can be introduced according to the characteristics of the data, aiming to improve the rationality and accuracy of the classification.

[0062] The flood classification system for performing the aforementioned flood classification method includes a data acquisition module, a data preprocessing module, a comparison module, and a classification output module.

[0063] The data acquisition module is used to collect the standardized flow time series data of floods in the target basin and the target time range.

[0064] The data preprocessing module is used to clean, round, interpolate, and standardize the standardized flow time series data of floods.

[0065] The comparison module is used to obtain the pairwise similarity between multiple floods.

[0066] The classification output module is used to determine the flood type and output.

Claims

1. A flood classification method based on dynamic time warping and multi-clustering coupling, characterized by: Step (1) collecting and processing flood flow time series data to obtain a dataset consisting of K standardized flood fields; Step (2) uses the dynamic time warping method to calculate the time warping distance between the two standardized floods in the K fields, and obtains the similarity matrix of the K floods; the details are as follows: (2-1) Select any two flood flow time series data and calculate the boundary value of the high and low flow time series data of the flood flow time series data that comes first based on the maximum flow value and the minimum flow value of the flood; (2-2) Set the distance coefficients of high and low flow time series data, and calculate the distance between any two points on the two flood flow time series data. The distance coefficient of high flow time series data is smaller than the distance coefficient of low flow time series data. (2-3) Using the dynamic time warping distance algorithm, the time warping distance between the two floods is calculated; (2-4) Calculate the time warp distance between the time series data of the K-field standardized flood flow according to the time sequence traversal, and finally obtain the similarity matrix of the K-field flood; Step (3) Similar flood classification, flood classification based on DTW coupled hierarchical clustering and K-medoids multi-clustering; (3-1) Using the DTW coupled hierarchical clustering method, the flood clustering process is calculated, isolated points are removed, the number of flood classifications R and the corresponding classifications are determined, and finally the silhouette coefficient S1 of the classification is calculated; (3-2) The number of categories R determined by hierarchical classification is selected as the number of categories of K-medoids‌, and any of the two most similar sets of standardized secondary flood flow time series data in each category after hierarchical classification is selected as the initial center of the K-medoids‌ cluster; the two most similar sets of standardized secondary flood flow time series data in a category are the two floods with the smallest dynamic time warping distance; (3-3) The K-medoids‌ clustering method is used to iteratively calculate the flood classification until the termination condition is met. The silhouette coefficient of the classification is recorded as S2; (3-4) Select the classification with the higher silhouette coefficient score between the silhouette coefficients S1 and S2 as the final flood classification number R', and obtain the corresponding classification.

2. The flood classification method based on dynamic time warping and multi-clustering coupling according to claim 1 is characterized in that: Step (1) is as follows: (1-1) Collect flow time series data within the target basin and target time range based on hydrological yearbook data; (1-2) Clean the flow time series data and remove abnormal data that does not conform to the hydrological laws; (1-3) If the extracted traffic data is hourly data, it shall be retained; if the extracted traffic data is non-hourly data, it shall be compared with the adjacent hourly data according to the rounding principle. If there is no data at the adjacent hourly data, the non-hourly data shall be stored as the adjacent hourly data; if there is data at the adjacent hourly data, the data shall be compared with the adjacent hourly data. If it is greater than the adjacent hourly data, the non-hourly data shall replace the original adjacent hourly data and be stored. Otherwise, the non-hourly data shall be discarded and the original hourly data shall be retained unchanged. (1-4) If the interval between adjacent flow time series data is greater than or equal to 2 hours, the intermediate process flow time series data is supplemented with the hourly flow time series data during the period by linear interpolation, and finally the hourly flow time series data for the entire period is obtained; (1-5) Based on the mean value and standard deviation of annual flow time series data, the hourly flow time series data are standardized according to the year; (1-6) The start and end time of each flood are verified by combining the original data, and a data set consisting of K standardized floods is obtained according to the start time.

3. The flood classification method based on dynamic time warping and multi-clustering coupling according to claim 1 is characterized in that: Steps (2-3) calculate the time warp distance between the two floods as follows: Perform similarity measurement on two flood flow time series data QA and QB with lengths m and n respectively, calculate the distance between each pair of points in the two series, and obtain the distance matrix based on QA and QB; The cumulative distance between the first i elements of the flood flow time series data QA and the first j elements of the flood flow time series data QB ; Represents the distance between the i-th element of the flood flow time series data QA and the j-th element of the flood flow time series data QB; According to the principle of dynamic programming, we backtrack from the end point (m,n) to the starting point (1,1) to find the optimal path r(m,n). The path selection after each data point matching follows continuity and monotonicity, and the boundaries are consistent. Normalize the total distance by the path length L to get the time warp distance , .

4. The flood classification method based on dynamic time warping and multi-clustering coupling according to claim 1 is characterized in that: Step (3-1) is as follows: (3-1-1) During the first classification, each flood is considered as Class 1, and the classification set is recorded as Q (0) , let Q (0) =Q, the time warped distance between the time series data of the normalized secondary flood flow in the known K fields, recorded as the initial distance matrix ; (3-1-2) Select the distance matrix D (0) The two floods corresponding to the minimum value in the , are merged into one category, recorded as new category C, , Represents the smallest number in the K×K matrix, that is, the minimum distance; after clustering is completed, the number of categories will be reduced by one, and the new category set will be recorded as ; (3-1-3) Calculate the new class C and the original Q (0) The dynamic time warping distance of each category, after clustering, the distance matrix between each category is updated to ; (3-1-4) Repeat (3-1-2) and (3-1-3) until they are clustered into one category; (3-1-5) Draw a cluster tree based on the system clustering results and remove isolated points; Isolation point judgment conditions: Assume that each type of flood event must meet ≥Y events, and those that do not meet the conditions are considered isolated points, 1≤Y≤3; Determine a reasonable number of flood classifications R, 2≤R≤5. If there are multiple options that meet the above conditions at the same time, calculate the silhouette coefficient of each option separately, and select the category with the largest silhouette coefficient S1 and the corresponding flood classification number R as the preferred option.

5. The flood classification method based on dynamic time warping and multi-clustering coupling according to claim 1 is characterized in that: Step (3-3) is as follows: (3-3-1) DTW-based classification and recognition; calculate the distance between all data points except the initial center medoids and each K-medoids, and assign them to the cluster Ci where the closest medoids are located. The distance metric uses the dynamic time warping algorithm; (3-3-2) Calculate the new center of each of the R categories; traverse all categories, assume that all data points Qi∈Ci are new center medoids for each category, and calculate the total distance Z of all other data points to the new center medoids, and select the center medoids corresponding to the minimum distance as the new center; the new center , Qi traverses all points in Ci, DTW(Qi,Qj) is the dynamic time warping distance function, the new center It is the point that minimizes the total distance inside Ci; (3-3-3) Repeat (3-3-1) and (3-3-2) until the termination condition is met; the termination condition is that the set number of iterations is reached or the K-medoids in all categories no longer change; (3-3-4) Calculate the classification silhouette coefficient S2 of DTW coupled hierarchical clustering and K-medoids.

6. A system for implementing the flood classification method based on dynamic time warping and multi-clustering coupling as described in any one of claims 1 to 5, characterized in that: Including data acquisition module, data preprocessing module, comparison module, classification output module; The data acquisition module is used to collect flow time series data within the target basin and target time range; The data preprocessing module is used to clean, process the flood flow time series data at the whole time, interpolate and standardize them; The comparison module is used to obtain pairwise similarities between multiple floods; The classification output module is used to obtain and output the determined flood type.

Citation Information

Patent Citations

  • A method and apparatus for sampling data

    CN109508350A

  • Flood classification forecasting method based on DTW hierarchical clustering

    CN115345244A