Hydropower station similar output query method based on fuzzy C-means clustering and dynamic time warping algorithm

By applying fuzzy C-mean clustering and dynamic time regularization algorithms in the output data query of hydropower stations, the problem that traditional methods are difficult to identify similar output patterns is solved, and accurate query of historical output data of hydropower stations is realized, which improves the operation management level and economic benefits.

CN120144636APending Publication Date: 2025-06-13CHINA YANGTZE POWER
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510219422.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional hydropower station output data query methods are difficult to accurately identify similar output modes, and there are problems of low efficiency and insufficient accuracy. They cannot effectively explore historical periods with similar output modes, and cannot provide a comprehensive reference for hydropower station operation decisions.

Method used

Using the methods based on fuzzy C mean clustering and dynamic time regularization algorithm, the steps of calculating similarity and similar output query results are realized through data acquisition and preprocessing, fuzzy C mean clustering, building clustering center time series, and dynamic time regularization algorithms to calculate the output of similar output query results.

Benefits of technology

It can more accurately find similar output situations in the historical output data of hydropower stations, provide more valuable reference, and improve the operation and management level and economic benefits of hydropower stations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144636A_ABST
    Figure CN120144636A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of hydropower station operation data processing, and particularly provides a hydropower station similar output query method based on fuzzy C-means clustering and a dynamic time warping algorithm. Fuzzy C-means clustering: determining a clustering number C, initializing a clustering center, calculating a membership matrix and updating the clustering center until convergence; constructing a clustering center time sequence: constructing a time sequence corresponding to each clustering center; calculating similarity by using a dynamic time warping algorithm: calculating a similarity index of the to-be-queried output time sequence and each clustering center time sequence by using the dynamic time warping algorithm; and outputting a similar output query result: sorting the clusters according to the similarity index, and outputting a similar output clustering result. Compared with a traditional method, the method has the advantages that the similar output condition in historical output data of the hydropower station can be found more accurately, and a more valuable reference basis is provided for operation optimization, fault diagnosis, scheduling decision and the like of the hydropower station.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of hydropower station operation data processing. Specifically, it relates to a method for querying similar output of a hydropower station based on the fuzzy C-means clustering and dynamic time warping algorithms. Background Art

[0002] Hydropower stations play an important role in the power system. The output situation of hydropower stations is of great significance for their efficient operation, power grid dispatching, etc. Its output situation is affected by various factors, such as basin inflow, water level drop, unit operation status, etc. Effective analysis of historical output data of hydropower stations and querying of similar working conditions are helpful for optimizing hydropower station operation scheduling, equipment maintenance, and fault diagnosis, etc.

[0003] Traditional output data queries often rely only on simple numerical comparisons and the like. When dealing with complex hydropower station output data, it is often difficult to accurately identify similar output patterns, and there are problems such as low efficiency and insufficient accuracy. It is difficult to effectively mine historical periods with similar output patterns and cannot provide a comprehensive reference for hydropower station operation decisions well. By using advanced data mining and analysis algorithms, similar output situations can be found more accurately, assisting dispatching and operation personnel to better grasp the operation characteristics of hydropower stations. However, there is currently a lack of technical solutions for applying relevant effective integration algorithms to this specific query scenario. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method for querying similar output of a hydropower station based on the fuzzy C-means clustering and dynamic time warping algorithms, which can more accurately find similar output situations in the historical output data of hydropower stations.

[0005] To solve the above technical problem, the technical solution adopted by the present invention is: A method for querying similar output of a hydropower station based on the fuzzy C-means clustering and dynamic time warping algorithms, comprising the following steps: Step 1: Data collection and preprocessing: Collect historical output data of the hydropower station, and then perform data cleaning and normalization processing; Step 2: Fuzzy C-means clustering: Determine the number of clusters C, initialize the cluster centers, calculate the membership matrix and update the cluster centers until convergence; Step 3: Construct the time series of the cluster centers: Construct the time series C i , where, i = 1, 2, …, C; Step 4: Calculate similarity using the dynamic time warping algorithm: Use the dynamic time warping algorithm to calculate the similarity index between the output time series to be queried and the time series of each cluster center; Step 5. Output of similar output query results: Sort the clusters according to the similarity index and output the similar output clustering results.

[0006] In the preferred solution, in the above Step 1, when performing data cleaning, for missing data, linear interpolation is used to fill it according to the output values at adjacent time points, and for outliers beyond the range, data elimination is performed.

[0007] In the preferred solution, the data elimination is to eliminate the data points beyond the rated output range.

[0008] In the preferred solution, Step 2 includes the following steps: S201. Determine the number of clusters C; S202. Initialize the cluster centers; S203. Calculate the membership matrix; S204. Update the cluster centers; S205. Repeat Steps S203 - 205 until the convergence condition is met.

[0009] In the preferred solution, in Step S201, the method for determining the number of clusters C includes the elbow method. By plotting the cluster error index curves for different numbers of clusters, the number of clusters corresponding to the inflection point of the curve is selected as the number of clusters C.

[0010] In the preferred solution, in Step S202, randomly select C data points as the initial cluster center vectors V = [v 1 , v 2 , v i , …, v C , where v i represents the i -th cluster center, i = 1, 2, …, C .

[0011] In the preferred solution, in Step S203, for the membership degree of each data point i belonging to the u ij -th cluster, it is calculated according to the following formula: (1); where m is the fuzzy index, ||·|| represents the distance metric between the data point and the cluster center, j = 1, 2, …, N , N is the total number of data points, represents thek a clustering center

[0012] In a preferred solution, in the step S204, the clustering center is updated according to the following formula: (2); wherein represents the membership degree of the data point belonging to the i th cluster when the fuzzy index is m; In a preferred solution, in the fourth step, for the output time series Q of the hydropower station to be queried, it is successively subjected to dynamic time warping calculation with each clustering center time series C i ; The dynamic time warping algorithm constructs a cost matrix D, where D(i,j) represents the distance between the i-th point of the sequence Q and the i -th point of the sequence C j , and then finds an optimal path from the upper left corner to the lower right corner of the matrix through dynamic programming, so that the sum of the distances of the point pairs on the path is minimized, and the minimum distance is the dynamic time warping distance between the sequence Q and the clustering center sequence C i DTW(Q,C i ) .

[0013] In a preferred solution, in the fourth step, the similarity index is calculated according to the dynamic time warping distance, and the similarity index S i is calculated according to the following formula: (3); wherein, max( DTM ) is the maximum dynamic time warping distance calculated between all clustering center sequences and the sequence to be queried.

[0014] A method for querying similar output of a hydropower station based on fuzzy C-means clustering and dynamic time warping algorithm provided by the present invention combines fuzzy C-means clustering and dynamic time warping algorithm, can fully explore similar patterns in the historical output data of the hydropower station, and compared with traditional methods, can more accurately find the situations similar to the target output period under different working conditions, provides a more valuable reference basis for various aspects such as operation optimization, fault diagnosis, and dispatching decision-making of the hydropower station, and helps to improve the overall operation management level and economic benefits of the hydropower station. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The present invention will be further described below with reference to the drawings and embodiments: Figure 1 is the overall flow block diagram of the method of the present invention; ​Figure 2 It is the program flow chart of the present invention; Figure 3 It is the similar output query result graph in the embodiment; Figure 4 It is the similar output comparison graph between the retrieval date and the similar date in the embodiment. Specific embodiments

[0016] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0017] Embodiment 1: As Figure 1 and 2 shown, a method for querying similar output of a hydropower station based on the fuzzy C-means clustering and dynamic time warping algorithms includes the following steps: Step 1. Data collection and preprocessing: Collect the historical output data of the hydropower station, and then perform data cleaning and normalization processing.

[0018] When performing data cleaning, for missing data, linear interpolation method is used to fill in according to the output values of adjacent time points, and for outliers beyond the rated output range, data elimination is performed.

[0019] After removing outliers and missing values, then perform normalization processing to map the data to a specific interval, such as the interval [0, 1], for subsequent clustering and similarity calculation.

[0020] Step 2. Fuzzy C-means clustering: Determine the number of clusters C, initialize the cluster centers, calculate the membership matrix and update the cluster centers until convergence.

[0021] Specifically, it includes the following steps: S201. Determine the number of clusters C: According to the actual operation characteristics of the hydropower station and the data analysis requirements, determine the appropriate number of clusters C. For example, through empirical methods such as the elbow method, draw the cluster error index curves under different numbers of clusters, and select the number of clusters corresponding to the inflection point of the curve as the optimal C value.

[0022] S202. Initialize the cluster centers: Randomly select C data points as the initial cluster center vector V = [v 1 , v 2 , v i , …, v C , where v i represents the i th cluster center,i =1,2,…, C .

[0023] S203, calculate the membership matrix: For each data point Belong to i The membership of the cluster u ij , calculated according to the following formula: (1); in, m is the fuzzy index, which is generally between 1.5 and 2.5, and ||·|| represents the distance measure between the data point and the cluster center, such as the Euclidean distance; j =1,2,…, N , N is the total number of data points, Indicates k Cluster centers.

[0024] S204, update cluster center: Cluster Center The update formula is: (2); in, Indicates that when the fuzzy index is m, the data point Belong to i The degree of membership of a cluster.

[0025] S205, repeating steps S203-205 until the convergence condition is met, for example, the change in the cluster center is less than a preset threshold ε or the number of iterations reaches a set maximum value M, then the iteration is stopped.

[0026] Step 3: Construct the time series of cluster centers: Construct the time series C corresponding to each cluster center i ,in, i =1,2,…,C.

[0027] The time series corresponding to each cluster center is constructed, which reflects the changing trend of the output mode represented by the cluster center over time.

[0028] Step 4: Calculate similarity using dynamic time warping algorithm: Use the dynamic time warping algorithm to calculate the similarity index between the output time series to be queried and the time series of each cluster center.

[0029] For the hydropower station output time series Q to be queried, it is sequentially compared with each cluster center time series C i Perform dynamic time warping calculation; the dynamic time warping algorithm constructs a cost matrix D, whereD(i,j) represents the distance between the i-th point of sequence Q and the i point of sequence C j . Then, by using the method of dynamic programming, an optimal path from the upper left corner to the lower right corner of the matrix is found, so that the sum of the distances between the point pairs on the path is minimized, and the minimum distance is the dynamic time warping distance between sequence Q and the clustering center sequence C i . DTW(Q,C i ) .

[0030] The similarity index is calculated based on the dynamic time warping distance, and the calculation formula is: (3); where, max( DTM ) is the maximum dynamic time warping distance calculated for all clustering center sequences and the sequence to be queried.

[0031] The similarity index S i ranges from [0, 1], and the closer it is to 1, the more similar it indicates.

[0032] Step Five, Output of Similar Output Query Results: According to the calculated similarity index, sort the clusters, select the top several clusters with higher similarity indices as the similar output clustering results, and output the hydropower station output data and corresponding time information in these clusters, etc., for further analysis and decision-making.

[0033] Example 2: In this example, the example scenario is set as follows: Assume that a hydropower station has output data recorded every 15 minutes in the past year, and its full-plant rated output range is between 0 and 3 million kilowatts. Now, it is necessary to query the historical periods similar to the output situation of a certain continuous 4-hour period recently to assist in analyzing whether the current operating state is normal and referring to historical dispatching strategies, etc.

[0034] Extract the output data of the past year from the hydropower station's database. After inspection, it is found that there are a small amount of missing data and individual data points that significantly exceed the range of 0 to 3 million kilowatts. For the missing data, linear interpolation is used to fill it according to the output values of adjacent time points; for the abnormal out-of-range data points, they are directly removed. After processing, a complete and accurate historical output data set is obtained.

[0035] Based on experience and the actual operating conditions of this hydropower station, the number of clusters C = 3 is determined, corresponding to three output operating condition categories of low (0 - 1 million kilowatts), medium (1.01 - 2 million kilowatts), and high (2.01 - 3 million kilowatts). The fuzzy index m = 2 is set. After randomly initializing the cluster centers, iterative calculations are performed according to the membership degree calculation and cluster center update formulas in Example 1, and the convergence threshold ε is set to 0.0001. After multiple iterations, clustering is completed, and the historical output data is divided into these 3 categories.

[0036] Determine the output time period of the continuous 4 hours that needs to be queried recently as the target output time period. For example, from 16:00 to 20:00 on July 1, 2024, the average output of the whole plant is 1.49 million kilowatts, the standard deviation is 50,000 kilowatts, the maximum value is 1.7 million kilowatts, the minimum value is 1.3 million kilowatts, etc. These statistical characteristics form the target feature vector.

[0037] For the historical output time periods in each cluster category, intercept them with a 4 - hour time window, extract the corresponding feature vectors, and then use the DTW (Dynamic Time Warping) algorithm to calculate the similarity with the target feature vector. The similarity threshold is set to 0.2. After calculation and screening, multiple historical 4 - hour output time periods with similarity higher than the threshold are found in different cluster categories.

[0038] Finally, sort the clusters according to the similarity index, select the top 3 clusters with higher similarity as the similar output query results, and visually display the time range of the found similar historical output time periods, the corresponding output data, etc. in a chart, and output it to the hydropower station operation and management personnel so that they can refer to the historical similar situations for current operation status analysis, scheduling decision - making and other operations.

[0039] The clustering evaluation index is an important measurement method for measuring the quality of clustering. Clustering evaluation indexes such as the Rand index and the silhouette coefficient are widely used in clustering analysis. The closer the value of the evaluation index is to 1, the better the clustering effect, and the more ideal the corresponding similar output query results. Combining the evaluation indexes, the query results of the combined model provided by the present invention are compared with those of the K - means and fuzzy C - means clustering models as shown in Table 1. It can be seen that the clustering evaluation index result of the method of the present invention is the best, which proves the effectiveness of the method provided by the present invention.

[0040]

[0041] In summary, the present invention applies a method for querying similar power outputs of a hydropower station based on the fuzzy C-means clustering and dynamic time warping algorithms to data mining of similar power outputs within a certain period of a certain hydropower station. Under the given conditions, the model provided by the present invention can fully mine the similar patterns in the historical power output data of the hydropower station. Compared with traditional methods, it can more accurately find the situations similar to the target power output period under different working conditions, providing a more valuable reference basis for various aspects such as operation optimization, fault diagnosis, and dispatching decision-making of the hydropower station, and helping to improve the overall operation management level and economic benefits of the hydropower station.

[0042] It is easy for those skilled in the art to understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for querying similar output of hydropower stations based on fuzzy C-means clustering and dynamic time warping algorithm, characterized in that: The following steps are involved: Step 1: Data collection and preprocessing: Collect historical output data of hydropower stations, and then perform data cleaning and normalization; Step 2: Fuzzy C-means clustering: determine the number of clusters C, initialize the cluster centers, calculate the membership matrix and update the cluster centers until convergence; Step 3: Construct the time series of cluster centers: Construct the time series C corresponding to each cluster center i ,in, i =1,2,…,C; Step 4: Calculate similarity using dynamic time warping algorithm: Use dynamic time warping algorithm to calculate the similarity index between the output time series to be queried and the time series of each cluster center; Step 5: Output of similar output query results: Sort clusters according to similarity indicators and output similar output cluster results.

2. According to claim 1, a method for querying similar output of hydropower stations based on fuzzy C-means clustering and dynamic time warping algorithm is characterized in that: In the step 1, when data cleaning is performed, missing data is filled in according to the output values ​​of adjacent time points using linear interpolation, and outliers that are out of range are removed.

3. A method for querying similar output of hydropower stations based on fuzzy C-means clustering and dynamic time warping algorithm according to claim 2, characterized in that: The data elimination is performed to eliminate data points that exceed the rated output range.

4. The method for querying similar output of hydropower stations based on fuzzy C-means clustering and dynamic time warping algorithm according to claim 1 is characterized in that: The step 2 includes the following steps: S201, determining the number of clusters C; S202, initializing cluster centers; S203, calculating the membership matrix; S204, updating cluster centers; S205. Repeat steps S203 to S205 until the convergence condition is met.

5. A method for querying similar output of hydropower stations based on fuzzy C-means clustering and dynamic time warping algorithm according to claim 4, characterized in that: In step S201, the method for determining the number of clusters C includes the elbow method, which is to draw clustering error index curves under different numbers of clusters and select the number of clusters corresponding to the inflection point of the curve as the number of clusters C.

6. A method for querying similar output of hydropower stations based on fuzzy C-means clustering and dynamic time warping algorithm according to claim 4, characterized in that: In step S202, C data points are randomly selected as the initial cluster center vector V=[v1,v2, v i ,…,v C ],in v i Indicates i Cluster centers, i =1,2,…, C .

7. A method for querying similar output of hydropower stations based on fuzzy C-means clustering and dynamic time warping algorithm according to claim 4, characterized in that: In step S203, for each data point Belong to i The membership of the cluster u ij , calculated according to the following formula: (1); in, m is the fuzzy index, ||·|| represents the distance measure between the data point and the cluster center, j =1,2,…, N , N is the total number of data points, Indicates k Cluster centers.

8. The method for querying similar output of hydropower stations based on fuzzy C-means clustering and dynamic time warping algorithm according to claim 4 is characterized in that: In step S204, the cluster center The update formula is: (2); in, Indicates that when the fuzzy index is m, the data point Belong to i The degree of membership of a cluster.

9. The method for querying similar output of hydropower stations based on fuzzy C-means clustering and dynamic time warping algorithm according to claim 1 is characterized in that: In the step 4, for the hydropower station output time series Q to be queried, each cluster center time series C i Perform dynamic time warping calculation; the dynamic time warping algorithm constructs a cost matrix D, where D(i,j) Represents the i-th point of sequence Q and sequence C i No. j The distance between the points is calculated, and then the optimal path from the upper left corner to the lower right corner of the matrix is ​​found through dynamic programming, so that the sum of the distances between the points on the path is minimized. The minimum distance is the distance between the sequence Q and the cluster center sequence C. i Dynamic Time Warping Distance DTW(Q,C i ) .

10. A method for querying similar output of hydropower stations based on fuzzy C-means clustering and dynamic time warping algorithm according to claim 9, characterized in that: In step 4, the similarity index is calculated based on the dynamic time warping distance. S i The calculation formula is: (3); Among them, max( DTM ) is the maximum dynamic time warping distance calculated between all cluster center sequences and the query sequence.

Citation Information

Cited By

  • Flood classification method and system based on dynamic time warping and multi-cluster coupling

    CN120493104A

  • District line loss rate abnormity diagnosis method based on data driving

    CN120850156A