A Classification Method for the Operating Conditions of a Turbine Main Shaft Based on Data Analysis

By constructing the spindle power data set and optimizing the objective function of the K-Means algorithm, combining the population optimization algorithm and individual variation strategy, the problem of local optimality in spindle operating condition classification is solved, accurate operating condition classification and reliable clustering results are achieved, and reliable data support is provided for the intelligent management of the turbine.

CN119357725BActive Publication Date: 2025-07-08WUHAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411383117.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2025-07-08
Estimated Expiration
2044-09-30

AI Technical Summary

Technical Problem

In the prior art, the spindle operating condition classification method is too simplified and cannot fully capture the complexity of the operating condition, resulting in insufficient reliability of the classification results. The K-Means clustering algorithm is susceptible to the initial clustering center, resulting in a local optimal solution, limiting its applicability in complex operating conditions.

Method used

By collecting spindle stress data to calculate spindle power, building spindle power data set, calculating the overall contour coefficient, optimizing the objective function and classification model of the K-Means algorithm, iterative training is performed in combination with the population optimization algorithm, selecting the optimal number of clusters and distance metrics, using random initialization, chaotic mapping and random reverse learning to initialize population position, introducing individual variation strategies, dynamically update the cluster center until the results converge.

Benefits of technology

The precise classification of spindle operating conditions is achieved, the clustering effect and parameter accuracy is improved, the local optimal dilemma is avoided, the stability and reliability of clustering results are ensured, and a detailed classification report is provided to support the intelligent management of the turbine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119357725B_ABST
    Figure CN119357725B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of main shaft condition classification, and proposes a method for classifying the working conditions of a water turbine main shaft based on data analysis, including the following steps: collecting main shaft stress data, calculating the main shaft power based on the main shaft stress data, and pre-classifying the main shaft working conditions according to the main shaft power; constructing a main shaft power data set through the main shaft power, calculating the overall silhouette coefficient of the main shaft power data set, and respectively constructing a first objective function of the K-Means algorithm and a second objective function of the main shaft condition classification model based on the overall silhouette coefficient; selecting the first optimal number of clusters, the optimal Euclidean distance, and the optimal Manhattan distance through the second objective function; setting the number of individuals and the total number of iterations, iteratively training the second objective function to obtain the optimized individual positions; inputting the optimized individual positions into the first objective function, and performing clustering iteration on the first objective function based on the first optimal number of clusters, the optimal Euclidean distance, and the optimal Manhattan distance to obtain the second final number of clusters, and performing optimal classification of the main shaft working conditions based on the second final number of clusters. The present invention improves the convergence accuracy and iteration speed of the main shaft condition classification, effectively avoids falling into the local optimal dilemma, and realizes the accurate classification of the main shaft working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of main shaft condition classification, and particularly to a method for classifying the conditions of a water turbine main shaft based on data analysis. Background Art

[0002] Currently, for the main shaft condition classification method, traditional empirical classification methods often appear to be overly simplified and unable to fully capture the complexity of the conditions. This single classification method affects the reliability of the classification results, thus limiting the in-depth analysis of the main shaft state and fatigue life. Although existing technologies such as the K-Means clustering algorithm have been applied to condition classification, the clustering effect and efficiency of this algorithm are greatly affected by the initial clustering centers, easily leading to the generation of local optimal solutions, which limits its applicability in complex conditions. Therefore, it is urgent to explore a better clustering method to achieve more accurate condition classification. Summary of the Invention

[0003] In view of this, the present invention proposes a method for classifying the conditions of a water turbine main shaft based on data analysis, which solves the problem that the existing technology is only sensitive to the initial clustering centers and easily generates local optima.

[0004] The technical solution of the present invention is realized as follows: The present invention provides a method for classifying the conditions of a water turbine main shaft based on data analysis, including the following steps:

[0005] S1. Collect main shaft stress data, calculate the main shaft power based on the main shaft stress data, and pre-classify the main shaft conditions according to the main shaft power;

[0006] S2. Construct a main shaft power data set through the main shaft power, calculate the overall silhouette coefficient of the main shaft power data set, and respectively construct a first objective function of the K-Means algorithm and a second objective function of the main shaft condition classification model based on the overall silhouette coefficient;

[0007] S3. Select the first optimal number of clusters, the optimal Euclidean distance, and the optimal Manhattan distance through the second objective function;

[0008] S4. Set the number of individuals and the total number of iterations, perform iterative training on the second objective function, and obtain the optimized individual positions;

[0009] S5. Input the optimized individual positions into the first objective function, perform clustering iteration on the first objective function based on the first optimal number of clusters, the optimal Euclidean distance, and the optimal Manhattan distance, obtain the second final number of clusters, and perform the best classification of the main shaft conditions based on the second final number of clusters.

[0010] Based on the above technical solution, preferably, step S1 includes:

[0011] S11. Set the main shaft stress sensor on the main shaft of the water turbine, and collect the original main shaft stress data through the main shaft stress sensor;

[0012] S12. Preprocess the original main shaft stress data to obtain the main shaft stress data. The preprocessing includes data filtering and normalization processing;

[0013] S13. Calculate the main shaft power by combining the main shaft stress data through the torque angular velocity method;

[0014] S14. Pre-classify the operating conditions of the main shaft of the water turbine based on the main shaft power. The pre-classification includes dividing the main shaft power into multiple preset power intervals, and dividing different preset power intervals into different pre-classified operating conditions.

[0015] Based on the above technical solutions, preferably, step S2 includes:

[0016] S21. Construct a main shaft power data set according to the main shaft power data. The main shaft power data set contains multiple main shaft power data points;

[0017] S22. Calculate the silhouette coefficient of each main shaft power data point. The silhouette coefficient is used to evaluate the relative position of the main shaft power data point in the clustering;

[0018] S23. Calculate the overall silhouette coefficient of the main shaft power data set based on the silhouette coefficient of each main shaft power data point;

[0019] S24. Construct the first objective function of the K-Means algorithm based on the overall silhouette coefficient. The first objective function is used to optimize the clustering effect;

[0020] S25. Construct the second objective function of the main shaft operating condition classification model based on the overall silhouette coefficient. The second objective function is used to optimize the clustering parameters.

[0021] Based on the above technical solutions, preferably, step S3 includes:

[0022] S31. Optimize the clustering parameters through the second objective function, calculate the overall silhouette coefficient under different clustering numbers, and select the best clustering number based on the clustering effect;

[0023] S32. Calculate the best Euclidean distance and the best Manhattan distance of the best clustering number as the distance metric standard of the K-Means algorithm;

[0024] S33. Input the first best clustering number, the best Euclidean distance, and the best Manhattan distance into the K-Means algorithm.

[0025] Based on the above technical solutions, preferably, step S4 includes:

[0026] S41. Set the number of individuals and the total number of iterations. The number of individuals is the number of candidate solutions used in the optimization process.

[0027] S42. Initialize the positions of the population. Adopt a population initialization strategy, which includes random initialization, chaotic mapping, and random reverse learning.

[0028] S43. Calculate the fitness value of each individual under the second objective function to evaluate its clustering effect.

[0029] S44. Sort according to the fitness value and select the individual with the highest fitness value as the current optimal solution.

[0030] S45. Execute the optimization iteration process, which includes an exploration stage, an exploitation stage, and an exploration stage, and update the positions of the individuals to find a better solution for the current optimal solution.

[0031] S46. Increase the population diversity through an individual mutation strategy and update the positions of the individuals.

[0032] S47. After each iteration, determine whether the set number of iterations is reached. When the set number of iterations is not reached, return to step S43 to continue the iteration. When the set number of iterations is reached, output the position of the current optimal optimized individual.

[0033] Based on the above technical solutions, preferably, step S4 further includes:

[0034] The calculation formula for the random initialization is:

[0035] χ i :x ij =ll j +r·(ul j -ll j ), i = 1, 2,..., N, j = 1, 2,..., M;

[0036] Among them, χ i is the position of the i-th candidate solution, x ij is the position coordinate of the i-th candidate solution, ll j and ul j respectively represent the lower bound and the upper bound of the j-th decision variable, r is a random value between 0 and 1, N is the number of candidate solutions, and M is the number of decision variables;

[0037]

[0038] The calculation formula for the chaotic mapping is:

[0039]

[0040] Among them, X k+1 is the number of individuals after the k-th chaotic mapping, and X k is the number of individuals before the k-th chaotic mapping, and mod(·) is the modulo function;

[0041] The calculation formula of the stochastic reverse learning is as follows:

[0042]

[0043] Among them, is the reverse solution of X k , U b and L b are respectively the upper and lower bounds of the problem to be optimized, and J is a random value between 0 and 1.

[0044] Based on the above technical solution, preferably, the calculation formula of the individual mutation strategy is:

[0045] G hippo = D hippo [1 + aC(0, σ 2 ) + bG(0, σ 2 )∑;

[0046]

[0047] Among them, G hippo is the position of the individual with the highest fitness after mutation, D hippo is the position of the individual with the highest fitness before mutation, C(0, σ 2 ) is a Cauchy distribution random variable, G(0, σ 2 ) is a Gaussian distribution random variable, a and b are respectively the Cauchy distribution dynamic parameter and the Gaussian distribution dynamic parameter, is the current individual, f(·) is the fitness function, exp(·) is the exponential function, and σ is the standard deviation.

[0048] Based on the above technical solution, preferably, step S5 includes:

[0049] S51, input the first best clustering number, the best Euclidean distance, and the best Manhattan distance into the K-Means algorithm;

[0050] S52, based on the input parameters, perform K-Means clustering iterative calculation on the spindle power data set;

[0051] S53, during the clustering iteration process, dynamically update the clustering center of each data point until the clustering result converges to obtain the second best clustering number;

[0052] S54. Classify the main shaft operating conditions finally according to the second best clustering number, and output the classification result.

[0053] On the basis of the above technical solutions, preferably, step S52 includes:

[0054] At the beginning of each iteration, calculate the distance between the current clustering center and each data point, and assign the data point to the nearest clustering center; at the end of each iteration, calculate the mean value of each clustering center to update the position of the clustering center, and record the clustering result and the position change of the clustering center for each iteration.

[0055] On the basis of the above technical solutions, preferably, step S54 includes:

[0056] According to the second best clustering number, calculate the silhouette coefficient of each cluster to evaluate the clustering effect; compare the clustering result with the preset operating condition standard to generate a classification report, pointing out the characteristics of each operating condition category; when outputting the classification result, provide the sample number and main statistical characteristics of each operating condition category, and classify the operating conditions based on the sample number and the main statistical characteristics.

[0057] A method for classifying the operating conditions of a water turbine main shaft based on data analysis according to the present invention has the following beneficial effects compared with the prior art:

[0058] (1) By collecting the main shaft stress data and calculating the main shaft power, realize the pre-classification of the main shaft operating conditions, construct the main shaft power data set, and calculate the overall silhouette coefficient, so as to optimize the objective function and classification model of the K-Means algorithm, thereby improving the clustering effect and the accuracy of parameters. By optimizing the objective function, select the best clustering number and distance metric, combine the population optimization algorithm for iterative training, and introduce the individual mutation strategy to increase the diversity of the population, thereby improving the convergence accuracy and iteration speed, and then effectively avoiding falling into the local optimal dilemma, and finally realizing the accurate classification of the main shaft operating conditions;

[0059] (2) By setting the number of individuals and the total number of iterations, use random initialization, chaotic mapping and random reverse learning to initialize the population position. By calculating the fitness value of each individual under the second objective function, evaluate the clustering effect and select the individual with the highest fitness value as the current optimal solution. Execute the exploration, exploration and exploitation stages in the optimization iteration process, which helps to continuously update the individual position and find a better solution. At the same time, the individual mutation strategy is used to increase the population diversity and improve the clustering optimization effect;

[0060] (3) By inputting the optimal clustering parameters into the K-Means algorithm, clustering is performed under optimal conditions to ensure the accuracy of the clustering results. During the clustering iteration process, the clustering centers of each data point are dynamically updated until the clustering results converge, obtaining a more stable and reliable second optimal number of clusters, laying a foundation for the final working condition classification. Record the clustering results of each iteration and the position changes of the clustering centers, providing a basis for subsequent analysis of the clustering results and helping to evaluate the clustering effect;

[0061] (4) The characteristics of each working condition category are pointed out through the generated classification report, providing the sample quantity and main statistical characteristics, enhancing the interpretability of the classification results. Based on the sample quantity and statistical characteristics, the working condition classification is carried out, fully mining the working condition information contained in the data, providing a reliable basis for the intelligent management of the water turbine. Brief Description of the Drawings

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0063] Figure 1 It is a flowchart of a method for classifying the working conditions of a water turbine spindle based on data analysis according to the present invention. Detailed Embodiments

[0064] The following will combine the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0065] Please refer to Figure 1 , and a method for classifying the working conditions of a water turbine spindle based on data analysis is provided, including the following steps:

[0066] S1. Collect the spindle stress data, calculate the spindle power based on the spindle stress data, and pre-classify the spindle working conditions according to the spindle power;

[0067] S2. Construct a spindle power data set through the spindle power, calculate the overall silhouette coefficient of the spindle power data set, and respectively construct the first objective function of the K-Means algorithm and the second objective function of the spindle working condition classification model based on the overall silhouette coefficient;

[0068] S3. Select the first optimal number of clusters, the optimal Euclidean distance, and the optimal Manhattan distance through the second objective function;

[0069] S4. Set the number of individuals and the total number of iterations, and perform iterative training on the second objective function to obtain the optimized individual positions;

[0070] S5. Input the optimized individual positions into the first objective function, perform clustering iteration on the first objective function based on the first optimal number of clusters, the optimal Euclidean distance, and the optimal Manhattan distance to obtain the second final number of clusters, and perform optimal classification on the main shaft working conditions based on the second final number of clusters.

[0071] Specifically, in this embodiment, by collecting the main shaft stress data and calculating the main shaft power, the pre-classification of the main shaft working conditions is realized, the main shaft power dataset is constructed, and the overall silhouette coefficient is calculated, so as to optimize the objective function and classification model of the K-Means algorithm, thereby improving the clustering effect and the accuracy of the parameters. By optimizing the objective function, selecting the optimal number of clusters and distance metrics, and combining the population optimization algorithm for iterative training, the individual mutation strategy is introduced to increase the diversity of the population, thereby improving the convergence accuracy and iterative speed, and effectively avoiding falling into the local optimal dilemma, and finally realizing the accurate classification of the main shaft working conditions.

[0072] Step S1 includes:

[0073] S11. Set the main shaft stress sensor on the water turbine main shaft, and collect the original main shaft stress data through the main shaft stress sensor;

[0074] S12. Perform preprocessing on the original main shaft stress data to obtain the main shaft stress data, and the preprocessing includes data filtering and normalization processing;

[0075] S13. Calculate the main shaft power by combining the main shaft stress data with the torque angular velocity method;

[0076] S14. Perform pre-classification on the main shaft working conditions of the water turbine based on the main shaft power, and the pre-classification includes dividing the main shaft power into multiple preset power intervals, and dividing different preset power intervals into different pre-classification working conditions.

[0077] In a specific embodiment, the torque angular velocity method calculates the power through the formula P = T·ω, where P is the power, T is the torque, and ω is the angular velocity. According to the calculated main shaft power, it is divided into multiple preset power intervals (such as low power, medium power, and high power), and different power intervals are corresponding to different pre-classification working conditions (such as normal working condition, slightly abnormal working condition, and severely abnormal working condition)

[0078] Specifically, in this embodiment, by installing a stress sensor on the main shaft of the water turbine and performing data filtering and normalization processing, the accuracy and reliability of the collected data are improved. The torque angular velocity method is used to calculate the main shaft power, which can more accurately reflect the actual operating state of the water turbine under different working conditions, ensuring the authenticity and effectiveness of the power data. By dividing the main shaft power into multiple preset power intervals, the operating state of the water turbine can be quickly identified, laying a foundation for accurate working condition classification.

[0079] Step S2 includes:

[0080] S21, constructing a main shaft power data set based on the main shaft power data, where the main shaft power data set contains multiple main shaft power data points;

[0081] S22, calculating the silhouette coefficient of each main shaft power data point, where the silhouette coefficient is used to evaluate the relative position of the main shaft power data point in the clustering;

[0082] S23, calculating the overall silhouette coefficient of the main shaft power data set based on the silhouette coefficient of each main shaft power data point;

[0083] S24, constructing a first objective function of the K-Means algorithm based on the overall silhouette coefficient, where the first objective function is used to optimize the clustering effect;

[0084] S25, constructing a second objective function of the main shaft working condition classification model based on the overall silhouette coefficient, where the second objective function is used to optimize the clustering parameters.

[0085] In a specific embodiment, according to the collected main shaft power data, a main shaft power data set is constructed. This data set contains multiple main shaft power data points, and each data point represents the main shaft power value within a specific time period.

[0086] Calculate the silhouette coefficient of each main shaft power data point. The calculation formula of the silhouette coefficient is:

[0087]

[0088] Among them, s(m) is the silhouette coefficient of the m-th data point, B(m) is the average distance from the m-th data point to the nearest other cluster, and A(m) is the average distance from the m-th data point to other points within the same cluster. The silhouette coefficient is used to evaluate the relative position of each data point in the clustering, and the value closer to 1 indicates a better clustering effect.

[0089] Based on the silhouette coefficient of each main shaft power data point, calculate the overall silhouette coefficient of the main shaft power data set. The overall silhouette coefficient is the average of the silhouette coefficients of all data points, reflecting the clustering effect of the entire data set.

[0090] Construct the first objective function of the K-Means algorithm based on the overall silhouette coefficient, and the form of the first objective function is:

[0091]

[0092] where O1 is the first objective function, Q is the number of data points, and the first objective function aims to maximize the overall silhouette coefficient to optimize the clustering effect.

[0093] Construct the second objective function of the main shaft working condition classification model based on the overall silhouette coefficient, and the form of the second objective function is:

[0094]

[0095] where O2 is the second objective function, C l is the l-th cluster, μ l is the cluster center, d(x, μ l ) is the distance from the data point to the cluster center, and the second objective function is used to optimize the clustering parameters.

[0096] Specifically, in this embodiment, by constructing a main shaft power data set, the collected power data is systematically sorted. By calculating the silhouette coefficient of each data point, the relative position of the data point in the clustering can be effectively evaluated, thereby providing a basis for selecting the optimal number of clusters and optimizing the clustering effect. Calculating the overall silhouette coefficient and constructing the objective function can quantify the clustering effect, ensure that the parameter optimization in the clustering process is based on the actual data performance, and improve the accuracy and reliability of the clustering.

[0097] Step S3 includes:

[0098] S31, optimize the clustering parameters through the second objective function, calculate the overall silhouette coefficient under different numbers of clusters, and select the optimal number of clusters based on the clustering effect.

[0099] S32, calculate the optimal Euclidean distance and the optimal Manhattan distance of the optimal number of clusters as the distance metric standard of the K-Means algorithm.

[0100] S33, input the first optimal number of clusters, the optimal Euclidean distance, and the optimal Manhattan distance into the K-Means algorithm.

[0101] In a specific embodiment, the clustering parameters are optimized through the second objective function. First, set a range of the number of clusters (for example, from 2 to 10), then perform K-Means clustering for each number of clusters, and calculate the corresponding overall silhouette coefficient. By comparing the overall silhouette coefficients under different numbers of clusters, select the number of clusters with the highest silhouette coefficient as the first optimal number of clusters.

[0102] After determining the optimal number of clusters, the optimal Euclidean distance and the optimal Manhattan distance under this number of clusters are further calculated. This can be achieved by statistically analyzing the distances between the sample points of each cluster and the cluster centers, and selecting the distance metric criterion that results in the optimal clustering effect.

[0103] Input the first optimal number of clusters, the optimal Euclidean distance, and the optimal Manhattan distance as parameters into the K-Means algorithm.

[0104] Specifically, in this embodiment, by calculating the silhouette coefficient under different numbers of clusters, the optimal number of clusters can be effectively identified, ensuring that the clustering results can truly reflect the internal structure of the data. By calculating the optimal Euclidean distance and the optimal Manhattan distance, a suitable data metric criterion can be provided for the K-Means algorithm, ensuring the adaptability of the clustering algorithm under different data distributions. Inputting the optimized clustering parameters into the K-Means algorithm ensures the systematicness and effectiveness of the entire classification process.

[0105] Step S4 includes:

[0106] S41, Set the number of individuals and the total number of iterations. The number of individuals is the number of candidate solutions used in the optimization process.

[0107] S42, Initialize the positions of the population. Adopt a population initialization strategy, which includes random initialization, chaotic mapping, and random reverse learning.

[0108] S43, Calculate the fitness value of each individual under the second objective function to evaluate its clustering effect.

[0109] S44, Sort according to the fitness value and select the individual with the highest fitness value as the current optimal solution.

[0110] S45, Execute the optimization iteration process. The optimization iteration process includes an exploration stage, an exploration phase, and an exploitation phase, and update the positions of the individuals to find a better solution for the current optimal solution.

[0111] S46, Increase the population diversity through an individual mutation strategy and update the positions of the individuals.

[0112] S47, After each iteration, determine whether the set number of iterations has been reached. When the set number of iterations has not been reached, return to step S43 to continue the iteration. When the set number of iterations has been reached, output the current optimal position of the optimized individual.

[0113] The calculation formula for random initialization is:

[0114] χ i :x ij =ll j +r·(ulj -ll j ), i = 1, 2, ..., N, j = 1, 2, ..., M;

[0115] Among them, χ i is the position of the i-th candidate solution, x ij is the position coordinate of the i-th candidate solution, ll j and ul j respectively represent the lower and upper bounds of the j-th decision variable, r is a random value between 0 and 1, N is the number of candidate solutions, and M is the number of decision variables;

[0116]

[0117] The calculation formula of the chaotic mapping is:

[0118]

[0119] Among them, X k+1 is the number of individuals after the k-th chaotic mapping, X k is the number of individuals before the k-th chaotic mapping, mod(·) is the modulo function;

[0120] The calculation formula of the stochastic reverse learning is:

[0121]

[0122] Among them, is the reverse solution of X k , U b and L b respectively represent the upper and lower bounds of the problem to be optimized, and J is a random value between 0 and 1.

[0123] The calculation formula of the exploration stage is:

[0124]

[0125] Among them, is the current individual, is the position coordinate of the current individual, D hippo is the position before mutation of the individual with the highest fitness, χ i is the current individual position, is the χ i position update in the exploration stage, I1 is an integer between 1 and 2, y1 is a random number between 0 and 1, and F i is the objective function value.

[0126] The calculation formula of the exploration phase is:

[0127]

[0128] Among them, Predator j is the position of the predator, is a random vector with a modulus between 0 and 1, is the position of the individual when facing the predator, is a random vector following the levy distribution, β is a uniform random number between 2 and 4, c is a uniform random number between 1 and 1.5; d is a uniform random number between 2 and 3, g is a uniform random number between -1 and 1, and are random vectors of dimension 1*m, is the position update of χ in the exploration stage i of.

[0129] The calculation formula for the exploitation stage is:

[0130]

[0131]

[0132] Among them, t is the current iteration number, T is the maximum iteration number, is the new upper bound of the variable decision, is the new lower bound of the variable decision, is the individual position of finding the nearest safe location, r3 is a random number between 0 and 1, ζ1 is the selected random vector or random number, is a random vector with a modulus between 0 and 1, r 12 is a random number of the normal distribution, r 13 is a random number between 0 and 1, is the position update of χ in the exploitation stage i of.

[0133] In a specific embodiment, the number of individuals and the total number of iterations are set. Suppose the number of individuals is 50, indicating that 50 candidate solutions are used in the optimization process; the total number of iterations is set to 100, indicating that the algorithm will find the optimal solution in 100 iterations.

[0134] Initialize the positions of the population. Adopt a random initialization strategy to generate the positions of 50 candidate solutions, and the decision variables of each candidate solution are within the preset lower and upper bounds.

[0135] Calculate the fitness value of each individual under the second objective function. The fitness value is used to evaluate the performance of each candidate solution in the clustering effect, and the higher the fitness value, the better the clustering effect.

[0136] Sort according to the fitness value, and select the individual with the highest fitness value as the current optimal solution. Record the position and fitness value of the current optimal solution.

[0137] Execute the optimization iteration process. This process includes an exploration stage, an exploration phase, and an exploitation phase, updating the positions of individuals to find a better solution for the current optimal solution.

[0138] Increase the population diversity through the individual mutation strategy and update the positions of individuals. The mutation strategy can include methods such as random mutation and crossover to prevent the algorithm from falling into a local optimum.

[0139] After each iteration, determine whether the set number of iterations is reached. If the set number of iterations is not reached, return to step S43 to continue the iteration; if the set number of iterations is reached, output the current optimal position of the optimized individual.

[0140] Specifically, in this embodiment, by reasonably setting the number of individuals and the number of iterations, it is ensured that the optimization process has sufficient search space and time to increase the possibility of finding the global optimal solution. By adopting the random initialization strategy, it can effectively increase the diversity of the population and avoid the concentration of the initial solution, thereby improving the exploration ability of the algorithm. By calculating the fitness value, the clustering effect of each candidate solution can be accurately evaluated. Through the exploration, exploration, and exploitation stages, the algorithm can effectively balance between the global and local. Through the individual mutation strategy, the diversity of the population can be maintained, preventing the algorithm from falling into the local optimal solution and improving the accuracy and reliability of the final clustering result.

[0141] The calculation formula of the individual mutation strategy is:

[0142] G hippo = D hippo [1 + aC(0, σ 2 ) + bG(0, σ 2 )];

[0143]

[0144] where G hippo is the position of the individual with the highest fitness after mutation, D hippo is the position of the individual with the highest fitness before mutation, C(0, σ 2 ) is a Cauchy distribution random variable, G(0, σ 2 ) is a Gaussian distribution random variable, a and b are the Cauchy distribution dynamic parameter and the Gaussian distribution dynamic parameter respectively, is the current individual, f(·) is the fitness function, exp(·) is the exponential function, and σ is the standard deviation.

[0145] In a specific embodiment, during the optimization process, in order to increase the diversity of the population and prevent the algorithm from falling into the local optimal solution, the individual mutation strategy is used to mutate the individual with the highest fitness.

[0146] The Cauchy distribution random variable C(0, σ 2 ) and the Gaussian distribution random variable G(0, σ 2 ) can be generated by a random number generator to ensure the randomness and diversity of each mutation.

[0147] The position G of the mutated individual hippo is used for fitness evaluation and optimization iteration.

[0148] After mutation, the fitness value of the mutated individual is recalculated to evaluate its performance in the clustering effect. If the fitness value of the mutated individual is higher than the current optimal solution, the current optimal solution is updated.

[0149] Specifically, in this embodiment, by mutating the individual with the highest fitness, the diversity of the population can be effectively increased, the algorithm can be prevented from falling into a local optimal solution, thereby improving the global search ability. The mutation strategy enables the algorithm to explore new solution spaces and find better solutions. By using the random variables of the Cauchy distribution and the Gaussian distribution, the mutation strategy has dynamic adaptability and can adjust the mutation amplitude according to the current state of the population. Through appropriate mutation, the convergence process of the algorithm can be accelerated, making the final clustering result more stable and reliable.

[0150] Step S5 includes:

[0151] S51, inputting the first best clustering number, the best Euclidean distance, and the best Manhattan distance into the K-Means algorithm;

[0152] S52, based on the input parameters, performing K-Means clustering iterative calculation on the main shaft power dataset;

[0153] S53, during the clustering iteration process, dynamically updating the clustering center of each data point until the clustering result converges to obtain the second best clustering number;

[0154] S54, according to the second best clustering number, performing final classification on the main shaft operating conditions and outputting the classification result.

[0155] Step S52 includes:

[0156] At the beginning of each iteration, calculate the distance between the current clustering center and each data point, and assign the data point to the nearest clustering center; at the end of each iteration, calculate the mean value of each clustering center to update the position of the clustering center, and record the clustering result and the position change of the clustering center for each iteration.

[0157] Step S54 includes:

[0158] Calculate the silhouette coefficient of each cluster according to the second-best number of clusters to evaluate the clustering effect; compare the clustering result with the preset working condition standard to generate a classification report, indicating the characteristics of each working condition category; when outputting the classification result, provide the sample number and main statistical features of each working condition category, and conduct working condition classification based on the sample number and the main statistical features.

[0159] In a specific embodiment, input the first-best number of clusters, the best Euclidean distance, and the best Manhattan distance into the K-Means algorithm. These parameters are obtained through step S3 to ensure that the K-Means algorithm can operate under the best conditions.

[0160] Based on the input parameters, perform K-Means clustering iterative calculation on the main shaft power dataset. The specific steps are as follows:

[0161] Initialize the cluster centers: randomly select the first-best number of data points as the initial cluster centers.

[0162] The iterative process includes:

[0163] At the beginning of each iteration, calculate the distance between the current cluster centers and each data point (using the best Euclidean distance and the best Manhattan distance).

[0164] Assign each data point to the nearest cluster center.

[0165] At the end of each iteration, calculate the mean of each cluster to update the position of the cluster centers.

[0166] Record the clustering results and the position changes of the cluster centers for each iteration.

[0167] During the clustering iteration process, dynamically update the cluster centers of each data point until the clustering result converges. The convergence criterion can be that the change in the position of the cluster centers is less than a preset threshold, or the maximum number of iterations is reached.

[0168] According to the second-best number of clusters, conduct a final classification of the main shaft working conditions and output the classification result. The output classification result includes:

[0169] The silhouette coefficient of each cluster to evaluate the clustering effect.

[0170] Compare the clustering result with the preset working condition standard to generate a classification report, indicating the characteristics of each working condition category, and provide the sample number and main statistical features of each working condition category.

[0171] Specifically, the steps of comparing the clustering result with the preset working condition standard to generate a classification report, indicating the characteristics of each working condition category, and providing the sample number and main statistical features of each working condition category include the following steps:

[0172] Compare the clustering results with the preset operating condition standards to generate a classification report. The classification report points out the characteristics of each operating condition category, such as the statistical characteristics of the main shaft power, the duration of the operating condition, etc., providing detailed information for the operating condition classification.

[0173] When outputting the classification results, provide the sample quantity and main statistical characteristics of each operating condition category, such as the average value, standard deviation, peak value, etc. These statistical characteristics can be used as the basis for the operating condition classification.

[0174] Based on the sample quantity and main statistical characteristics of each operating condition category, conduct a final classification of the main shaft operating conditions. For example, a power threshold can be set to divide the power data into different operating condition categories; or according to the power fluctuation characteristics, the operating conditions can be divided into types such as stable and fluctuating.

[0175] Finally, output the classification results, including the feature description of each operating condition category, the sample quantity statistics, and the corresponding main shaft power data. These information can provide a basis for the operation monitoring, fault diagnosis, and maintenance decision-making of the water turbine.

[0176] Specifically, in this embodiment, by inputting the first optimal clustering quantity, the optimal Euclidean distance, and the optimal Manhattan distance into the K-Means algorithm, clustering can be performed under optimal conditions to ensure the accuracy and effectiveness of the clustering results. During the clustering iteration process, the clustering center of each data point is dynamically updated, which can better adapt to the changes in the data. By obtaining the second optimal clustering quantity, the best classification of the main shaft operating conditions is carried out.

[0177] At the beginning of each iteration, calculate the distance between the current clustering center and each data point to ensure that the data points are accurately assigned to the nearest clustering center. At the end of each iteration, update the position of the clustering center by calculating the mean value of each clustering center, which can make the clustering center gradually move towards the concentrated area of the data points, enhancing the convergence of the clustering. Record the clustering results and the position changes of the clustering center for each iteration, which helps to analyze the evolution of the clustering process.

[0178] By calculating the silhouette coefficient of each cluster, evaluate the clustering effect. Compare the clustering results with the preset operating condition standards to generate a detailed classification report, which can clearly point out the characteristics of each operating condition category. When outputting the classification results, provide the sample quantity and main statistical characteristics of each operating condition category, which can provide more abundant information support for the operating condition classification. Based on the sample quantity and main statistical characteristics for the operating condition classification, the operating condition information contained in the data can be fully mined, providing reliable data support for the intelligent management and maintenance decision-making of the water turbine.

[0179] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for classifying the operating conditions of a water turbine main shaft based on data analysis, characterized in that, It includes the following steps: S1. Collect the spindle stress data, calculate the spindle power based on the spindle stress data, and pre-classify the spindle working conditions according to the spindle power; S2. Construct a spindle power data set through the spindle power, calculate the overall silhouette coefficient of the spindle power data set, and respectively construct a first objective function of the K-Means algorithm and a second objective function of the spindle working condition classification model based on the overall silhouette coefficient; S3. Select the first optimal number of clusters, the optimal Euclidean distance, and the optimal Manhattan distance through the second objective function; S4. Set the number of individuals and the total number of iterations, and perform iterative training on the second objective function to obtain the optimized individual positions; S5. Input the optimized individual positions into the first objective function, perform clustering iteration on the first objective function based on the first optimal number of clusters, the optimal Euclidean distance, and the optimal Manhattan distance to obtain the second final number of clusters, and perform optimal classification of the spindle working conditions based on the second final number of clusters; Step S2 includes: S21. Construct a spindle power data set according to the spindle power data, and the spindle power data set contains multiple spindle power data points; S22. Calculate the silhouette coefficient of each spindle power data point, and the silhouette coefficient is used to evaluate the relative position of the spindle power data point in the clustering; S23. Calculate the overall silhouette coefficient of the spindle power data set based on the silhouette coefficient of each spindle power data point; S24. Construct a first objective function of the K-Means algorithm based on the overall silhouette coefficient, and the first objective function is used to optimize the clustering effect; S25. Construct a second objective function of the spindle working condition classification model based on the overall silhouette coefficient, and the second objective function is used to optimize the clustering parameters.

2. The method for classifying the operating conditions of a water turbine main shaft based on data analysis according to claim 1, wherein, Step S1 includes: S11. Set the spindle stress sensor on the water turbine spindle, and collect the original spindle stress data through the spindle stress sensor; S12. Perform preprocessing on the original spindle stress data to obtain the spindle stress data, and the preprocessing includes data filtering and normalization processing; S13. Calculate the spindle power by combining the spindle stress data through the torque angular velocity method; S14. Pre-classify the spindle working conditions of the water turbine based on the spindle power, and the pre-classification includes dividing the spindle power into multiple preset power intervals, and dividing different preset power intervals into different pre-classified working conditions.

3. The method for classifying the working conditions of a water turbine main shaft based on data analysis according to claim 2, wherein, Step S3 includes: S31. Optimize the clustering parameters through the second objective function, calculate the overall silhouette coefficient under different numbers of clusters, and select the optimal number of clusters based on the clustering effect; S32. Calculate the optimal Euclidean distance and the optimal Manhattan distance of the optimal number of clusters as the distance metric standard of the K-Means algorithm; S33. Input the first optimal number of clusters, the optimal Euclidean distance, and the optimal Manhattan distance into the K-Means algorithm.

4. The method for classifying the operating conditions of a water turbine main shaft based on data analysis according to claim 3, wherein Step S4 includes: S41. Set the number of individuals and the total number of iterations, and the number of individuals is the number of individuals of the candidate solutions used in the optimization process; S42. Initialize the positions of the population, adopting a population initialization strategy, where the population initialization strategy includes random initialization, chaotic mapping, and random reverse learning; S43. Calculate the fitness value of each individual under the second objective function to evaluate its clustering effect; S44. Sort according to the fitness value and select the individual with the highest fitness value as the current optimal solution; S45. Execute the optimization iteration process, where the optimization iteration process includes an exploration stage, an exploitation stage, and an exploitation phase, and update the positions of the individuals to find a better solution to the current optimal solution; S46. Increase the population diversity through an individual mutation strategy and update the positions of the individuals; S47. After each iteration, determine whether the set number of iterations is reached. When the set number of iterations is not reached, return to step S43 to continue the iteration. When the set number of iterations is reached, output the current optimal optimized individual position.

5. The method for classifying the operating conditions of a turbine main shaft based on data analysis according to claim 4, wherein Step S4 also includes: The calculation formula for the random initialization is: χ i :x i,j =ll j +r·(ul j -ll j ),i=1,2,...,N,j=1,2,...,M; where χ i is the position of the i-th candidate solution, and x i,j is the position coordinate of the i-th candidate solution, ll j and ul j represent the lower and upper bounds of the j-th decision variable respectively, r is a random value between 0 and 1, N is the number of candidate solutions, and M is the number of decision variables; The calculation formula for the chaotic mapping is: Among them, X k+1 is the number of individuals after the k-th chaotic mapping, and X k is the number of individuals before the k-th chaotic mapping, and mod(·) is the modulo function; The calculation formula for the random reverse learning is: Among them, is the reverse solution of X, k U b and L b are the upper and lower boundaries of the problem to be optimized respectively, and J is a random value between 0 and 1.

6. The method for classifying the working conditions of a water turbine main shaft based on data analysis according to claim 5, wherein The calculation formula for the individual mutation strategy is: G hippo = D hippo [1 + aC(0, σ 2 ) + bG(0, σ 2 )]; Among them, G hippo is the position after mutation of the individual with the highest fitness, D hippo is the position of the individual with the highest fitness before mutation, C(0, σ 2 ) is a Cauchy distribution random variable, G(0, σ 2 ) is a Gaussian distribution random variable, a and b are the Cauchy distribution dynamic parameter and the Gaussian distribution dynamic parameter respectively, is the current individual, f(·) is the fitness function, exp(·) is the exponential function, and σ is the standard deviation.

7. The method for classifying the operating conditions of a water turbine main shaft based on data analysis according to claim 6, characterized in that, Step S5 includes: S51. Input the first best clustering number, the best Euclidean distance, and the best Manhattan distance into the K-Means algorithm; S52. Based on the input parameters, perform K-Means clustering iterative calculation on the main shaft power dataset; S53. During the clustering iteration process, dynamically update the clustering centers of each data point until the clustering result converges to obtain the second best clustering number; S54. According to the second best clustering number, perform a final classification on the main shaft operating conditions and output the classification result.

8. A method for classifying the operating conditions of a water turbine main shaft based on data analysis according to claim 7, characterized in that, Step S52 includes: At the beginning of each iteration, calculate the distance between the current clustering center and each data point and assign the data point to the nearest clustering center; at the end of each iteration, calculate the mean of each clustering center to update the position of the clustering center, and record the clustering result and the position change of the clustering center for each iteration.

9. The method for classifying the operating conditions of a water turbine main shaft based on data analysis according to claim 8, wherein Step S54 includes: According to the second best clustering number, calculate the silhouette coefficient of each cluster to evaluate the clustering effect; compare the clustering result with the preset operating condition standard to generate a classification report, indicating the characteristics of each operating condition category; when outputting the classification result, provide the sample number and the main statistical characteristics of each operating condition category, and perform operating condition classification based on the sample number and the main statistical characteristics.

Citation Information

Patent Citations

  • Office building load data analysis method, device and equipment and storage medium

    CN118194071A

  • Pathological data analysis method and apparatus, and device and storage medium

    WO2021135063A1