A wind power generation monitoring data collection and clustering method, medium and system

Through multi-layer lightweight neural networks and data quality assessment matrices, the problem of data collection and clustering separation in wind turbine monitoring is solved, and adaptive sampling strategy optimization is achieved to ensure the accurate collection of key data and efficient use of resources.

CN119884798BActive Publication Date: 2025-09-12BEIJING NANCAL RUIYUAN DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510224376.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-09-12
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

In the existing technology, wind turbine monitoring data collection and clustering are separated, resulting in waste of sampling resources or information loss, making it difficult to achieve differentiated sampling and analysis of data of different importance.

Method used

A multi-layer lightweight neural network is used for feature extraction and fusion, combined with the data quality assessment matrix and acquisition contribution value to achieve adaptive sampling frequency adjustment. The data is divided into acquisition data clusters through component decomposition and multi-layer lightweight clustering methods to determine the optimal sampling time window length.

Benefits of technology

It achieves accurate evaluation and differentiated sampling of different monitoring parameters, ensures accurate sampling of key data clusters and moderate sampling of secondary data clusters, reduces computational complexity, and is suitable for large-scale engineering applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119884798B_ABST
    Figure CN119884798B_ABST
Patent Text Reader

Abstract

The present invention provides a wind power generation monitoring data collection and clustering method, medium and system, which belongs to the technical field of wind power generation monitoring data collection, including: setting an initial sampling frequency based on the operating parameters of the wind turbine generator set to collect and preprocess data, performing component decomposition to obtain stable components and variable components, establishing an acquisition data quality assessment matrix to calculate the acquisition contribution value, using a multi-layer lightweight clustering method to divide the data into acquisition data clusters, the multi-layer lightweight clustering method performs multi-level feature extraction and fusion through multiple lightweight neural networks, determines the optimal sampling time window length based on the acquisition data cluster, updates the sampling interval and applies it to the wind turbine generator set monitoring data collection, solves the technical problem that the existing technology often separates the collection and clustering of wind power generation monitoring data, making it difficult to achieve differentiated sampling and analysis of data of different importance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wind power generation monitoring data acquisition, and in particular relates to a wind power generation monitoring data acquisition clustering method, medium and system. Background Art

[0002] As wind power continues to grow in the energy mix, the importance of monitoring the operating status of wind turbines is becoming increasingly prominent. Currently, wind turbine monitoring primarily involves two steps: data acquisition and data analysis. Data acquisition is responsible for acquiring parameters such as voltage, current, speed, vibration, and temperature, while data analysis processes and analyzes the collected data through methods such as clustering. In existing technologies, these two steps are often implemented separately. Data acquisition primarily utilizes a fixed sampling frequency, sampling various parameters at preset, fixed time intervals. The sampling frequency is typically set based on empirical values ​​or conservative estimates, lacking a feedback mechanism for adjusting the data analysis results. This fixed sampling method is difficult to adapt to the monitoring needs of wind turbines under varying operating conditions. It generates a large amount of redundant data when the turbine is operating steadily, and can miss important information when the turbine's status changes rapidly.

[0003] In terms of data analysis, existing technologies primarily use traditional algorithms such as K-means and hierarchical clustering to cluster collected data and identify the unit's operating modes and status characteristics. However, these clustering methods only process historical data and are unable to guide or optimize the data collection process. Even if it is discovered that certain data clusters require more frequent sampling or that certain data samples are redundant, it is impossible to adjust the sampling strategy in a timely manner.

[0004] This separation of data collection and analysis leads to the following problems: First, data collection cannot be adaptively adjusted based on cluster analysis results, resulting in wasted sampling resources or missing information. Second, cluster analysis cannot fully utilize the temporal characteristics and parameter correlations during data collection, affecting the accuracy of analysis results. Third, the lack of a feedback mechanism makes it difficult to achieve differentiated sampling of data of varying importance, reducing the efficiency of the monitoring system.

[0005] Especially in the monitoring of large wind farms, as the scale of data increases, the problem caused by the separation of collection and analysis becomes more prominent. How to effectively feed back the cluster analysis results to the data collection process and realize intelligent adjustment of the sampling strategy has become a technical problem that needs to be solved urgently in the field of wind power monitoring. Summary of the Invention

[0006] In view of this, the present invention provides a wind power generation monitoring data collection and clustering method, medium and system, which can solve the technical problem that the existing technology often separates the collection and clustering of wind power generation monitoring data, making it difficult to achieve differentiated sampling and analysis of data of different importance.

[0007] The present invention is achieved in that:

[0008] The first aspect of the present invention provides a wind power generation monitoring data collection and clustering method, including: setting an initial sampling frequency based on the operating parameters of the wind turbine generator set to collect voltage data, current data, speed data, vibration data, and temperature data, preprocessing the voltage data, current data, speed data, vibration data, and temperature data to obtain preprocessed data, component decomposing the preprocessed data to obtain stable components and variable components, establishing an acquisition data quality assessment matrix to calculate the acquisition contribution value, and using a multi-layer lightweight clustering method to divide the data into acquisition data clusters. The multi-layer lightweight clustering method performs multi-level feature extraction and fusion through electrical feature lightweight neural networks, mechanical feature lightweight neural networks, environmental feature lightweight neural networks, motor feature lightweight neural networks, mechanism feature lightweight neural networks, and cluster feature lightweight neural networks, determines the optimal sampling time window length based on the acquisition data cluster, updates the sampling interval, and applies it to wind turbine generator set monitoring data collection.

[0009] The collected data includes voltage data, current data, speed data, vibration data, and temperature data, and the preprocessing includes denoising and standardization.

[0010] Among them, the component decomposition obtains voltage stability component, voltage variation component, current stability component, current variation component, speed stability component, speed variation component, vibration stability component, vibration variation component, temperature stability component, and temperature variation component, and establishes wind turbine monitoring data collection evaluation indicators.

[0011] The acquisition data quality assessment matrix includes data integrity, data volatility and data correlation. The acquisition features are evaluated for importance based on the acquisition data quality assessment matrix to calculate the acquisition contribution value.

[0012] Among them, the multi-layer lightweight clustering method includes constructing an electrical feature lightweight neural network, a mechanical feature lightweight neural network, and an environmental feature lightweight neural network. The electrical feature lightweight neural network inputs a voltage stability component, a voltage variation component, a current stability component, and a current variation component.

[0013] Among them, the mechanical feature lightweight neural network inputs the speed stability component, speed variation component, vibration stability component, and vibration variation component, and the environmental feature lightweight neural network inputs the temperature stability component and temperature variation component.

[0014] Among them, the first network output of the electrical feature lightweight neural network and the second network output of the mechanical feature lightweight neural network are jointly input into the motor feature lightweight neural network, and the second network output of the mechanical feature lightweight neural network and the third network output of the environmental feature lightweight neural network are jointly input into the mechanism feature lightweight neural network.

[0015] Among them, the fourth network output of the motor feature lightweight neural network and the fifth network output of the mechanism feature lightweight neural network are jointly input into the cluster feature lightweight neural network, and the collected data cluster is determined according to the sixth network output of the cluster feature lightweight neural network.

[0016] On the basis of the above technical solution, the wind power generation monitoring data collection and clustering method of the present invention can also be improved as follows:

[0017] The multi-layer lightweight clustering method includes:

[0018] Constructing an electrical feature lightweight neural network, inputting the voltage stability component, the voltage variation component, the current stability component, and the current variation component into the electrical feature lightweight neural network; constructing a mechanical feature lightweight neural network, inputting the speed stability component, the speed variation component, the vibration stability component, and the vibration variation component into the mechanical feature lightweight neural network; constructing an environmental feature lightweight neural network, inputting the temperature stability component and the temperature variation component into the environmental feature lightweight neural network;

[0019] The first network output of the electrical feature lightweight neural network and the second network output of the mechanical feature lightweight neural network are jointly input into the motor feature lightweight neural network; the second network output of the mechanical feature lightweight neural network and the third network output of the environmental feature lightweight neural network are jointly input into the mechanism feature lightweight neural network;

[0020] The fourth network output of the motor feature lightweight neural network and the fifth network output of the mechanism feature lightweight neural network are jointly input into the clustering feature lightweight neural network;

[0021] The collected data cluster is determined according to a sixth network output of the clustering feature lightweight neural network.

[0022] Furthermore, the electrical feature lightweight neural network adopts a three-layer perceptron structure, specifically: the number of neurons in the input layer is 4, corresponding to the voltage stability component, voltage variation component, current stability component, and current variation component; the number of neurons in the hidden layer is 8, using the ReLU activation function; the number of neurons in the output layer is 2, using the Sigmoid activation function.

[0023] Furthermore, the mechanical feature lightweight neural network adopts a three-layer perceptron structure, specifically: the number of neurons in the input layer is 4, corresponding to the speed stability component, speed change component, vibration stability component, and vibration change component; the number of neurons in the hidden layer is 8, using the ReLU activation function; the number of neurons in the output layer is 2, using the Sigmoid activation function.

[0024] Furthermore, the environmental feature lightweight neural network adopts a three-layer perceptron structure, specifically: the number of neurons in the input layer is 2, corresponding to the temperature stability component and the temperature change component, the number of neurons in the hidden layer is 4, and the ReLU activation function is adopted; the number of neurons in the output layer is 2, and the Sigmoid activation function is adopted.

[0025] Furthermore, the step of jointly inputting the first network output of the electrical feature lightweight neural network and the second network output of the mechanical feature lightweight neural network into the motor feature lightweight neural network is specifically: splicing the first network output and the second network output into a 4-dimensional vector, and inputting it into the input layer of the motor feature lightweight neural network. The motor feature lightweight neural network adopts a two-layer structure, the number of hidden layer neurons is 6, the ReLU activation function is adopted, and the number of output layer neurons is 2.

[0026] Furthermore, the second network output of the mechanical feature lightweight neural network and the third network output of the environmental feature lightweight neural network are jointly input into the mechanism feature lightweight neural network step, specifically: the second network output and the third network output are spliced ​​into a 4-dimensional vector, and input into the input layer of the mechanism feature lightweight neural network. The mechanism feature lightweight neural network adopts a two-layer structure, the number of hidden layer neurons is 6, the ReLU activation function is adopted, and the number of output layer neurons is 2.

[0027] Furthermore, the fourth network output of the motor feature lightweight neural network and the fifth network output of the mechanism feature lightweight neural network are jointly input into the cluster feature lightweight neural network step, specifically: the fourth network output and the fifth network output are spliced ​​into a 4-dimensional vector, which is input into the input layer of the cluster feature lightweight neural network. The cluster feature lightweight neural network adopts a three-layer structure, with 8 neurons in the first hidden layer and 4 neurons in the second hidden layer. The number of neurons in the output layer is the desired number of clusters. The data collection cluster step is determined based on the sixth network output of the cluster feature lightweight neural network, specifically: the sixth network output is subjected to Softmax normalization processing to obtain the probability distribution of each sample belonging to each cluster, and the samples are divided into the cluster category with the highest probability to form the collection data cluster.

[0028] A second aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores program instructions, and when the program instructions are run in a computer, they are used to execute the above-mentioned wind power generation monitoring data collection and clustering method.

[0029] A third aspect of the present invention provides a wind power generation monitoring data collection and clustering system, which includes the above-mentioned computer-readable storage medium.

[0030] Compared with the prior art, the wind power generation monitoring data collection and clustering method, medium and system provided by the present invention have the following beneficial effects:

[0031] First, this invention overcomes the limitations of traditional methods that separate data collection and clustering, establishing a comprehensive data quality assessment system. By calculating the collection quality score and collection contribution value, it accurately assesses the importance of different monitoring parameters, providing a reliable basis for adaptive adjustment of the sampling frequency. This sampling optimization mechanism based on clustering feedback ensures that the monitoring system can perform differentiated sampling based on the importance of the data.

[0032] Secondly, the present invention processes monitoring data using component decomposition and performs feature extraction and fusion via a multi-layer lightweight neural network. This processing approach not only effectively captures the essential characteristics of the data but also achieves a deep integration of the acquisition process and cluster analysis. By establishing three foundational networks for electrical, mechanical, and environmental characteristics, as well as two intermediate layers for motor and mechanism characteristics, the system can adaptively adjust the sampling strategy for different parameters.

[0033] Of particular note is the innovative way in which clustering results are fed back into the sampling control process. By calculating the temporal and spatial correlations of data clusters and combining them with the acquisition contribution value to determine the optimal sampling window length, this method enables intelligent adjustment of the sampling frequency. The design of a sampling interval that is inversely proportional to the acquisition contribution value ensures accurate sampling of important data clusters and appropriate sampling of less important data clusters.

[0034] Furthermore, the lightweight neural network architecture employed in this paper offers low computational complexity. Through the rational design of the number of network layers and neurons, it achieves real-time coordinated optimization of data collection and clustering. This lightweight design enables the system to rapidly respond to changes in clustering results and adjust sampling strategies in a timely manner, making it particularly suitable for large-scale deployment and application in practical engineering projects.

[0035] In summary, the present invention solves the technical problem that the existing technology often separates the collection and clustering of wind power generation monitoring data, making it difficult to achieve differentiated sampling and analysis of data of different importance. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 A flow chart of the method provided by the present invention;

[0037] Figure 2 Schematic diagram of voltage data components after EMD decomposition in Example 2;

[0038] Figure 3 Schematic diagram of the temporal and spatial correlation variation trends of different data clusters in Example 2;

[0039] Figure 4 Schematic diagram of the optimal sampling time window length for each data cluster in Example 2;

[0040] Figure 5 This is a dynamic change diagram of the sampling frequency of each parameter under extreme weather conditions in Example 2. DETAILED DESCRIPTION

[0041] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0042] like Figure 1 FIG. 1 is a flow chart of a wind power generation monitoring data collection and clustering method provided by the first aspect of the present invention. The method comprises the following steps:

[0043] S01. Set the initial sampling frequency according to the operating parameters of the wind turbine generator set, and collect voltage data, current data, speed data, vibration data, and temperature data during the operation of the wind turbine generator set in real time;

[0044] S02. Preprocess the voltage data, current data, speed data, vibration data, and temperature data, including denoising and standardization, to obtain preprocessed data;

[0045] S03. Decompose the preprocessed data into components to obtain voltage stable component and voltage variation component of voltage data, current stable component and current variation component of current data, speed stable component and speed variation component of speed data, vibration stable component and vibration variation component of vibration data, and temperature stable component and temperature variation component of temperature data, and establish wind turbine generator monitoring data acquisition evaluation indicators;

[0046] S04. Calculate the collection quality scores of voltage data, current data, speed data, vibration data, and temperature data based on the wind turbine generator monitoring data collection evaluation indicators, and establish a collection data quality assessment matrix. The collection data quality assessment matrix includes data integrity, data volatility, and data correlation;

[0047] S05. Based on the data collection quality assessment matrix, evaluate the importance of the collection features of voltage data, current data, speed data, vibration data, and temperature data, and calculate the collection contribution values ​​of voltage data, current data, speed data, vibration data, and temperature data to the wind turbine operating status monitoring;

[0048] S06. Weighting the voltage data, current data, speed data, vibration data, and temperature data according to the collected contribution values, and dividing the weighted voltage data, current data, speed data, vibration data, and temperature data into a plurality of collected data clusters using a multi-layer lightweight clustering method;

[0049] S07. Based on the collected data quality assessment matrix, for each collected data cluster, calculate the temporal correlation and spatial correlation of the data within the collected data cluster;

[0050] S08. Determine the optimal sampling time window length corresponding to the collected data cluster based on the time correlation and space correlation and the collection contribution value;

[0051] S09. Adaptively adjust the sampling frequency of the voltage data, current data, speed data, vibration data, and temperature data in the collected data cluster based on the optimal sampling time window length, and update the sampling interval of the voltage data, current data, speed data, vibration data, and temperature data. The sampling interval is inversely proportional to the collection contribution value.

[0052] S10. Applying the sampling interval to wind turbine generator set monitoring data collection to achieve cluster adaptive adjustment of wind turbine generator set monitoring data collection.

[0053] The specific implementation of the above steps is described in detail below:

[0054] In step S01, the initialization process for wind turbine monitoring data collection first requires determining an appropriate sampling frequency based on the actual operating conditions of the wind turbine. During wind turbine operation, wind speed and power output are two important operating parameters, and changes in these parameters directly affect the turbine's operating status. When the wind speed is relatively stable, the turbine's operating parameters are also relatively stable, and a relatively low sampling frequency can be used for data collection. However, when the wind speed fluctuates dramatically, the turbine's operating status can change rapidly, requiring a higher sampling frequency to accurately capture these changes. Similarly, when the turbine's power output fluctuates significantly, the sampling frequency should also be increased accordingly. This ensures that sufficient data information is obtained under different operating conditions. After determining the initial sampling frequency, the system will collect various data during wind turbine operation in real time, including voltage, current, speed, vibration, and temperature data. This data reflects different aspects of the turbine's operation and is an important basis for condition monitoring and fault diagnosis.

[0055] In step S02, to ensure the accuracy of subsequent analysis, the collected raw data must be preprocessed. The preprocessing process mainly includes two steps: denoising and standardization. During the actual data collection process, due to the influence of various factors, the collected data often contains some noise signals. This noise may come from environmental interference, errors in the sensor itself, interference during signal transmission, and other aspects. If this noise is not processed, it will affect the accuracy of subsequent analysis. The denoising process can adopt various methods. For example, the median filter method can effectively remove sudden impulse noise, while the wavelet denoising method can effectively remove background noise while retaining the important characteristics of the signal. After completing the denoising process, the data must be standardized. This is because different types of data may have different dimensions and numerical ranges. Directly using the raw data for analysis may cause certain parameters with larger values ​​to dominate the analysis process, obscuring the influence of other important parameters. Through standardization, all data can be unified into the same scale range, eliminating the influence of dimensions, making the subsequent analysis more objective and accurate.

[0056] In step S03, component decomposition of the preprocessed data is a key processing step. The main purpose of this step is to decompose each type of data into stable and varying components. This decomposition can better reflect the essential characteristics of the data. Taking the empirical mode decomposition method as an example, this method can adaptively decompose nonlinear, non-stationary signals into a series of components with different characteristics. During the decomposition process, all local extreme points in the signal are first identified. Then, upper and lower envelopes are constructed from these extreme points. The average of the upper and lower envelopes forms an average envelope. This average envelope is subtracted from the original signal to obtain a new signal. If this new signal meets the conditions for an intrinsic mode function (IMF), it is considered an IMF. If not, the above process is repeated until an IMF is obtained. This gradual decomposition ultimately yields a series of IMFs and a residual signal. Among these decomposed components, low-frequency components generally reflect the overall trend of the signal and can be considered stable components, while high-frequency components reflect the local fluctuation characteristics of the signal and can be considered varying components. Through this decomposition, the stable and variable components of voltage, current, speed, vibration, and temperature data can be obtained. These components provide an important basis for the subsequent establishment of monitoring data collection and evaluation indicators.

[0057] In step S04, a comprehensive data quality assessment system is established based on the stable and variable components of each type of data obtained in the previous steps. This assessment system primarily evaluates data quality based on three dimensions: data integrity, data volatility, and data correlation. Data integrity focuses on the relationship between the number of data points actually collected and the theoretical number of data points collected. This metric can reflect whether there were any data loss or interruptions during the data collection process. In practical applications, data collection may be interrupted or lost due to various reasons. Therefore, assessing data integrity is crucial for ensuring data quality. Data volatility focuses on the degree of data dispersion and can reflect data stability. If the data volatility of a parameter is large, it indicates that the parameter may have undergone significant changes during the monitoring period. Such changes may be related to the operating status of the equipment and require special attention. Data correlation focuses on the degree of correlation between data points and can reflect the continuity and reliability of the data. There should be a certain degree of correlation between adjacent data points. If the correlation is too low, it may indicate anomalies in the data collection process. These three dimensions of indicators together form a data quality assessment matrix, which can be used to comprehensively assess the quality of the data for each monitoring parameter.

[0058] In step S05, the importance of each monitoring parameter is evaluated based on the data quality assessment matrix, and the contribution value of each parameter to wind turbine operating status monitoring is calculated. The core of this step is to comprehensively determine the importance of each parameter in the overall monitoring system by evaluating its performance across different quality dimensions. In practical applications, different monitoring parameters have different importance in reflecting the turbine's operating status. For example, changes in some parameters may directly reflect the turbine's key performance indicators, while other parameters may only serve as auxiliary monitoring indicators. Therefore, it is necessary to comprehensively evaluate the importance of each parameter based on its performance across three dimensions: data integrity, data volatility, and data correlation, combined with the operational characteristics of the wind turbine. This assessment should not only consider the characteristics of the parameters themselves, but also the relationships between them. For example, some parameters may have strong correlations, in which case the importance weights of some parameters should be appropriately reduced to avoid duplication of information. The collected contribution values ​​calculated in this way can truly reflect the importance of each parameter in the monitoring process.

[0059] In step S06, classifying the weighted data using a multi-layer lightweight clustering method is a complex process. This process first requires the construction of three basic lightweight neural networks, one for processing electrical features, one for processing mechanical features, and one for processing environmental features. The design of these three networks all adopts a simple three-layer perceptron structure. By reasonably setting the number of neurons in the input layer, hidden layer, and output layer, the key information of each type of feature can be effectively extracted. On this basis, two intermediate lightweight neural networks are used to extract motor features and mechanism features. Finally, a clustering feature lightweight neural network is used to synthesize the previously extracted features to obtain the final clustering results. This multi-level feature extraction and fusion method can fully consider the correlation between different types of data, thereby obtaining more reasonable data classification results. At the same time, due to the use of a lightweight network structure, the computational burden of the entire processing process is relatively small, making it suitable for application in practical engineering.

[0060] In step S07, for each data cluster obtained by multi-layer lightweight clustering, an in-depth correlation analysis needs to be performed. This analysis process mainly includes two aspects: time correlation analysis and spatial correlation analysis. Time correlation analysis focuses on the degree of correlation between data points in the time series within the same data cluster. This correlation may be manifested as periodic changes, trend changes or other types of time dependence of the data. By analyzing these time correlations, the changing pattern of data within the data cluster can be understood, providing a basis for subsequent sampling frequency adjustment. Spatial correlation analysis focuses on the degree of correlation between different monitoring points within the same data cluster. Since a wind turbine is a complex electromechanical system, there may be complex correlations between the monitoring data of different parts. By analyzing these spatial correlations, the mutual influence between the various parts of the system can be better understood, which helps to optimize the data acquisition strategy.

[0061] In step S08, determining the optimal sampling time window length is a process that requires comprehensive consideration of multiple factors. This process needs to consider the time correlation, spatial correlation and collection contribution value of each parameter of the data cluster at the same time. When a data cluster shows a strong time correlation, it means that the data in the cluster has good time continuity. In this case, the length of the sampling time window can be appropriately increased and the sampling frequency can be reduced. On the contrary, if the time correlation is weak, the sampling time window needs to be shortened and the sampling frequency needs to be increased to better capture the changes in the data. Similarly, spatial correlation will also affect the selection of the optimal sampling time window. Strong spatial correlation indicates that there is a close relationship between different monitoring points. In this case, the sampling strategy can be optimized by collaboratively considering the data of multiple monitoring points. In addition, the collection contribution value of each parameter needs to be considered. For parameters with higher contribution values, a higher sampling frequency should be guaranteed to ensure that sufficient effective information is obtained.

[0062] In step S09, the sampling frequency of the monitoring parameters in each data cluster is adaptively adjusted based on the determined optimal sampling time window length. The core principle of this adjustment process is that the sampling frequency is inversely proportional to the acquisition contribution value. For parameters with higher acquisition contribution values, it means that these parameters are more important to the monitoring system, and a higher sampling frequency needs to be maintained to capture their detailed change information. For parameters with lower acquisition contribution values, the sampling frequency can be appropriately reduced, which can reduce the burden of data storage and processing without significantly affecting the overall performance of the monitoring system. This adaptive adjustment mechanism enables the system to optimize resource utilization efficiency while ensuring monitoring results. In addition, the adjustment of the sampling frequency also needs to take into account actual engineering constraints, including sensor performance limitations, data transmission bandwidth limitations, etc., to ensure that the adjusted sampling frequency is feasible.

[0063] In step S10, the adaptive sampling strategy determined above is applied to the actual wind turbine monitoring data collection process. This is a dynamic process, and the system needs to continuously monitor the changes in the characteristics of each data cluster and adjust the sampling strategy in a timely manner according to these changes. During actual operation, the working state of the wind turbine may change, and these changes will be reflected in the changes in the characteristics of the monitoring data. The system needs to be able to identify these changes in a timely manner and adjust the sampling strategy accordingly. This dynamic adjustment mechanism ensures that the monitoring system can always maintain the optimal data collection effect and provide reliable data support for the status monitoring and fault diagnosis of the wind turbine. At the same time, this clustering-based adaptive adjustment method also has good scalability, and new monitoring parameters can be added or evaluation indicators can be adjusted according to actual needs.

[0064] The calculation process involved in the present invention is described in detail below:

[0065] 1. The component decomposition calculation adopts the empirical mode decomposition method to decompose the signal x(t), which is specifically expressed as follows:

[0066]

[0067] Where c i (t) is the eigenmode function; r n (t) is the residual term; n is the number of decomposition levels; t is the sampling time.

[0068] 2. The calculation process of the acquisition data quality assessment matrix Q is as follows:

[0069]

[0070] in:

[0071]

[0072] Where, v j is the sampling data value; N is the actual number of sampling points; N t is the theoretical number of sampling points; is the data average, and i ranges from 1 to 5.

[0073] 3. The calculation process of the collection contribution value is as follows:

[0074]

[0075] Where C i is the collection contribution value of the i-th type of data; α1, α2, α3 are weight coefficients, and satisfy q imax is the maximum value of column i.

[0076] 4. Forward propagation calculation of electrical characteristics lightweight neural network:

[0077] h1=f(W1x+b1);

[0078] y1=g(W2h1+b2);

[0079] Where x is the input vector; W1, W2 are weight matrices; b1, b2 are bias vectors; f(·) is the ReLU activation function; g(·) is the Sigmoid activation function; h1 is the hidden layer output; and y1 is the network output.

[0080] 5. The structure and calculation method of the mechanical feature network and the environmental feature network are the same, the difference lies in the input dimension.

[0081] 6. The feature fusion process is calculated as follows:

[0082] z = [y1; y2];

[0083] h f =f(W f z+b f );

[0084] y f =g(W o z+b o );

[0085] Where z is the splicing vector; y1, y2 are the features to be fused; W f ,W o is the fusion layer weight; b f ,b o is the fusion layer bias; h f is the fusion hidden layer output; y f The fusion output.

[0086] 7. The final clustering process uses Softmax to calculate the category probability:

[0087]

[0088] Where p k is the probability that the sample belongs to the kth class; y k is the output value of the network for the kth class; K is the total number of cluster categories.

[0089] In this plan:

[0090] 1. Component decomposition using EMD method can effectively process nonlinear and non-stationary signals;

[0091] 2. The quality assessment matrix comprehensively considers the three dimensions of completeness, volatility and correlation;

[0092] 3. The contribution value is collected and weight coefficients are introduced to achieve adaptive balance in different dimensions;

[0093] 4. The neural network structure adopts a lightweight design to reduce computational complexity;

[0094] 5. Feature fusion adopts a multi-level fusion strategy to retain feature information at different levels.

[0095] A second aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores program instructions, and when the program instructions are run in a computer, they are used to execute the above-mentioned wind power generation monitoring data collection and clustering method.

[0096] A third aspect of the present invention provides a wind power generation monitoring data collection and clustering system, which includes the above-mentioned computer-readable storage medium.

[0097] Existing methods for clustering wind power generation monitoring data collection suffer from fixed data collection strategies, limited feature extraction capabilities, and insufficient clustering algorithm performance. In terms of data collection, traditional methods generally use a fixed sampling frequency for data acquisition, which is unable to adaptively adjust the data characteristics of wind turbines under different operating conditions. Due to the significant differences in the changing characteristics of various parameters during wind turbine operation, fixed sampling methods cannot balance the relationship between data value and collection costs, often resulting in insufficient data collection under critical conditions and a large amount of redundant data under stable conditions. In terms of feature extraction, existing technologies often process different types of monitoring data, such as electrical characteristics, mechanical characteristics, and environmental characteristics, independently, ignoring the coupling relationship and physical connection between them. As a result, the extracted features cannot fully reflect the system status. In terms of clustering algorithms, traditional methods such as K-means and hierarchical clustering have high requirements for data distribution, and the algorithm complexity increases significantly with the data size, making it difficult to meet real-time processing requirements.

[0098] The wind power generation monitoring data acquisition clustering method proposed in the present invention achieves a good balance between data processing capabilities and computing efficiency through a multi-layer lightweight neural network. First, the adoption of a lightweight network structure reduces the complexity of the model. By streamlining the number of network layers and neurons, the computational load is significantly reduced while ensuring the feature extraction capability. This lightweight design enables the model to adapt to the edge computing environment and improves the practicality of the system. Secondly, the hierarchical feature extraction architecture constructed based on physical mechanisms fully considers the correlation between different types of data. The electrical feature network focuses on extracting the characteristic patterns of electrical quantities such as voltage and current, the mechanical feature network focuses on capturing the dynamic characteristics of mechanical parameters such as speed and vibration, and the environmental feature network focuses on the influence of environmental factors such as temperature. This network design based on domain knowledge improves the interpretability and physical meaning of the features.

[0099] At the feature fusion level, the present invention realizes multi-level feature fusion through the motor feature network and the mechanism feature network. The motor feature network combines electrical features with mechanical features, reflecting the coupling relationship between the electromagnetic characteristics and mechanical characteristics of the motor; the mechanism feature network fuses mechanical features with environmental features, reflecting the impact of environmental factors on the mechanical system. This progressive feature fusion strategy can gradually extract higher-level system features and provide more effective feature expression for subsequent clustering. The final clustering feature network realizes intelligent classification of monitoring data by comprehensively analyzing the features of each layer, providing a basis for optimizing the sampling strategy.

[0100] Another innovation of this invention lies in its adaptive sampling mechanism driven by acquisition contribution values. By establishing a data quality assessment matrix encompassing integrity, volatility, and correlation, the system comprehensively assesses the importance of different types of data. The acquisition contribution values ​​calculated based on these assessments directly guide the dynamic adjustment of sampling parameters, achieving optimal allocation of sampling resources. This data-driven sampling strategy enables the system to ensure the collection of critical information while avoiding the generation of redundant data.

[0101] Compared to existing technologies, the lightweight neural network clustering method presented in this paper offers significant theoretical advantages. Its lightweight design reduces model complexity, making the system more suitable for engineering practice. Its hierarchical structure, based on physical mechanisms, enhances feature interpretability. Multi-level feature fusion fully exploits inter-data relationships. Its adaptive sampling mechanism enables intelligent allocation of data collection resources. These innovations give this method significant theoretical value and promising application prospects in the field of wind power monitoring data collection.

[0102] Furthermore, this invention improves the system's adaptability and maintainability by optimizing the network structure and computational process. The modular design facilitates system expansion and upgrades, while the simplified network structure reduces the difficulty of debugging and maintenance. These features enable this invention to better meet the needs of engineering practice and provide a new technical path for the intelligent upgrade of wind farm monitoring systems.

[0103] Specifically, the principles of this invention are as follows: First, at the data processing level, the invention uses component decomposition to process monitoring data, breaking down complex time series signals into stable and variable components. This decomposition method, based on signal processing theory, effectively separates the basic and dynamic characteristics of the data. The stable component reflects the basic operating state of the system and is suitable for a lower sampling frequency; the variable component, on the other hand, contains the dynamic characteristics of the system and requires a higher sampling frequency to capture state changes.

[0104] Secondly, the present invention designs a multi-layer lightweight neural network structure, including three basic feature networks and three fusion networks. This hierarchical network structure design follows the principle of "divide and conquer" and achieves the coordinated optimization of collection and clustering through layer-by-layer feature extraction and fusion. The basic feature network adopts a three-layer perceptron structure, provides nonlinear mapping capabilities through the ReLU activation function, and maps features to a fixed interval through the Sigmoid activation function.

[0105] At the feature fusion level, the motor feature network and the mechanism feature network combine features at different levels to establish correlations between parameters. This feature fusion approach aligns with the physical operating principles of wind turbines and accurately reflects the impact of different parameters on system status. The resulting clustering feature network achieves deep feature fusion through a multi-layered structure and achieves reliable clustering results through Softmax normalization.

[0106] In terms of sampling optimization, this paper designs an adaptive sampling mechanism based on cluster feedback, drawing on information theory and control theory. By calculating the temporal and spatial correlations of data clusters, the system can assess the importance and changing patterns of different data. The design of a sampling interval that is inversely proportional to the acquisition contribution ensures the rational allocation of sampling resources. This design not only complies with the Nyquist sampling theorem but also enables optimization based on the actual importance of the data.

[0107] In summary, this invention achieves collaborative optimization of data collection and clustering through a multi-layer lightweight neural network. Its principles are well-founded and its logic is rigorous. The system can adaptively adjust the sampling strategy based on the clustering results. Cluster analysis also fully utilizes the temporal characteristics of the sampling process, forming a virtuous cycle in which data collection and clustering mutually promote and optimize each other. This design approach not only solves the problem of the separation between data collection and clustering in traditional methods, but also provides a scalable and implementable engineering solution.

[0108] A specific embodiment 1 of the present invention is provided below. The specific implementation of each step in this embodiment 1 is described in detail as follows: The specific implementation of step S01 is to first determine the initial sampling frequency based on the operating parameters of the wind turbine generator set, such as wind speed, power output, etc. The purpose of this step is to ensure the real-time collection of important parameters during the operation of the unit. Specifically, different sampling frequencies will be set according to different operating conditions. For example, when the wind speed changes drastically or the power output fluctuates greatly, the sampling frequency will be appropriately increased to ensure that the time resolution of the data is within 0.5 seconds. At the same time, key parameter data such as voltage, current, speed, vibration and temperature during the operation of the unit will also be collected. The collection of these parameter data can provide basic data support for the subsequent monitoring and diagnosis of the unit's operating status.

[0109] The specific implementation method of step S02 is to first pre-process the collected voltage, current, speed, vibration and temperature data. The pre-processing includes two steps: denoising and standardization. The purpose of denoising is to eliminate the noise components introduced by factors such as environmental interference and instrument accuracy, and retain the essential characteristics of the signal. The purpose of standardization is to unify the original parameter data of different dimensions into the same dimension, eliminate the influence of the dimension, and facilitate subsequent feature extraction and analysis. Denoising can use median filtering, in which the filter window length is set to 5; standardization uses z-score standardization, that is, the data is subtracted from the mean and divided by the standard deviation. The purpose of this step is to ensure that subsequent data analysis is more accurate and reliable.

[0110] The specific implementation method of step S03 is to perform component decomposition on the pre-processed voltage, current, speed, vibration and temperature data. The purpose of component decomposition is to extract the stable component and the variable component in each parameter as the basic indicator for subsequent data evaluation. The specific component decomposition method is to use the empirical mode decomposition (EMD) algorithm. The EMD algorithm can adaptively decompose the nonlinear and non-stationary original signal into a series of intrinsic mode functions (IMFs), each IMF representing an oscillation component on a different time scale. By performing EMD decomposition on the voltage, current, speed, vibration and temperature data respectively, the respective stable components (composed of low-frequency IMFs) and variable components (composed of high-frequency IMFs) can be obtained. The mathematical expression is as follows:

[0111]

[0112] Among them, c i (t) is the eigenmode function, r n (t) is the residual term, n is the number of decomposition levels, and t is the sampling time. These components reflect the stability and dynamic characteristics of the key parameters of the unit and provide an important basis for subsequent data quality evaluation.

[0113] The specific implementation of step S04 is to construct a 5×3 data quality assessment matrix Q based on the stable and variable components of voltage, current, speed, vibration and temperature extracted in step S03. Each element of the matrix represents a quality characteristic of a certain parameter, including: data integrity (q i1 ), data volatility (q i2 ) and data correlation (q i3 ). The calculation formula is as follows:

[0114]

[0115] Among them, v j is the sampling data value, N is the actual number of sampling points, N t is the theoretical number of sampling points, is the data average. Data integrity reflects the ratio of the actual number of sampling points to the theoretical number of sampling points, characterizing the completeness of the data. Data volatility reflects the degree of data dispersion and characterizes data stability. Data correlation reflects the correlation between adjacent data points and characterizes the temporal continuity of the data. By comprehensively considering these three dimensions, the quality of each parameter data can be comprehensively assessed.

[0116] The specific implementation of step S05 is to calculate the contribution value C of each parameter data (voltage, current, speed, vibration, temperature) to the operation status monitoring of the wind turbine generator set based on the data quality evaluation matrix Q constructed in step S04. i This contribution value is obtained by weighted average of the three dimensions of the Q matrix. The weight coefficients α1, α2 and α3 can be adjusted according to the actual application requirements and meet the requirements of Contribution value C i It reflects the relative importance of the parameter data in the overall monitoring system and provides a basis for subsequent adaptive sampling. The calculation formula is:

[0117]

[0118] Among them, q imax Indicates the maximum value of the index in column i. According to actual application requirements, α1 = 0.4, α2 = 0.3, and α3 = 0.3 can be set.

[0119] The specific implementation of step S06 is to first calculate the collection contribution value C of each parameter data obtained in step S05. i , weighting the original data. Then a multi-level lightweight clustering method is used to divide the weighted voltage, current, speed, vibration and temperature data into several collected data clusters. This multi-level clustering method includes:

[0120] 1) Construct a lightweight neural network for electrical characteristics and input voltage stability component v s 、Voltage variation component v d , current stability component i s and the current variation component i d , using a three-layer perceptron structure, the number of hidden layer neurons is 8, and the output is a 2-dimensional electrical feature vector y1. The forward propagation calculation is as follows:

[0121] h1=f(W1x+b1);

[0122] y1=g(W2h1+b2);

[0123] Among them, x is the input vector, W1, W2 are weight matrices, b1, b2 are bias vectors, f(·) is the ReLU activation function, g(·) is the Sigmoid activation function, and h1 is the hidden layer output.

[0124] 2) Construct a lightweight neural network based on mechanical characteristics and input the speed stability component n s , speed variation component n d , vibration stability component v b,s and the vibration variation component v b,d ,The structure and calculation method are the same as above, and the output is a 2D mechanical feature vector y2.

[0125] 3) Construct a lightweight neural network based on environmental characteristics and input the temperature stability component T s and temperature variation component T d ,The structure and calculation method are the same as above, and the output is a 2D environmental feature vector y3.

[0126] 4) Concatenate y1 and y2 and input them into the motor feature lightweight network, which adopts a two-layer structure with 6 hidden layer neurons and outputs a 2D motor feature vector y4.

[0127] 5) Concatenate y2 and y3 and input them into the lightweight mechanism feature network, which adopts a two-layer structure with 6 hidden layer neurons and outputs a 2-dimensional mechanism feature vector y5.

[0128] 6) Concatenate y4 and y5 and input them into the clustering feature lightweight network, which adopts a three-layer structure. The number of neurons in the first hidden layer is 8, the number of neurons in the second hidden layer is 4, the number of neurons in the output layer is the required number of clusters K, and the clustering result y6 is output.

[0129] This multi-level feature fusion method can fully explore the intrinsic connections between different measurement parameters and extract more discriminative clustering features, thereby obtaining more accurate data collection cluster division results.

[0130] The specific implementation of step S07 is to calculate the time correlation and spatial correlation of the data within each cluster of collected data obtained in step S06. t reflects the temporal dependence and spatial correlation ρ between data points within the same cluster. s It reflects the degree of correlation between different sensor measurement points within the same cluster. These two indicators can measure the overall quality characteristics of the collected data cluster and provide a basis for the subsequent adaptive adjustment of the sampling frequency.

[0131] The specific implementation of step S08 is to calculate the time correlation ρ obtained in step S07. t and spatial correlation ρ s, combined with the collection contribution value C of each parameter data calculated in step S05 i , determine the optimal sampling time window length T corresponding to each collected data cluster win The specific method is, when ρ t and ρ s When it is higher, it means that the data has strong continuity and correlation, and the sampling frequency can be appropriately reduced, that is, T is increased. win On the contrary, it is necessary to increase the sampling frequency to capture more dynamic change information and reduce T win At the same time, the contribution value C i The higher the parameter, the higher the sampling frequency should be. win By comprehensively considering these factors, the optimal sampling time window length of each data cluster is finally determined, where T win The value range is [0.5,10] seconds.

[0132] The specific implementation of step S09 is to determine the optimal sampling time window length T according to step S08. win , the sampling frequency of voltage, current, speed, vibration and temperature data in each data cluster is adaptively adjusted. Specifically, the sampling frequency f s and the collection contribution value C i Inversely proportional, that is, the higher the contribution value of the parameter, the higher the sampling frequency; the lower the contribution value of the parameter, the lower the sampling frequency can be appropriately reduced. The calculation formula is:

[0133]

[0134] This adaptive adjustment mechanism can ensure the high-time resolution acquisition of key parameter data while also taking into account the acquisition efficiency of secondary parameters, thereby optimizing resource utilization.

[0135] The specific implementation of step S10 is to apply the adaptive sampling frequency determined in step S09 to the actual wind turbine monitoring data collection process. During operation, the system will monitor the time correlation ρ of each collected data cluster in real time. t and spatial correlation ρ s , and dynamically adjust the sampling frequency f accordingly s ,ensures high quality collection of key parameter data while taking into account the overall collection efficiency.,This adaptive clustering adjustment method can meet the data collection requirements under ,complex operating environments of wind turbines, and provide reliable basic data support for ,subsequent state monitoring and fault diagnosis.

[0136] To better understand and implement the present invention, Example 2 of a specific application scenario is provided below: A wind farm consists of 100 3-MW wind turbines, with a total installed capacity of 300 MW. This wind farm is located in a remote area, subject to harsh environmental conditions, and frequently exposed to wind, cold temperatures, ice, and snow, presenting significant challenges for monitoring the turbines' operation. To ensure the safe and stable operation of the turbines, operations and maintenance personnel urgently require an intelligent monitoring data collection method that can adapt to complex environments.

[0137] Based on the proposed clustering method for wind power generation monitoring data collection, the wind farm's operations and maintenance team, after thorough demonstration and practical validation, decided to deploy this monitoring system throughout the entire farm. First, they determined the initial sampling frequency for monitoring parameters based on the actual operating characteristics of the wind turbines. Because the wind farm's environment is complex and volatile, with the potential for dramatic changes in wind speed and power output, to ensure high temporal resolution of key parameter data, the initial sampling frequency was set to 0.5 Hz, or two samples per second. Simultaneously, the monitoring system collects real-time data on five key parameters: voltage, current, speed, vibration, and temperature, for each of the 100 turbines.

[0138] Next, the operations and maintenance personnel preprocessed the collected raw monitoring data. First, they used a median filter to remove noise, with a filter window length of 5. Then, they used z-score normalization to normalize the parameter data of different dimensions to the range [-1, 1], eliminating the impact of dimensional differences. After these two preprocessing steps, the quality of the monitoring data was initially improved.

[0139] Next, the operation and maintenance personnel perform EMD component decomposition on the pre-processed monitoring data. Taking voltage data as an example, after EMD decomposition, the stable component v of the voltage can be obtained. s and the variable component v d Similarly, EMD decomposition was performed on the other four parameter data to extract the corresponding stable components and variable components, as shown in Table 1:

[0140] Table 1 Monitoring parameter components after EMD decomposition

[0141] parameter Stable component Variable component Voltage <![CDATA[v s ]]> <![CDATA[v d ]]> Current <![CDATA[i s ]]> <![CDATA[i d ]]> Speed <![CDATA[n s ]]> <![CDATA[n d ]]> vibration <![CDATA[v b,s ]]> <![CDATA[v b,d ]]> temperature <![CDATA[T s ]]> <![CDATA[T d ]]>

[0142] With these component data, the operation and maintenance personnel then constructed a 5×3 data quality assessment matrix Q. The specific calculation process is as follows:

[0143]

[0144] Among them, v j is the data value of the jth sampling point, N is the actual number of sampling points, N t is the theoretical number of sampling points, is the average value of the data. Figure 2 As shown in Figure 1, the voltage data components after EMD decomposition are displayed, including the time domain waveforms of the original voltage signal, stable component, and variable component.

[0145] Through calculation, the data quality assessment matrix Q of the 100 wind farm units is obtained, as shown in Table 2:

[0146] Table 2 Data quality assessment matrix Q;

[0147] parameter Data integrity Data volatility Data Correlation Voltage 0.98 0.12 0.85 Current 0.95 0.15 0.82 Speed 0.92 0.18 0.79 vibration 0.90 0.22 0.76 temperature 0.88 0.25 0.72

[0148] As can be seen from Table 2, the overall monitoring data quality of the wind farm is good and the data integrity is high, but there is still room for improvement in data volatility and correlation.

[0149] Based on this evaluation matrix, the operation and maintenance personnel then calculated the collection contribution value C of each parameter data. i , specifically as follows:

[0150]

[0151] Among them, α1 = 0.4, α2 = 0.3, and α3 = 0.3. According to the calculation results, the collection contribution values ​​of 100 units are shown in Table 3:

[0152] Table 3 Collection contribution value C of each parameter data i ;

[0153] parameter <![CDATA[Collection contribution value C i > Voltage 0.92 Current 0.87 Speed 0.82 vibration 0.76 temperature 0.70

[0154] As can be seen from Table 3, voltage and current data are more important in the entire monitoring system and contribute significantly to their acquisition, while vibration and temperature data are relatively less important. This lays the foundation for subsequent adaptive sampling frequency adjustment.

[0155] Next, the operation and maintenance personnel used the multi-level lightweight neural network clustering method proposed in this invention to classify the monitoring data. Specifically, three lightweight neural networks were first constructed for electrical characteristics, mechanical characteristics, and environmental characteristics, respectively, to extract the feature vectors of different types of parameters. The input of the electrical characteristic network is v s 、v d 、i s and i d , a three-layer perceptron structure is used, the number of hidden layer neurons is 8, and the output is a 2-dimensional electrical feature vector y1. The input of the mechanical feature network is n s 、n d 、v b,s and v b,d, the structure and calculation method are the same as above, and the output is a 2D mechanical feature vector y2. The input of the environmental feature network is T s and T d ,It also adopts a three-layer perceptron structure, with 4 hidden layer neurons, and outputs a 2-dimensional environmental feature vector y3.

[0156] Next, concatenate y1 and y2 and input them into the motor feature network, which uses a two-layer structure with 6 hidden layer neurons. The output is a two-dimensional motor feature vector y4. Concatenate y2 and y3 and input them into the mechanism feature network, which also uses a two-layer structure with 6 hidden layer neurons. The output is a two-dimensional mechanism feature vector y5. Finally, concatenate y4 and y5 and input them into the cluster feature network, which uses a three-layer structure with 8 neurons in the first hidden layer, 4 neurons in the second hidden layer, and the number of neurons in the output layer equal to the desired number of clusters, K = 10. The output is the clustering result y6.

[0157] Based on this hierarchical feature extraction and fusion method, the monitoring data of the 100 wind turbines in the wind farm were finally divided into 10 data collection clusters. Figure 3 As shown, the temporal correlation and spatial correlation changing trends of different data clusters are demonstrated.

[0158] Next, the operation and maintenance personnel calculated the time correlation ρ of the internal data for each data cluster. t and spatial correlation ρ s , as shown in Table 4:

[0159] Table 4 Correlation index of each collected data cluster

[0160] Cluster number <![CDATA[Time correlation ρ t > <![CDATA[Spatial correlation ρ s <!-- 14 -->]]> 1 0.92 0.88 2 0.89 0.85 3 0.86 0.82 4 0.83 0.79 5 0.80 0.75 6 0.77 0.72 7 0.74 0.68 8 0.71 0.64 9 0.68 0.60 10 0.65 0.56

[0161] Table 4 shows that there are significant differences in the temporal and spatial correlations between different data clusters. Cluster 1 has the highest correlation, indicating that the data within this cluster have strong continuity and correlation; while cluster 10 has the lowest correlation, indicating that the data within this cluster have large discreteness and uncertainty.

[0162] Combined with the collection contribution value C of each parameter data obtained in step S05 i , the operation and maintenance personnel determined the optimal sampling time window length T corresponding to each data collection cluster win , as shown in Table 5:

[0163] Table 5 Optimal sampling time window length T for each data cluster win ;

[0164] Cluster number Collect the parameters with the highest contribution value <![CDATA[T win (seconds)]]> 1 Voltage 0.5 2 Voltage, current 0.8 3 Current 1.0 4 Current, speed 1.5 5 Speed 2.0 6 Speed, vibration 2.5 7 vibration 3.0 8 Vibration, temperature 4.0 9 temperature 5.0 10 temperature 6.0

[0165] As can be seen from Table 5, for collecting voltage and current data with higher contribution values, the optimal sampling time window length is shorter, ranging from 0.5 to 1.0 seconds; while for collecting vibration and temperature data with lower contribution values, the optimal sampling time window length is longer, ranging from 3.0 to 6.0 seconds. Such an adaptive adjustment mechanism can not only ensure the high time resolution collection of key parameter data, but also take into account the collection efficiency of secondary parameters, thus optimizing the performance of the overall monitoring system. Figure 4 As shown in the figure, the optimal sampling time window length distribution of each data cluster is displayed in the form of a histogram.

[0166] Finally, the operation and maintenance personnel of the wind farm will apply the above-determined adaptive sampling frequency to the actual monitoring data collection process. During operation, the monitoring system will monitor the time correlation ρ of each collected data cluster in real time. t and spatial correlation ρ s , and dynamically adjust the sampling frequency f according to the corresponding relationship between Table 4 and Table 5 s , the specific calculation formula is:

[0167]

[0168] Through this adaptive clustering adjustment method, the wind farm monitoring system can ensure high-quality collection of key parameter data in a complex operating environment, while taking into account the optimization of overall collection efficiency, providing a reliable data basis for subsequent status monitoring and fault diagnosis.

[0169] For example, during operation of a 3MW wind turbine under extreme weather conditions (wind speeds reaching 25 meters per second, temperatures dropping to -20 degrees Celsius, and severe snow and ice cover), the monitoring system collected the following data:

[0170] 1. The initial sampling frequency is set to 0.5 Hz, or 2 samples per second.

[0171] 2. After EMD component decomposition and data quality assessment, the monitoring parameter data of the unit is divided into cluster 3. According to Table 4 and Table 5, the time correlation of this cluster ρ t =0.86, spatial correlation ρ s =0.82, optimal sampling time window length T win =1.0 sec.

[0172] 3. Since the current data collection contribution value C in this cluster i =0.87 is the highest, so the sampling frequency f s =1 / T win = 1 Hz. The contribution of speed data collection is second, f s = 0.8 Hz; the contribution of vibration and temperature data is low, fs =0.6 Hz.

[0173] 4. In the next 1 second, the monitoring system collected 2 voltage samples, 1 current sample, 1 speed sample, 1 vibration sample and 1 temperature sample respectively.

[0174] 5. As the unit's operating status changed over time, the monitoring system detected a decrease in the temporal and spatial correlations of various parameters. Consequently, the system automatically adjusted the sampling frequency to 1.2 Hz for voltage and current, and 0.8 Hz for speed, vibration, and temperature.

[0175] 6. Through continuous adaptive adjustment, the monitoring system ensures high-resolution acquisition of key parameter data while taking into account the acquisition efficiency of secondary parameters, comprehensively reflecting the changes in the operating status of the unit in harsh environments, and providing a reliable basis for subsequent fault diagnosis. Figure 5 As shown in Figure 2, the dynamic changes of the sampling frequency of each parameter under extreme weather conditions are demonstrated.

[0176] It should be noted that the explanations of the variables involved in the present invention are shown in Table 6.

[0177] Table 6 Variable explanation table

[0178]

[0179] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered by the scope of protection of the present invention.

Claims

1. A wind power generation monitoring data collection and clustering method, characterized in that: include: The initial sampling frequency is set based on the operating parameters of the wind turbine generator set to collect voltage data, current data, speed data, vibration data, and temperature data. The voltage data, current data, speed data, vibration data, and temperature data are preprocessed to obtain preprocessed data. The preprocessed data are component-decomposed to obtain stable components and variable components. A collection data quality assessment matrix is ​​established. The collection data quality assessment matrix includes data integrity, data volatility, and data correlation. The importance of collection features is evaluated based on the collection data quality assessment matrix, and the collection contribution value is calculated. A multi-layer lightweight clustering method is used to divide the data into collection data clusters. The multi-layer lightweight clustering method performs multi-level feature extraction and fusion through electrical feature lightweight neural networks, mechanical feature lightweight neural networks, environmental feature lightweight neural networks, motor feature lightweight neural networks, mechanism feature lightweight neural networks, and clustering feature lightweight neural networks. The optimal sampling time window length is determined based on the collection data cluster, and the sampling interval is updated in a manner that the sampling interval is inversely proportional to the collection contribution value and is applied to wind turbine generator set monitoring data collection.

2. The wind power generation monitoring data collection and clustering method according to claim 1, characterized in that: The collected data includes voltage data, current data, rotation speed data, vibration data, and temperature data, and the preprocessing includes denoising and standardization.

3. The wind power generation monitoring data collection and clustering method according to claim 1, characterized in that: The component decomposition obtains voltage stability component, voltage variation component, current stability component, current variation component, speed stability component, speed variation component, vibration stability component, vibration variation component, temperature stability component, and temperature variation component, establishes wind turbine generator set monitoring data collection evaluation index, calculates collection quality scores of voltage data, current data, speed data, vibration data, and temperature data based on the wind turbine generator set monitoring data collection evaluation index, and establishes a collection data quality assessment matrix, which includes data integrity, data volatility, and data correlation.

4. The wind power generation monitoring data collection and clustering method according to claim 1, characterized in that: The multi-layer lightweight clustering method includes constructing an electrical feature lightweight neural network, a mechanical feature lightweight neural network, and an environmental feature lightweight neural network. The electrical feature lightweight neural network inputs a voltage stability component, a voltage variation component, a current stability component, and a current variation component.

5. The wind power generation monitoring data collection and clustering method according to claim 4, characterized in that: The mechanical feature lightweight neural network inputs a speed stability component, a speed variation component, a vibration stability component, and a vibration variation component, and the environmental feature lightweight neural network inputs a temperature stability component and a temperature variation component.

6. The wind power generation monitoring data collection and clustering method according to claim 5, characterized in that: The first network output of the electrical feature lightweight neural network and the second network output of the mechanical feature lightweight neural network are jointly input into the motor feature lightweight neural network, and the second network output of the mechanical feature lightweight neural network and the third network output of the environmental feature lightweight neural network are jointly input into the mechanism feature lightweight neural network.

7. The wind power generation monitoring data collection and clustering method according to claim 6, characterized in that: The fourth network output of the motor feature lightweight neural network and the fifth network output of the mechanism feature lightweight neural network are jointly input into the cluster feature lightweight neural network, and the collected data cluster is determined according to the sixth network output of the cluster feature lightweight neural network.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program instructions, and when the program instructions are run in a computer, they are used to execute the wind power generation monitoring data collection and clustering method according to any one of claims 1 to 7.

9. A wind power generation monitoring data collection and clustering system, characterized in that: Contains the computer-readable storage medium of claim 8.

Citation Information

Patent Citations

  • Report data sampling method for wastewater treatment plant

    CN105243127A

  • Hydropower station gate detection method and system

    CN119394620A