Method for monitoring fan main shaft bearing based on multi-sensor data fusion and deep learning
By using multi-sensor data fusion and deep learning to monitor the main shaft bearing of the wind turbine, the problem of low accuracy in bearing fault diagnosis under low-speed conditions has been solved, achieving efficient and accurate fault identification and early warning, and improving the operating efficiency and economic benefits of the wind turbine.
Patent Information
- Application Number
- CN202511119267.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-21
AI Technical Summary
Traditional bearing fault diagnosis methods are not very accurate under low-speed conditions and are easily affected by noise. Existing research focuses on data processing of single sensors, neglecting the correlation between multiple sensors, which limits the diagnostic performance in complex low-speed environments.
A method for monitoring the main shaft bearing of a wind turbine is adopted by multi-sensor data fusion and deep learning. The acoustic emission signal is filtered by multi-level filters, a Sensor-CNN network model is constructed, and the multi-sensor data is fused by self-attention mechanism to construct an MSISFNet model for comprehensive analysis.
It significantly improves the diagnostic accuracy of bearing faults under low-speed operating conditions, enhances the robustness of the model in noisy environments and the accuracy of fault classification, enables early fault identification and warning, and reduces equipment downtime and maintenance costs.
Smart Images

Figure CN120995081A_ABST
Abstract
Description
Technical fields:
[0001] This invention relates to the field of wind turbine operation and maintenance management, specifically a method for wind turbine bearing fault diagnosis using multi-sound emission sensor data fusion and Inception-SE neural network. Background technology:
[0002] With the continuous growth of global energy demand and the increasing emphasis on renewable energy, wind power has developed rapidly. The main shaft bearing of a wind turbine, as a critical component, bears continuous mechanical loads and torques, and its operating status directly affects the efficiency and reliability of the entire system. Traditional bearing fault diagnosis methods mainly rely on vibration analysis and acoustic emission technology. Vibration analysis is highly effective under high-speed operating conditions, but under low-speed conditions, fault signals are weak and easily interfered with by noise, leading to decreased diagnostic accuracy. Existing research largely focuses on data processing from single sensors, neglecting the correlation between multiple sensors, thus limiting diagnostic performance in complex low-speed environments. Therefore, there is an urgent need for a comprehensive diagnostic method that combines multi-sensor data fusion with advanced deep learning models to improve the accuracy and robustness of bearing fault detection under low-speed conditions. Summary of the Invention:
[0003] The purpose of this invention is to provide a method for monitoring wind turbine main shaft bearings based on multi-sensor data fusion and deep learning. This method is used to solve the problems of low accuracy and low efficiency in detecting bearing faults under low-speed operating conditions.
[0004] The technical solution adopted by this invention to solve its technical problem is as follows: This method for monitoring the main shaft bearing of a wind turbine based on multi-sensor data fusion and deep learning includes the following steps:
[0005] S1. Collect acoustic emission signal data from multiple acoustic emission sensors under normal operating conditions and fault operating conditions of the fan main shaft. The fault operating conditions include: damage to the outer ring of the main shaft, damage to the inner ring, damage to the rolling elements, and damage to multiple rolling elements.
[0006] S2. Apply a multi-stage filter to screen and filter the acoustic emission signal collected by S1, retaining signals with obvious fault characteristics;
[0007] S3. Divide the filtered acoustic emission signal into training set and validation set according to each sensor. For each sensor's training set and validation set, construct and train the corresponding Sensor-CNN network model respectively.
[0008] S4. The obtained Sensor-CNN network models are fused into a single network model, MSISFNet, using a self-attention mechanism to achieve comprehensive analysis of multi-sensor data.
[0009] S5. The filtered signal from S2 is divided into training and validation sets without distinguishing between sensors. The preprocessed training set samples are input into the MSISFNet network model for feature extraction, training, and validation. A comprehensive evaluation metric is used to evaluate the performance of the MSISFNet network model.
[0010] The preprocessed training set samples are input into the Sensor-CNN network model for feature extraction, training, and validation, and the performance of the Sensor-CNN network model is evaluated using a comprehensive evaluation metric.
[0011] S6. Use the trained network model MSISFNet to identify spindle faults, output fault categories, and display the diagnostic results and warning information in real time through the display and warning module.
[0012] In the above scheme, the construction and training of the corresponding Sensor-CNN network model in S3 are specifically as follows:
[0013] S31. Each Sensor CNN contains three cascaded improved Inception modules. The Inception module contains the following four branches: a small-scale convolution branch extracts local features through 1×1 convolution; a medium-scale convolution branch first reduces dimensionality through 1×1 convolution, and then extracts the main patterns through 3×3 depthwise separable convolution; a large-scale convolution branch reduces dimensionality through 1×1 convolution, and then uses 5×5 depthwise separable convolution to perceive global features; and a pooling branch uses max pooling to compress the time dimension, and then adjusts the number of channels through 1×1 convolution to retain global key information.
[0014] S32. Batch normalization is combined after each branch to improve convergence speed and enhance generalization performance. The outputs of the four branches are fused in the channel dimension to form a multi-scale feature map, capturing the multi-scale features of the signal.
[0015] S33. Add a channel attention module, i.e., an SE module, after each Inception module. The SE module focuses on key features through an adaptive weighting mechanism, thereby improving the feature representation capability of the network.
[0016] S34. Based on the performance of the validation set, the ReduceLROnPlateau learning rate scheduler is adopted. When the loss function of the validation set no longer decreases significantly within 10 consecutive training cycles, the current learning rate is automatically reduced by 50%.
[0017] S35. Use the Adam optimizer to update the parameters of the Sensor-CNN network model;
[0018] S36. Train the Sensor-CNN model using the training set and validate and tune the parameters using the validation set.
[0019] In the above scheme, S4 is as follows: Construct a neural network model MSISFNet, and input the signals processed by each acoustic emission sensor into an improved convolutional neural network Sensor CNN. This network integrates the Inception module and the Squeeze-and-Excitation module to achieve multi-scale feature extraction and channel attention optimization. The features extracted by each Sensor-CNN network model are fused at the feature level through a self-attention mechanism to generate a global feature representation. Finally, the fused feature vector is passed through a dropout layer, a fully connected layer, and a Softmax activation function to output the probability distribution of bearing fault categories, thereby achieving fault classification and diagnosis.
[0020] Furthermore, in the above scheme, S4 specifically refers to:
[0021] S41 Multi-source Feature Parallel Extraction: The network front end deploys n independent and structurally consistent Sensor-CNN networks in parallel, where n is the same as the number of acoustic emission sensors deployed on site. Each Sensor-CNN sequentially completes local, mesoscale, and global feature extraction and channel attention weighting. The 64-dimensional feature vectors output by each Sensor-CNN network are stacked along the sensor dimension to form an initial fusion feature map of shape [N, n, 64], which completes the complementary integration of multi-sensor information at the parameter level.
[0022] S42 Global Feature Self-Attention Fusion: The [N,n,64] feature map is directly fed into the self-attention fusion unit. Through query-key-value mapping and weighted summation, the correlation between the features of each sensor is calculated at once and the weights are dynamically assigned. The output is a global feature representation of shape [N,64].
[0023] S43 Classification Decision and Overfit Suppression: The fused 64-dimensional global features are first randomly suppressed by the Dropout layer, then raised to 128 dimensions by the fully connected hidden layer, and finally mapped to the 4-dimensional Softmax output space to output the bearing fault category;
[0024] S44 end-to-end joint training: MSISFNet integrates the fusion unit, each Sensor-CNN and the classification layer into a single differentiable network, updating all parameters simultaneously through a single backpropagation; joint training enables each Sensor-CNN network to be optimized in tandem with the global task, preserving the fine-grained features of a single sensor and achieving complementary gains across sensors.
[0025] In the above scheme, S2 specifically refers to:
[0026] S21. Background noise removal: In the acoustic emission analysis software, a threshold of 35dB is set to remove background noise in the signal with an amplitude lower than the threshold.
[0027] S22. Select time-domain features: Use time-domain statistical features to filter out effective signal segments containing fault information. Time-domain statistical features include mean, variance, power, root mean square (RMS), kurtosis, and skewness.
[0028] S23. Smoothing the signal: Apply a high-pass filter with a cutoff frequency of 50Hz and a Savitzky-Golay filter to smooth the signal. Set the polynomial order of the filter to 3 and the window size to 21. The Savitzky-Golay filter smooths the data through local polynomial fitting. Given a window of length 2m+1, the output of the filter is the weight coefficients calculated through polynomial fitting.
[0029] S24. Signal Extraction: Extract a signal segment of 1024 data points around the point of maximum amplitude to ensure focused analysis of fault characteristics. For insufficient segments, use zero padding.
[0030] In the above scheme, when S5 uses comprehensive evaluation metrics to evaluate the performance of the network model MSISFNet, the following evaluation metrics are used to comprehensively evaluate the network model MSISFNet: accuracy, loss value, precision, and recall. Accuracy is the ratio of the number of correctly predicted samples to the total number of samples. Loss value is the difference between the predicted class and the true class. Precision is the proportion of positive class predictions for each class out of all samples predicted as positive. Recall is the proportion of positive class predictions for each class out of all samples actually predicted as positive. The F1 score is the harmonic mean of precision and recall, which represents the model's overall ability to predict positive classes.
[0031] Beneficial effects:
[0032] 1. By integrating data from multiple acoustic emission sensors, MSISFNet can fully capture and fuse multidimensional feature information from different sensors, significantly improving the diagnostic accuracy of bearing faults under low-speed conditions. Experimental results show that MSISFNet achieves an accuracy of 99.15% on the test set, far exceeding traditional single-sensor methods.
[0033] 2. The multi-level signal filtering technology of this invention effectively removes background noise and high-frequency interference. Combined with the self-attention mechanism for feature-level fusion, the model can still accurately identify fault characteristics in noisy environments, thus improving the robustness and reliability of the system under complex working conditions.
[0034] 3. To address the problem of weak fault signals in wind turbine main shaft bearings during low or very low speed operation, this invention optimizes the signal processing flow and deep learning model design, enabling effective extraction and identification of fault features under low-speed conditions. This effectively solves the problems of weak signals and severe noise interference in traditional diagnostic techniques under low-speed conditions, and has broad application prospects.
[0035] 4. This invention employs an improved Inception module combined with the SE (Squeeze-and-Excitation) module to achieve efficient extraction of multi-scale features and channel attention optimization, thereby enhancing the model's ability to select key features and improving the accuracy of fault classification.
[0036] 5. By enabling early and accurate bearing fault diagnosis, this invention can promptly identify potential problems, arrange maintenance and replacement in advance, reduce equipment downtime and maintenance costs, and improve the operating efficiency and economic benefits of wind turbines.
[0037] 6. Validation results on different experimental platforms show that the MSISFNet model has strong generalization ability and can adapt to different experimental environments and working conditions, ensuring stability and reliability in practical applications.
[0038] 7. By combining existing multi-sensor technology and advanced deep learning algorithms, this invention automates and automates complex fault diagnosis tasks, simplifies the fault diagnosis process, and reduces the complexity of technical implementation.
[0039] 8. This invention utilizes multi-sound emission sensor signal fusion technology combined with deep learning to perform online fault status monitoring and analysis of wind turbine main shaft bearings, enabling early detection and effective warning of wind turbine main shaft bearing faults, improving the accuracy and robustness of bearing fault detection under low-speed conditions, and aiming to solve the problems of low accuracy and low efficiency in bearing fault detection under low-speed operating conditions. Attached image description:
[0040] Figure 1 This is a flowchart illustrating the overall process of wind turbine fault diagnosis according to the present invention.
[0041] Figure 2 This diagram shows the experimental platform layout used for verifying the process of this invention. This invention uses two sets of experimental equipment to verify the proposed monitoring method. The two diagrams on the left show a small bearing test bench; the diagram on the right shows the bearing used in an actual fan. Both test benches are prefabricated with four types of internal defects.
[0042] Figure 3 The image shows a replacement bearing for a faulty bearing on the small experimental platform of this invention. There are four types of pre-fabricated defects, from left to right: multiple rolling elements, rolling elements, inner ring, and outer ring.
[0043] Figure 4 This is a flowchart of the multi-stage filtering process of the present invention;
[0044] Figure 5 This is a schematic diagram of the zero-padding step in this invention;
[0045] Figure 6 This is a schematic diagram illustrating the effect of applying the filter before and after the present invention;
[0046] Figure 7 This is a structural diagram of the feature extraction module of the present invention;
[0047] Figure 8 This is a structural diagram of the data fusion module of the present invention;
[0048] Figure 9 To compare the results of the fusion model with the single Sensor-CNN model in this invention, three fusion models and a Sensor-CNN model are used as comparison models. The three fusion models use different methods to fuse data from multiple sensors, while the single Sensor-CNN is trained using data from different acoustic emission sensors.
[0049] Figure 10 This shows the accuracy curve and evaluation matrix of the MSISFNet model of this invention. Detailed implementation method:
[0050] The present invention will be further described below with reference to the accompanying drawings:
[0051] This method for monitoring wind turbine main shaft bearings based on multi-sensor data fusion and deep learning first preprocesses bearing vibration signals from multiple sensors using multi-level signal filtering techniques. This includes background noise removal, time-domain feature selection, signal smoothing, and fault feature extraction. Specifically, the steps are: filtering out background noise signals with amplitudes below 35dB using a set threshold; selecting effective signal segments based on time-domain statistical features such as mean, variance, energy, root mean square (RMS), kurtosis, and skewness; smoothing the signal using a Savitzky-Golay filter to reduce high-frequency noise interference; and extracting fault feature intervals centered on the signal's maximum amplitude point. Subsequently, a neural network model is constructed, inputting the processed signals from each sensor into an improved convolutional neural network. This network integrates an Inception module and a Squeeze-and-Excitation module to achieve multi-scale feature extraction and channel attention optimization. Features extracted from multiple sensors are fused at the feature level through a self-attention mechanism to generate a global feature representation. Finally, the fused feature vector is processed through a dropout layer, a fully connected layer, and a Softmax activation function to output the probability distribution of bearing fault categories, achieving fault classification and diagnosis.
[0052] The present invention specifically includes the following steps:
[0053] S1. Collect acoustic emission signal data from multiple acoustic emission sensors under normal and fault conditions of the fan main shaft. The fault conditions include: damage to the outer ring of the main shaft, damage to the inner ring, damage to the rolling elements, and damage to multiple rolling elements.
[0054] S11, Multi-sound emission sensor:
[0055] In this embodiment, three acoustic emission sensors are installed on the outer ring of the main shaft bearing of the wind turbine generator set, with each sensor arranged at an angle of 120° (see appendix). Figure 2 (Side view of a small experimental bench). These sensors are used to collect acoustic emission signals of the spindle bearing in real time under normal and fault conditions, including damage to the outer ring, inner ring, rolling elements, and multiple rolling elements.
[0056] S12, Signal conditioning circuit:
[0057] The signal conditioning circuit includes a preamplifier and a filter:
[0058] Preamplifier: A low-noise, high-gain operational amplifier is used to amplify the weak acoustic emission signal from the sensor.
[0059] Filter: Configure a bandpass filter, set the passband frequency to 20kHz to 1MHz, filter out unwanted low-frequency and high-frequency noise, and ensure signal quality.
[0060] S13, Data Acquisition Module:
[0061] Equipment: Employs a high sampling rate data acquisition card, supporting a sampling rate of at least 1MHz to ensure high-fidelity acquisition of acoustic emission signals.
[0062] Interface: Connects to a computer via USB 3.0 interface, supports multi-channel synchronous acquisition, and ensures time consistency of signals from various sensors.
[0063] S14, Control and Communication Module:
[0064] Embedded controller: The Arduino Mega 2560 is used as the control core to coordinate the working status of the sensors and the data acquisition process.
[0065] Communication interface: Data transfer with a computer is achieved via a USB interface, ensuring real-time performance and stability.
[0066] S15, Power Management Module:
[0067] Power supply: A voltage regulator module is used to provide a stable voltage and ensure the normal operation of all components.
[0068] Protection circuit: Built-in overvoltage and overcurrent protection to prevent abnormal power supply from damaging the system.
[0069] S16. Housing and mounting structure:
[0070] Protective housing: The protective housing is made of metal and is dustproof and waterproof, adapting to the working environment of the wind turbine main shaft bearing.
[0071] Mounting bracket: Equipped with an adjustable mounting bracket to ensure that the sensor and data acquisition module are securely installed and to avoid signal acquisition errors caused by vibration.
[0072] Multiple acoustic emission sensors acquire acoustic emission signals from the main shaft bearing in real time. The acquired digital acoustic emission signals are stored and transmitted to the signal processing module, providing a high-quality data foundation for subsequent feature extraction and fault diagnosis. Through the aforementioned signal acquisition device, efficient and accurate acoustic emission signal acquisition of the wind turbine main shaft bearing under normal and fault conditions can be achieved, providing reliable data support for subsequent multi-sensor data fusion and deep learning fault diagnosis.
[0073] S2. Apply a multi-stage filter to screen and filter the acoustic emission signal collected by S1, retaining signals with obvious fault characteristics.
[0074] Signal preprocessing, such as Figure 4 As shown, signal processing involves four steps:
[0075] S21. Background noise removal:
[0076] Set a threshold of 35dB in the acoustic emission analysis software to remove background noise with an amplitude below this threshold, ensuring that only valid signals are retained in subsequent analyses.
[0077] S22, Time-domain feature selection:
[0078] Valid signal segments containing fault information are selected using time-domain statistical characteristics such as mean, variance, energy, power, root mean square (RMS), kurtosis, and skewness. Specific selection thresholds are shown in the table below.
[0079]
[0080] S23, Signal Smoothing:
[0081] A high-pass filter with a cutoff frequency of 50Hz is applied, and a Savitzky-Golay filter is used to smooth the signal and reduce noise interference. The polynomial order of the filter is set to 3, and the window size is set to 21. The Savitzky-Golay filter smooths the data through local polynomial fitting. Given a window of length 2m+1, the filter output y[i] can be calculated using the following formula:
[0082]
[0083] Where y[i] is the original data point, It is the smoothed output value, c j These are the weighting coefficients calculated through polynomial fitting; coefficient c j The following formula can be used to calculate the order of a polynomial of degree n:
[0084] C = (X T X)- 1 X T Y
[0085] Where X is the design matrix, containing the values of the polynomial basis functions within the window (e.g., 1, x, x). 2 ...x n Y is the signal vector, and C is the coefficient vector. The design matrix X is constructed based on the relative positions within the window and the polynomial order n. See the before and after images. Figure 6 The goal of smoothing is to present the slow changes in values so that it is easier to see the trend of the data and to reduce identification errors caused by sudden changes in the data.
[0086] S24, Signal interception:
[0087] A signal segment of 1024 data points was extracted around the point of maximum amplitude to ensure focused analysis of fault characteristics. For any insufficient segments, zero-padding was used to bring the total to 1024 data points. The zero-padding method is as follows: Figure 5 As shown.
[0088] S3. Divide the filtered acoustic emission signal into training and validation sets according to each sensor. For each sensor's training and validation sets, construct and train the corresponding Sensor-CNN network model.
[0089] S31. Constructing a Sensor-CNN network model:
[0090] Feature extraction module, such as Figure 7As shown, an independent convolutional neural network architecture is designed for each sensor to ensure that the model can effectively extract the signal features specific to that sensor. Each Sensor-CNN contains three cascaded improved Inception modules, with the specific structure as follows:
[0091] Small-scale convolution branch: Extracts local features through 1×1 convolution.
[0092] Mesoscale convolution branch: First, dimensionality is reduced by 1×1 convolution, and then the main patterns are extracted by 3×3 depthwise separable convolution.
[0093] Large-scale convolution branch: Dimensionality is reduced using 1×1 convolution, followed by 5×5 depthwise separable convolution to perceive global features.
[0094] Pooling branch: Max pooling is used to compress the time dimension, and then the number of channels is adjusted by 1×1 convolution to retain global key information.
[0095] S32, Batch Normalization:
[0096] Each branch is followed by batch normalization to improve convergence speed and generalization performance. The outputs of the four branches are fused along the channel dimension to form a multi-scale feature map, thereby effectively capturing the multi-scale features of the signal and improving the overall feature representation capability.
[0097] S33, Channel Attention Module (SE Module):
[0098] A channel attention module is added after each Inception module. Through an adaptive weighting mechanism, it enhances the focus on key features and improves the network's feature representation ability. This mechanism enables the model to dynamically adjust the importance of each channel, further improving diagnostic accuracy.
[0099] S34, Learning Rate Scheduling:
[0100] Based on the validation set performance, the ReduceLROnPlateau learning rate scheduler is employed. When the loss function of the validation set no longer decreases significantly over 10 consecutive training epochs, the current learning rate is automatically reduced by 50%. This strategy dynamically adjusts the global learning rate by monitoring changes in the validation set loss, thereby promoting fine-tuning of the model in the later stages of training and improving convergence speed and model performance.
[0101] S35, Optimizer Settings:
[0102] The Adam optimizer was used to update the model parameters, with an initial learning rate of 0.001. The Adam optimizer combines the advantages of momentum and adaptive learning rate adjustment, dynamically adjusting the learning rate of each parameter to accelerate convergence, improve training stability, and effectively avoid the model getting trapped in local optima. During training, a batch size of 32 was set to balance training efficiency and model performance.
[0103] S36. Model Training and Validation:
[0104] During training, the Sensor-CNN model is trained using the training set and validated and its parameters tuned using the validation set. By monitoring the loss and accuracy on the validation set, the learning rate and other training parameters are dynamically adjusted to ensure that the model achieves optimal performance at different training stages.
[0105] S4. The obtained Sensor-CNN network models are fused into a single network model, MSISFNet, using a self-attention mechanism to achieve comprehensive analysis of multi-sensor data.
[0106] like Figure 8 The above Sensor-CNN is shown as being fused together.
[0107] S41, Feature Vector Fusion:
[0108] First, the feature vectors extracted from the three sensors are fused to form an initial fused feature map of shape [N, 3, 64]. Here, N is the number of samples, 3 is the number of sensors, and 64 is the feature dimension of each sensor. This step integrates multi-scale features from different sensors as the basis for subsequent fusion.
[0109] S42, Self-attention mechanism fusion:
[0110] Feature fusion is further achieved through a self-attention mechanism. The specific process is as follows:
[0111] The fused feature maps are mapped to the Query, Key, and Value vector spaces, respectively.
[0112] The dot product operation is used to calculate the correlation between the query and the key.
[0113] The value vectors are weighted and summed, and the final output fused feature shape is [N, 64].
[0114] This mechanism can dynamically adjust the weights of different sensor features, thereby improving the expressive power of fused features and diagnostic accuracy.
[0115] S43. Overfitting Prevention and Classification Layer:
[0116] To reduce the risk of overfitting, the fused feature vectors undergo random neuron dropout in a dropout layer before entering the classification layer. Subsequently, the feature vectors are mapped to 128 dimensions through a hidden layer and finally mapped to the class space, with an output shape of [N, 4], where 4 represents the number of fault categories. The application of the dropout layer effectively prevents model overfitting and improves the model's generalization ability.
[0117] S44, End-to-end Joint Training:
[0118] The entire fusion process achieves end-to-end joint training. Through joint training, the fusion model MSISFNet and the Sensor-CNN models of each sensor can be optimized collaboratively, further improving the accuracy of fault diagnosis and the robustness of the model.
[0119] S4 is more specifically:
[0120] S41 Parallel Extraction of Multi-Source Features: Three independent and structurally consistent Sensor-CNN sub-networks are deployed in parallel at the network front end, the number of which is the same as the number of acoustic emission sensors deployed on-site. Each Sensor-CNN follows the three-level Inception-SE structure described in S3 of the instruction manual, and sequentially completes local, mesoscale, and global feature extraction and channel attention weighting; the 64-dimensional feature vectors output by each sub-network are stacked along the sensor dimension to form an initial fusion feature map of shape [N,3,64], thereby completing the complementary integration of multi-sensor information at the parameter level.
[0121] S42 Global Feature Self-Attention Fusion: The [N,3,64] feature map is directly fed into the self-attention fusion unit. Through query-key-value mapping and weighted summation, the correlation between the features of each sensor is calculated at once and the weights are dynamically assigned. The output is a global feature representation of shape [N,64]. This process does not require manual setting of weights. The network automatically learns the importance of different sensors in different fault modes during training, which significantly improves the robustness in low-speed and low signal-to-noise ratio environments.
[0122] S43 Classification Decision and Overfit Suppression: The fused 64-dimensional global features are first randomly suppressed by the Dropout layer, then raised to 128 dimensions by the fully connected hidden layer, and finally mapped to the 4-dimensional Softmax output space, corresponding to the four bearing fault categories described in S1 of the specification; this path is short and efficient, reducing the computational load while ensuring accuracy.
[0123] S44 end-to-end joint training: MSISFNet integrates the fusion unit, each Sensor-CNN and the classification layer into a single differentiable network, updating all parameters simultaneously through a single backpropagation; joint training enables each sub-network to be optimized in conjunction with the global task, preserving the fine-grained features of a single sensor while achieving complementary gains across sensors.
[0124] S5. The filtered signal from S2 is directly divided into training and validation sets without distinguishing between sensors. The preprocessed training set samples are then input into the MSISFNet network model for feature extraction, training, and validation. A comprehensive evaluation metric is used to evaluate the performance of the MSISFNet network model.
[0125] The input is fed into the Sensor-CNN base model for feature extraction, training, and validation, and the performance of the Sensor-CNN base model is evaluated using a comprehensive evaluation metric.
[0126] Evaluation metrics are used to comprehensively evaluate the model. These metrics include accuracy, loss value, precision, and recall. Accuracy is the ratio of the number of correctly predicted samples to the total number of samples, i.e.:
[0127]
[0128] Among them, a correct prediction refers to a sample whose predicted label matches the true label.
[0129] The loss value is calculated as the difference between the predicted class and the true class, i.e.:
[0130]
[0131] Among them, y i It's a real label. It is the predicted probability value.
[0132] Precision is the proportion of positive predictions for each class out of all samples predicted as positive, i.e.:
[0133]
[0134] Recall is the proportion of positive predictions for each class out of all actual positive samples, i.e.:
[0135]
[0136] In this context, TP is the true class and FN is the false negative class.
[0137] The F1 score is the harmonic mean of precision and recall, representing the model's overall ability to handle the positive class.
[0138]
[0139] TP indicates that the predicted class is positive and the actual class is positive; FP indicates that the predicted class is positive but the actual class is negative; TN indicates that the predicted class is negative and the actual class is negative; FN indicates that the predicted class is negative but the actual class is positive.
[0140] The model performance is evaluated using a comprehensive set of evaluation metrics, including accuracy, loss value, precision, and recall. Figure 9 The results of the comparison between the fusion model and the single Sensor-CNN model are presented. Figure 10 The accuracy curve and evaluation matrix of the MSISFNet model are shown.
[0141] S6. Use the trained network model MSISFNet to identify spindle faults, output fault categories, and display the diagnostic results and warning information in real time through the display and warning module.
[0142] The trained MSISFNet network model is used to identify spindle faults, output fault categories, and display the diagnostic results and early warning information in real time through the display and early warning module so that maintenance measures can be taken in a timely manner.
[0143] Overall process of this invention:
[0144] Data Acquisition (S1): The multi-sound emission sensor acquires the acoustic emission signal of the spindle bearing in real time, and the signal is amplified and filtered by the signal conditioning circuit to improve the signal quality.
[0145] Signal preprocessing (S2): background noise removal, time-domain feature selection, signal smoothing and signal truncation are performed to extract the effective signal segment containing fault information.
[0146] Feature extraction and model training (S3): Construct and train Sensor-CNN network models for each sensor to extract multi-scale features.
[0147] Data fusion and model building (S4): The MSISFNet model is constructed by fusing features from various sensors through a self-attention mechanism to achieve comprehensive analysis of multi-sensor data.
[0148] Model evaluation and fault identification (S5, S6): The model performance is evaluated using comprehensive evaluation metrics, and the spindle fault is identified using the MSISFNet model, outputting diagnostic results and early warning information.
[0149] Through the above specific implementation methods, the present invention can achieve efficient and accurate online fault status monitoring and analysis of the main shaft bearing of wind turbine generator under low-speed operating environment, meeting the technical requirements for early detection and effective early warning.
Claims
1. A method for monitoring the main shaft bearing of a wind turbine based on multi-sensor data fusion and deep learning, characterized in that... Includes the following steps: S1. Collect acoustic emission signal data from multiple acoustic emission sensors under normal operating conditions and fault operating conditions of the fan main shaft. The fault operating conditions include: damage to the outer ring of the main shaft, damage to the inner ring, damage to the rolling elements, and damage to multiple rolling elements. S2. Apply a multi-stage filter to screen and filter the acoustic emission signal collected by S1, retaining signals with obvious fault characteristics; S3. Divide the filtered acoustic emission signal into training set and validation set according to each sensor. For each sensor's training set and validation set, construct and train the corresponding Sensor-CNN network model respectively. S4. The obtained Sensor-CNN network models are fused into a single network model, MSISFNet, using a self-attention mechanism to achieve comprehensive analysis of multi-sensor data. S5. The filtered signal from S2 is divided into training and validation sets without distinguishing between sensors. The preprocessed training set samples are input into the MSISFNet network model for feature extraction, training, and validation. A comprehensive evaluation metric is used to evaluate the performance of the MSISFNet network model. The preprocessed training set samples are input into the Sensor-CNN network model for feature extraction, training, and validation, and the performance of the Sensor-CNN network model is evaluated using a comprehensive evaluation metric. S6. Use the trained network model MSISFNet to identify spindle faults, output fault categories, and display the diagnostic results and warning information in real time through the display and warning module.
2. The method for monitoring wind turbine main shaft bearings based on multi-sensor data fusion and deep learning according to claim 1, characterized in that: The specific steps for constructing and training the corresponding Sensor-CNN network model in S3 are as follows: S31. Each Sensor CNN contains three cascaded improved Inception modules. Each Inception module contains the following four branches: a small-scale convolution branch extracts local features through 1×1 convolutions; a medium-scale convolution branch first reduces dimensionality through 1×1 convolutions and then extracts the main patterns through 3×3 depthwise separable convolutions; a large-scale convolution branch reduces dimensionality through 1×1 convolutions and then uses 5×5 depthwise separable convolutions to perceive global features; and a pooling branch uses max pooling to compress the time dimension and then adjusts the number of channels through 1×1 convolutions to retain key global information. S32. Batch normalization is combined after each branch to improve convergence speed and enhance generalization performance. The outputs of the four branches are fused in the channel dimension to form a multi-scale feature map, capturing the multi-scale features of the signal. S33. Add a channel attention module, i.e., an SE module, after each Inception module. The SE module focuses on key features through an adaptive weighting mechanism, thereby improving the feature representation capability of the network. S34. Based on the performance of the validation set, the ReduceLROnPlateau learning rate scheduler is adopted. When the loss function of the validation set no longer decreases significantly within 10 consecutive training epochs, the current learning rate is automatically reduced by 50%. S35. Use the Adam optimizer to update the parameters of the Sensor-CNN network model; S36. Train the Sensor-CNN model using the training set and validate and tune the parameters using the validation set.
3. The method for monitoring wind turbine main shaft bearings based on multi-sensor data fusion and deep learning according to claim 2, characterized in that: S4 is as follows: Construct a neural network model MSISFNet, and input the signals processed by each acoustic emission sensor into an improved convolutional neural network Sensor CNN. This network integrates the Inception module and the Squeeze-and-Excitation module to achieve multi-scale feature extraction and channel attention optimization. The features extracted by each Sensor-CNN network model are fused at the feature level through a self-attention mechanism to generate a global feature representation. Finally, the fused feature vectors are processed through a dropout layer, a fully connected layer, and a Softmax activation function to output the probability distribution of bearing fault categories, thereby achieving fault classification and diagnosis.
4. The method for monitoring wind turbine main shaft bearings based on multi-sensor data fusion and deep learning according to claim 3, characterized in that: Specifically, S4 is: S41 Parallel Extraction of Multi-Source Features: n independent and structurally consistent Sensor-CNN networks are deployed in parallel at the network front end, where n is the same as the number of acoustic emission sensors deployed on site. Each Sensor-CNN sequentially completes local, mesoscale, and global feature extraction and channel attention weighting. The 64-dimensional feature vectors output by each Sensor-CNN network are stacked along the sensor dimension to form an initial fusion feature map of shape [N, n, 64], which completes the complementary integration of multi-sensor information at the parameter level. S42 Global Feature Self-Attention Fusion: The [N, n, 64] feature map is directly fed into the self-attention fusion unit. Through query-key-value mapping and weighted summation, the correlation between the features of each sensor is calculated at once and the weights are dynamically assigned. The output is a global feature representation of shape [N, 64]. S43 Classification Decision and Overfit Suppression: The fused 64-dimensional global features are first randomly suppressed by the Dropout layer, then raised to 128 dimensions by the fully connected hidden layer, and finally mapped to the 4-dimensional Softmax output space to output the bearing fault category; S44 End-to-End Joint Training: MSISFNet integrates the fusion unit, each Sensor-CNN, and the classification layer into a single differentiable network, updating all parameters simultaneously through a single backpropagation; joint training enables each Sensor-CNN network to co-optimize with the global task, preserving fine-grained features of a single sensor and achieving complementary gains across sensors.
5. The method for monitoring wind turbine main shaft bearings based on multi-sensor data fusion and deep learning according to claim 4, characterized in that: S2 specifically refers to: S21. Background noise removal: By setting a threshold of 35 dB in the acoustic emission analysis software, background noise with an amplitude lower than the threshold is removed from the signal; S22. Select time-domain features: Use time-domain statistical features to filter out effective signal segments containing fault information. Time-domain statistical features include mean, variance, power, root mean square RMS, kurtosis, and skewness. S23. Smoothing the signal: Apply a high-pass filter with a cutoff frequency of 50 Hz and a Savitzky-Golay filter to smooth the signal. Set the polynomial order of the filter to 3 and the window size to 21 by adjustment. The Savitzky-Golay filter smooths the data by local polynomial fitting. Given a window of length 2m+1, the output of the filter is the weight coefficients calculated by polynomial fitting. S24. Signal Extraction: Extract a signal segment of 1024 data points around the point of maximum amplitude to ensure focused analysis of fault characteristics. For any insufficient segments, use zero padding.
6. The method for monitoring wind turbine main shaft bearings based on multi-sensor data fusion and deep learning according to claim 5, characterized in that: When S5 uses comprehensive evaluation metrics to evaluate the performance of the MSISFNet network model, the following evaluation metrics are used: accuracy, loss value, precision, and recall. Accuracy is the ratio of the number of correctly predicted samples to the total number of samples. Loss value is the difference between the predicted class and the true class. Precision is the proportion of positive class predictions for each class out of all predicted positive samples. Recall is the proportion of positive class predictions for each class out of all actual positive samples. The F1 score is the harmonic mean of precision and recall, representing the model's overall ability to predict positive classes.
Citation Information
Cited By
Post-welding welding spot structure integrity detection method and system
CN121830906A