Fan gearbox fault diagnosis method based on multi-core wavelet denoising network

Through a multi-core wavelet denoising network, combined with wavelet core convolution, timing channel attention and Transformer encoder, the problem of noise masking characteristics and long-term dependency in gearbox fault diagnosis of wind turbine units is solved, and high-precision fault diagnosis is achieved.

CN120257046APending Publication Date: 2025-07-04NORTHEAST DIANLI UNIVERSITY
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510308243.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the fault diagnosis of gearbox of wind turbines, the prior art is easily masked by noise because the fault characteristics are difficult to capture long-term dependencies, resulting in low diagnostic accuracy and accuracy.

Method used

A multi-core wavelet denoising network is adopted, and multi-scale pulse features are extracted through the wavelet core convolution module, adaptive fusion is performed in combination with the timing channel attention mechanism, multi-scale adaptive soft threshold module is used for denoising, and global information is captured using the feature-driven Transformer encoder, and finally fault diagnosis is performed through the classifier.

Benefits of technology

It effectively improves the accuracy and accuracy of fault diagnosis in noisy environments, maintains high performance when processing long-term series signals, and solves the problems of insufficient feature extraction capabilities and difficult to capture long-term dependencies in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257046A_ABST
    Figure CN120257046A_ABST
Patent Text Reader

Abstract

The invention discloses a fan gearbox fault diagnosis method based on a multi-core wavelet denoising network, and belongs to the technical field of fault diagnosis. The method comprises the following steps: 1, collecting vibration signals to construct a data set, dividing a training set, a test set and a verification set, and carrying out normalization processing; 2, multi-scale pulse features are extracted by fusing a wavelet kernel convolution module, adaptive fusion of a time sequence channel attention mechanism is proposed, and redundant features are suppressed; a third step of integrating a multi-scale adaptive soft threshold module to realize accurate denoising and processing multi-scale information; 4, a feature-driven Transform encoder is adopted, and global information is captured by using a multi-head attention mechanism; and 5, inputting the extracted features into a classifier for classification, performing fine adjustment on the model by using a verification set, performing test set evaluation, and completing fault diagnosis. The method effectively solves the problems of low accuracy and low diagnosis precision caused by the fact that fault features are easily covered by noise and the long-term dependency relationship is difficult to capture in the fault diagnosis of the gearbox of the wind turbine generator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of fault diagnosis of wind turbine gearboxes, and particularly relates to a fault diagnosis method for a wind turbine gearbox based on a multi-core wavelet denoising network. Background Art

[0002] As an important component of a wind turbine, the fault diagnosis technology of the gearbox has important research significance for improving the operation efficiency of the wind turbine and reducing the operation cost. In order to ensure the normal operation of the wind turbine and prevent economic losses caused by the shutdown of the wind turbine unit, it is necessary to establish an efficient and accurate fault diagnosis model for the wind turbine gearbox.

[0003] Currently, the mainstream methods for fault diagnosis of wind turbine gearboxes are divided into two types: fault diagnosis methods based on signal processing technology and fault diagnosis methods based on data-driven. The methods based on signal processing technology extract multi-domain features characterizing the wind turbine state signals and analyze them to achieve fault diagnosis. For example, wavelet transform, Fourier transform, empirical mode decomposition, etc. However, multiple steps of the above methods rely on manual decision-making and have strong pertinence. For complex working conditions (such as multiple faults occurring simultaneously inside the gearbox), the self-learning ability and generalization ability are weak, and it is difficult to provide high-efficiency diagnostic performance. To overcome the limitations of the methods based on signal processing technology, data-driven fault diagnosis methods based on deep learning have gradually become a research hotspot. For example, one-dimensional convolutional network (1D-CNN), multi-scale convolutional network (MSCNN), etc. However, the existing models still have certain limitations in feature extraction ability and insufficient robustness to fault signals. Therefore, how to effectively separate the vibration signals of faulty components in a high-noise environment has become the first research difficulty at present. Although there are noise reduction models such as the deep residual shrinkage network (DRSN) that can effectively extract local features and short-term dependencies in a high-noise environment, when dealing with tasks containing complex and long-term dependencies, how the model fully mines and integrates long-distance information has become the second research difficulty at present.

[0004] Therefore, there is an urgent need for a new technical solution in the existing technology to solve this problem. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a fault diagnosis method for a wind turbine gearbox based on a multi-core wavelet denoising network, which is used to solve the problems of low accuracy and diagnostic precision in the fault diagnosis of wind turbine gearboxes in the existing technology due to the fact that fault features are easily masked by noise and it is difficult to capture long-term dependency relationships.

[0006] The technical solution adopted by the present invention is to provide a fault diagnosis method for a wind turbine gearbox based on a multi-core wavelet denoising network, including the following steps:

[0007] In the first step, a dataset is constructed using the original vibration signals collected by sensors at different positions, divided into a training set, a test set, and a validation set, and normalized.

[0008] In the second step, the training set samples are input into the fusion wavelet kernel convolution (FWKC) module to extract multi-scale pulse features, and a temporal channel attention mechanism (TCA) is proposed for adaptive fusion to suppress redundant features.

[0009] In the third step, the output multi-scale pulse features are input into the multi-scale adaptive soft threshold (MSAS) module to further process multi-scale information on the basis of achieving accurate signal denoising.

[0010] In the fourth step, a feature-driven Transformer (FD-Transformer) encoder is used to capture the global information of the samples using the multi-head attention mechanism and further extract global features.

[0011] In the fifth step, the extracted features are input into a classifier for classification, the model parameters are fine-tuned using the validation set, and tested using the test set to obtain the fault diagnosis results.

[0012] The specific steps of the first step include:

[0013] Step 1: Collect the vibration signal sample data of the wind turbine gearbox:

[0014] Specifically, a large number of labeled sample data are collected. The collected labeled sample data are the vibration signals of the wind turbine gearbox in normal state, planetary gear spalling, and inner ring cracking of high-speed and medium-speed shaft bearings, and are divided into a training set, a validation set, and a test set.

[0015] Step 2: Normalize the training set, validation set, and test set. The specific formula is as follows:

[0016]

[0017] In the formula, is denoted as the data value after normalization, x is denoted as the original data value, min(x) is denoted as the minimum data value, and max(x) is denoted as the maximum data value.

[0018] The specific steps of the second step include:

[0019] Step 1: Use wavelet basis functions to replace the convolution kernel, and combine the characteristics of continuous wavelet transform to perform multiple convolutions at different scales α and time translations β:

[0020] Specifically, three wavelet kernels of different scales, namely the Laplace wavelet, the Morlet wavelet, and the MexHat wavelet, are used to obtain the feature maps p1, p2, and p3. The continuous wavelet transform and the continuous wavelet kernel convolution p formula are as follows:

[0021]

[0022] p = ψ α,β (τ) * x(τ)

[0023] where τ represents time, x(τ) represents the input signal, ψ represents the wavelet basis function, * represents the conjugate operation of complex numbers, α represents the scale parameter, and α represents the time shift parameter;

[0024] The Laplace wavelet, Morlet wavelet, and MexHat wavelet kernel convolution formulas are as follows:

[0025]

[0026] where A and C are the normalization coefficients of the wavelet basis function, ζ is the viscous damping ratio, f0 is the sampling frequency, and δ is the delay parameter;

[0027] Step 2: Concatenate the feature maps output by the three wavelet kernel convolutions along the channel dimension to obtain the fused feature map p 123 , and the formula is as follows:

[0028] p 123 = Concat(ψ1(τ) * x(τ), ψ2(τ) * x(τ), ψ3(τ) * x(τ))

[0029]

[0030] where Concat represents concatenation;

[0031] Step 3: Use the TCA attention mechanism to fuse the features of the outputs of the continuous wavelet kernel convolutions of three different scales:

[0032] Specifically, through a lightweight convolutional network, for each time step in the time step dimension, the channel weight w is dynamically generated 123 , and the adaptive weight w 123 is used to perform weighted summation with p 123 , and the global spatial information is extracted through cross-channel global average pooling (GAP) to obtain the fused feature map p' 123 , and the backpropagation (BP) algorithm is used for backpropagation update. The specific formula is as follows:

[0033] w 123 = Sigmoid(Conv1d(p 123 ))

[0034]

[0035] In the formula, each feature map is regarded as a whole, and N k represents the number of channels of the feature map, τ represents the number of time steps, c is the number of channels of the weighted feature map, and the weighted feature map performs average pooling on the channel k values at the spatial position (N c and τ), and the dimension of the finally generated fused feature map is (N k and τ).

[0036] The specific steps of the third step include:

[0037] Step 1: Divide the MSAS module into two branches: branch_1 and branch_2; connect the two branches in parallel:

[0038] Specifically, branch_1 uses two layers of narrow kernel convolutional layers (kernel size 1×3) to extract local information containing high-frequency features, and branch_2 uses two layers of wide kernel convolutional layers (kernel size 7×7) to extract global information containing low-frequency features;

[0039] Step 2: Send the fused feature map p 123 into branch_1 and branch_2 respectively, and obtain the feature map F(p) after two convolutional operations and ReLU activation;

[0040] Step 3: Obtain the feature map F'(p) through residual connection, and the specific formula is as follows:

[0041] F'(p) = F(p) + p 123

[0042] In the formula, the feature map after two convolutional operations and ReLU activation is F(p), and the output after residual connection is F'(p);

[0043] Step 4: Introduce the SE attention mechanism to dynamically adjust the soft threshold function:

[0044] Specifically, perform global average pooling (GAP) on F'(p) to calculate the global feature representation of each channel N, that is, the global average value Z of the absolute values of the features of each channel N , generate the attention weight of each channel through a fully connected layer, and the attention coefficient S N of the Nth channel represents the importance of this channel, and the specific formula is as follows:

[0045] S N = Sigmoid(W a ReLU(W b Z N ))

[0046] Wherein, W a and W b are the weight matrices of two fully connected layers respectively, and the activation function ReLU is used to introduce non-linear transformation, and the Sigmoid function normalizes the generated weights to between [0, 1] for S N to obtain the dynamic soft threshold λ of the Nth group N ;

[0047] Step 5: Use the soft threshold function to perform dynamic denoising on the features of the two branches respectively to obtain feature maps T1 and T2, and the formula is as follows:

[0048] λ N =Z N S N

[0049] T N = sign(F N '(p)) max(|F N '(p)| - λ N , 0)

[0050]

[0051] Wherein, T N represents the feature map after soft threshold denoising, F N '(p) represents the feature of the Nth channel, and λ N represents the dynamic soft threshold;

[0052] Step 6: Linearly sum the two obtained feature maps to obtain a fused feature map T 12 , and the formula is as follows:

[0053] T 12 = T1 + T2

[0054] Wherein, T1 and T2 are the feature maps after denoising of Branch_1 and Branch_2 respectively, and T 12 is the fused feature map.

[0055] The specific steps of the fourth step include:

[0056] Step 1: Extract the feature vectors corresponding to each time step and input them into the FD-Transformer encoder;

[0057] Step 2: Stack 3 FD-Transformer encoders;

[0058] Step 3: Reduce the output feature channels to half of the input features;

[0059] Step 4: Flatten the output feature matrix.

[0060] The specific steps of the fifth step include:

[0061] Step 1: Input the flattened feature matrix into the MLP classifier for classification;

[0062] Step 2: Repeat the iteration until the training loss is the lowest;

[0063] Step 3: Input the validation set into the trained diagnostic model to fine-tune the parameters;

[0064] Step 4: Input the test set into the tuned model to obtain the diagnostic result and complete the fault diagnosis.

[0065] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.

[0066] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0067] Through the above design, the present invention can bring the following beneficial effects:

[0068] This method has made remarkable progress in solving the problems of difficult signal feature extraction in the noise environment of wind turbine gearboxes and insufficient capture of long-term dependence relationships in time series signals. Compared with existing methods, it realizes the extraction of multi-scale pulse features through the fusion of wavelet kernel convolution (FWKC) modules, proposes a time series channel attention mechanism (TCA) for adaptive fusion to suppress redundant features, and further processes multi-scale features and dynamically adjusts the soft threshold for denoising using an adaptive soft threshold module (MSAS). It effectively extracts pulse features in a high-noise environment and inputs them into a feature-driven Transformer (FD-Transformer) encoder. Combining the local feature extraction ability of the convolutional network, it uses the multi-head attention mechanism to capture the long-term dependence relationships of samples and further extracts global features, effectively improving the accuracy and precision of fault diagnosis and still maintaining high performance when processing long time series signals. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 It is the overall flowchart of a method for fault diagnosis of a wind turbine gearbox based on a multi-core wavelet denoising network according to the present invention;

[0070] Figure 2 It is a schematic diagram of the FWKC module of a method for fault diagnosis of a wind turbine gearbox based on a multi-core wavelet denoising network according to the present invention;

[0071] Figure 3Schematic diagram of the FD-Transformer encoder for a fault diagnosis method of a wind turbine gearbox based on a multi-core wavelet denoising network according to the present invention;

[0072] Figure 4 Schematic diagram of the performance of different methods in a noise environment in the embodiment of a fault diagnosis method of a wind turbine gearbox based on a multi-core wavelet denoising network according to the present invention;

[0073] Figure 5 Schematic diagram of the visualization of the training process in the embodiment of a fault diagnosis method of a wind turbine gearbox based on a multi-core wavelet denoising network according to the present invention;

[0074] Figure 6 Schematic diagram of the visualization of the attention scores of different attention heads in the embodiment of a fault diagnosis method of a wind turbine gearbox based on a multi-core wavelet denoising network according to the present invention;

[0075] Figure 7 Schematic diagram of the visualization of the attention scores of signals of different lengths in the embodiment of a fault diagnosis method of a wind turbine gearbox based on a multi-core wavelet denoising network according to the present invention. Detailed implementation manners

[0076] The following further describes the present invention in conjunction with the accompanying drawings and specific implementation manners:

[0077] In the first step, a data set is constructed using the original vibration signals collected by sensors at different positions, divided into a training set, a test set, and a validation set, and normalized. The specific steps are as follows:

[0078] S101: Collect sample data of the vibration signals of the wind turbine gearbox.

[0079] Specifically, 5000 label sample data are collected, the data length is 1024 data points, and the collected label sample data are the vibration signals of the wind turbine gearbox in normal state, planetary gear spalling, inner ring cracking of high-speed shaft and medium-speed shaft bearings. 1250 samples are collected for each type, and the data set is divided into a training set, a validation set, and a test set according to the ratio of 8:1:1.

[0080] S102: Normalize the training set, validation set, and test set, map the data values to between [0,1], accelerate the convergence of the model, and avoid gradient explosion or gradient disappearance. The specific formula is as follows:

[0081]

[0082] In the formula, is denoted as the data value after normalization, x is denoted as the original data value, min(x) is denoted as the minimum data value, max(x) is denoted as the maximum data value.

[0083] In the second step, the training set samples are input into the FWKC module to extract multi-scale pulse features, and the TCA attention mechanism is proposed for adaptive fusion to suppress redundant features. The specific steps are as follows:

[0084] S201: As Figure 1 shown, according to the specific adaptability of different wavelet basis functions to different types of pulse features, wavelet basis functions of different scales are used to replace the convolution kernels, and combined with the characteristics of continuous wavelet transform, multiple convolutions are implemented at different scales α and time translations β, so as to extract features in different frequency bands.

[0085] Specifically, three different scales of wavelet kernels, namely Laplace wavelet, Morlet wavelet and MexHat wavelet, are used to obtain feature maps p1, p2 and p3. The continuous wavelet transform and the convolution p formula of the continuous wavelet kernel are as follows:

[0086]

[0087] p = ψ α,β (τ) * x(τ)

[0088] In the formula, τ represents time, x(τ) represents the input signal, ψ represents the wavelet basis function, * represents the conjugate operation of complex numbers, α represents the scale parameter, and α represents the time translation parameter.

[0089] The convolution formulas of Laplace wavelet, Morlet wavelet and MexHat wavelet kernels are as follows:

[0090]

[0091] In the formula, A and C are the normalization coefficients of the wavelet basis function, ζ is the viscous damping ratio, f0 is the sampling frequency, and δ is the delay parameter.

[0092] S202: The feature maps output by the three wavelet kernel convolutions are concatenated according to the channel dimension to obtain the fused feature map p 123 , and its formula is as follows:

[0093] p 123 = Concat(ψ1(τ) * x(τ), ψ2(τ) * x(τ), ψ3(τ) * x(τ))

[0094]

[0095] In the formula, Concat represents concatenation.

[0096] S203: Use the TCA attention mechanism to perform feature fusion on the outputs of the three different scales of continuous wavelet kernel convolutions.

[0097] Specifically, as Figure 2 shown, through a lightweight convolutional network, in the time step dimension, channel weights w are dynamically generated for each time step 123 to effectively capture the changes in the signal at different time points and highlight the pulse features related to faults. Using the adaptive weights w 123 and p 123 to perform weighted summation, and extracting global spatial information through cross-channel global average pooling (GAP) to obtain the fused feature map p' 123 , and using the BP algorithm for backpropagation update. The specific formula is as follows:

[0098] w 123 = Sigmoid(Conv1d(p 123 ))

[0099]

[0100] In the formula, each feature map is regarded as a whole, N k represents the number of channels of the feature map, τ represents the number of time steps, c is the number of channels of the weighted feature map. The weighted feature map performs average pooling on the values of the channels at the spatial positions (N k , τ), and the dimension of the finally generated fused feature map is (N c , τ). This temporal weighting strategy effectively enhances the fusion effect of multi-scale features, thereby more accurately extracting the key pulse features in the temporal signal. k

[0101] Step 3: Input the output multi-scale pulse features into the MSAS module to further process the multi-scale information on the basis of realizing accurate signal denoising. The specific steps are as follows:

[0102] S301: Divide the MSAS module into two branches: Branch_1 and Branch_2; connect the two branches in parallel.

[0103] Specifically, as Figure 1 shown, Branch_1 uses two layers of narrow kernel convolutional layers (kernel size 1×3) to extract local information containing high-frequency features, and Branch_2 uses two layers of wide kernel convolutional layers (kernel size 7×7) to extract global information containing low-frequency features.

[0104] S302: Feed the fused feature map p 123 into Branch_1 and Branch_2 respectively, and obtain the feature map F(p) after two convolutional operations and ReLU activation.

[0105] S303: Obtain the feature map F'(p) through residual connection, retain a part of the original features, and add them to the extracted features. The specific formula is as follows:

[0106] F'(p) = F(p) + p 123

[0107] In the formula, the feature map after two convolutional operations and ReLU activation is F(p), and the output after residual connection is F'(p). The residual connection mechanism helps to maintain the information flow during training, prevent gradient vanishing or gradient explosion, and thus improve the stability of the network.

[0108] S304: Introduce the SE attention mechanism to dynamically adjust the soft threshold function.

[0109] Specifically, perform global average pooling (GAP) on F'(p) to calculate the global feature representation of each channel N, that is, the global average value Z of the absolute value of the features of each channel N , generate the attention weight of each channel through a fully connected layer. Usually, the first fully connected layer compresses the number of channels to a smaller dimension, and the second fully connected layer restores the compressed features to the original number of channels N. The attention coefficient S N of the Nth channel represents the importance of this channel. The specific formula is as follows:

[0110] S N = Sigmoid(W a ReLU(W b Z N ))

[0111] In the formula, W a and W b are the weight matrices of the two fully connected layers respectively, and the activation function ReLU is used to introduce non-linear transformation. The Sigmoid function normalizes the generated weights to between [0,1] for S N to obtain the Nth group of dynamic soft threshold λ N , and then adaptively adjust this threshold to denoise and shrink the features, improving the network's focusing ability on important features.

[0112] S305: Use the soft threshold function to perform dynamic denoising on the features of the two branches respectively to obtain the feature maps T1 and T2. The formula is as follows:

[0113] λ N = Z N S N

[0114] T N = sign(F N '(p))max(|F N'(p)|-λ N ,0)

[0115]

[0116] Wherein, T N represents the feature map after soft threshold denoising, and F N '(p) represents the feature of the Nth channel, and λ N represents the dynamic soft threshold.

[0117] S306: Linearly sum the two obtained feature maps to obtain the fused feature map T 12 , and its formula is as follows:

[0118] T 12 = T1 + T2

[0119] Wherein, T1 and T2 are the feature maps after denoising of branch_1 and branch_2 respectively, and T 12 is the fused feature map. This method can effectively retain multi-scale information and adaptively denoise different fault signals by dynamically adjusting the soft threshold.

[0120] Fourth step, perform phase space reconstruction on the labeled data set after noise reduction, and fine-tune the trained encoder. The specific steps are as follows:

[0121] S401: Extract the feature vectors corresponding to each time step and input them into the FD-Transformer encoder.

[0122] Specifically, as Figure 3 shown, for the subsequent classification task, only retain its encoder part. The fused feature map T 12 can be regarded as a feature matrix of [N i , L s , where N i is the number of channels, and L s is the sequence length. Adjust the dimension of the feature matrix to [L s , N i , extract the N i -dimensional feature vectors corresponding to each time step, and regard each feature vector as an embedding vector, with a total of L sThe embedded vectors are input into the FD-Transformer encoder. Through the multi-head self-attention mechanism, the features at each position (time step) of the input are weighted and summed to capture global dependencies. Through the MLP containing two fully connected layers, the features at each position are independently transformed to enhance the expression ability of the network. During the model calculation process, residual connections, layer normalization, and regularization processing are added to prevent overfitting caused by the overly complex network. Among them, the self-attention mechanism can effectively capture the relationships between inputs by calculating the influence of each position in the input sequence on other positions. In the multi-head self-attention mechanism, by processing multiple attention heads in parallel, the model's ability to capture different context information is further enhanced, thereby improving the model's expression ability and performance. Generally, the input sequence is set as where n e represents the length of the input sequence, and d represents the dimension of the sequence. Mathematically, for each self-attention head h, it is expressed as:

[0123]

[0124] In the formula, for each attention head h, three different weights are used for linear transformation to generate Q h query matrix, K h key matrix, V h value matrix. Q h and K h are dot-product calculated to obtain the similarity matrix To prevent the gradient disappearance of the Softmax function due to the too large dot-product result, is used for scaling, where d k represents the dimension of the query vector and the key vector. At the same time, the softmax function is used to convert the similarity matrix into a probability distribution, making the sum of each row equal to 1, representing the attention weight distribution relative to each query. According to the obtained attention weights, is weighted and summed to generate the final output matrix, where d v represents the dimension of V h . The output matrices of each self-attention head h are concatenated together to obtain the final result C. The specific formula is defined as follows:

[0125] C = [Attention1,Attention2,...Attention h W O

[0126] Among them, the attention results of each head are concatenated together and linearly transformed through the weight to map the concatenated features back to the original feature space;

[0127] Typically, 1D Transformer or 2D Transformer uses class tokens to segment input data and generate an independent embedding representation for each part. Each token represents the information of an independent part. However, in this patent, the input data has been converted into an embedding vector containing multi-channel feature maps, and these vectors are essentially already the embedding representations of different parts, so there is no need to add additional tokens or embedding operations. As for positional encoding, in the fault diagnosis task of vibration signals, positional encoding is not necessary. Therefore, this patent does not add positional encoding to the model.

[0128] S402: Stack three FD-Transformer encoders.

[0129] S403: Reduce the output feature channels to half of the input features.

[0130] Specifically, input the flattened feature matrix into the MLP classifier for classification.

[0131] S404: Flatten the output feature matrix.

[0132] The fifth step is to input the extracted features into the classifier for classification, fine-tune the model parameters using the validation set, and conduct tests through the test set to obtain the fault diagnosis results. The specific steps are as follows:

[0133] S501: Input the extracted features into the MLP classifier with three linear layers for classification. The dimension of linear layer 1 is 128, the dimension of linear layer 2 is 32, the dimension of linear layer 3 is the number of classifications, and the activation function is the softmax function.

[0134] S502: Repeat the iteration until the training loss is the lowest.

[0135] S503: Input the validation set into the trained diagnostic model to fine-tune the parameters.

[0136] S504: Input the test set into the tuned model to obtain the diagnostic results and complete the fault diagnosis.

[0137] Embodiment:

[0138] The experimental model of the present invention runs on Python 3.11, Windows 11, Intel(R) Core(TM) i7-13700H, RTX 4060 GPU. And the publicly available datasets: Case Western Reserve University (CWRU) dataset and real wind turbine (WT) dataset are used for evaluation and verification.

[0139] Collect 5000 large numbers of labeled sample data with a data length of 1024 data points. The collected labeled sample data are vibration signals of the gearbox of a wind turbine in normal state, planetary gear spalling, and inner ring cracking of high-speed and medium-speed shaft bearings. 1250 samples are collected for each type, constituting the real WT dataset. The dataset is divided into a training set, a validation set, and a test set according to the ratio of 8:1:1, and normalization processing is performed. The specific descriptions of the WT dataset and the CWRU dataset are shown in Tables 1 and 2 respectively.

[0140] Table 1 Description of WT Dataset

[0141]

[0142] Table 2 Description of CWRU Dataset

[0143]

[0144]

[0145] To verify the performance of the method involved in the present invention, the experiment will be carried out on the WT dataset for comparative evaluation. The number of iterations is 30 times. And to reduce the contingency of the experiment, each group of experiments is repeated 10 times, and the average value is taken as the reference for the final result. The multi-core wavelet denoising network (MKWDN) of the method involved in the present invention will be compared with the following models: supervised learning models: CNN with three convolutional layers, MSCNN that can process multi-scale features, DRSN-CW with adaptive denoising; unsupervised models: SimCLR, BYOL. The experimental results are shown in Table 3. MKWDN achieves the best results in terms of accuracy, precision, and F1 score, further verifying the effectiveness of the method involved in the present invention.

[0146] Table 3 Accuracy of Different Models on WT Dataset

[0147]

[0148] To verify the anti-noise robustness of the method involved in the present invention, the experiment is designed based on the CWRU dataset. The number of iterations is 30 times. And to reduce the contingency of the experiment, each group of experiments is repeated 10 times, and the average value is taken as the reference for the final result. Gaussian white noise and pink noise with a signal-to-noise ratio from -6 dB to 6 dB are added to the data. The experimental results are as Figure 4As shown in the figure, the 1D-CNN performs the worst, mainly because this network lacks feature enhancement and denoising mechanisms and cannot effectively extract fault features in a high-noise background. Although the MSCNN can extract multi-scale features, the overall effect still fails to meet the expectations. In contrast, due to the adaptive denoising function of the DRSN-CW, the accuracy still reaches 92.79% when -6dB pink noise is added, but it still fails to meet the requirements. SimCLR and BYOL improve the feature extraction ability through data augmentation strategies, enabling the model to have a certain anti-noise performance. However, as the noise intensity increases, the accuracy shows an obvious downward trend. Based on enhancing the feature extraction ability using the FWKC module, this patent introduces the MSAS module for denoising processing. Therefore, it shows the best denoising effect under both noise conditions, further verifying the robustness of the MKWDN and enabling it to successfully extract weak fault features in a high-noise environment.

[0149] To further evaluate the performance of the model, this experiment uses the CWRU dataset and visualizes it using the t-SNE algorithm in the first training cycle and the last training cycle. 100 samples of each fault type are visualized. The experimental results are as Figure 5 shown. After 30 iterations, the model can already complete the classification task well.

[0150] To verify the attention of the FD-Transformer encoder to fault features, this section of the experiment selects the OR021 fault type in the CWRU dataset and visualizes the scores of each attention head in the last layer of the FD-Transformer encoder and maps them to the original signal. Since the fault features have been effectively extracted before inputting into the FD-Transformer encoder, the complexity of fault diagnosis is reduced when using the FD-Transformer encoder. To reduce the computational burden of the model, the experiment selects four attention heads for analysis. The experimental results are as Figure 6As shown, different attention heads have different regions of interest in pulse features. The first attention head mainly focuses on key fault features, while the second and third attention heads fail to effectively focus on these fault features. Although the fourth attention head pays attention to pulse features, it does not focus on the main fault features. This is because the multi-head self-attention mechanism focuses on features from different perspectives. Since the input of the FD-Transformer comes from the features extracted by the previous network structure, the first attention head can effectively focus on the main fault features, which also indirectly reflects the feature extraction ability of the previously described module. To verify the model's ability to process long signals, subsequent experiments will conduct a visual analysis of the attention of the first attention head on OR021 samples of different lengths (sample lengths are 512, 1024, 2048). As Figure 7 shown, the first attention head successfully focuses on the main fault features in all three samples of different lengths. The experimental results show that the model can effectively capture the long-span dependency information in time series signals, and as the signal length increases, the model performance does not significantly decline. In practical engineering applications, the method involved in the present invention can process long signals and analyze them according to their attention, which helps to improve the accuracy of fault diagnosis.

[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; thus, these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope defined by the claims of the present invention.

Claims

1. A fault diagnosis method for a fan gearbox based on a multi-core wavelet denoising network, characterized in that: It includes the following steps: First step, construct a data set using the original vibration signals collected by sensors at different positions, divide it into a training set, a test set and a validation set, and perform normalization processing; Second step, input the training set samples into the Fusion Wavelet Kernel Convolution (FWKC) module to extract multi-scale pulse features, and propose a Temporal Channel Attention mechanism (TCA) for adaptive fusion to suppress redundant features; Third step, input the output multi-scale pulse features into the Multi-Scale Adaptive Soft Threshold (MSAS) module to further process multi-scale information on the basis of realizing accurate signal denoising; Fourth step, adopt a Feature-Driven Transformer (FD-Transformer) encoder to capture the global information of the samples using the multi-head attention mechanism, and further extract global features; Fifth step, input the extracted features into a classifier for classification, fine-tune the model parameters using the validation set, and perform tests through the test set to obtain the fault diagnosis results.

2. The fault diagnosis method of a fan gearbox based on a multi-core wavelet denoising network according to claim 1, wherein: The specific steps of the first step include: Step 1: Collect the vibration signal sample data of the wind turbine gearbox: Specifically, collect a large number of labeled sample data. The collected labeled sample data are the vibration signals of the wind turbine gearbox in normal state, planetary gear spalling, inner ring cracking of high-speed shaft and medium-speed shaft bearings, and divide them into a training set, a validation set, and a test set; Step 2: Perform normalization processing on the training set, validation set, and test set. The specific formula is as follows: In the formula, is denoted as the data value after normalization processing, \(x\) is denoted as the original data value, \(\min(x)\) is denoted as the minimum data value, and \(\max(x)\) is denoted as the maximum data value.

3. A fault diagnosis method for a fan gearbox based on a multi-core wavelet denoising network according to claim 1, characterized in that: The specific steps of the second step include: Step 1: Use wavelet basis functions instead of convolution kernels, and combine the characteristics of continuous wavelet transform to perform multiple convolutions at different scales α and time shift β: Specifically, use three different scales of wavelet kernels, namely Laplace wavelet, Morlet wavelet, and MexHat wavelet, to obtain feature maps p1, p2, and p3. The continuous wavelet transform and the continuous wavelet kernel convolution p formula are as follows: p = ψ α,β (τ) * x(τ) In the formula, τ represents time, x(τ) represents the input signal, ψ represents the wavelet basis function, * represents the conjugate operation of complex numbers, α represents the scale parameter, and α represents the time translation parameter; The convolution formulas of Laplace wavelet, Morlet wavelet, and MexHat wavelet kernels are as follows: In the formula, A and C are the normalization coefficients of the wavelet basis function, ζ is the viscous damping ratio, f0 is the sampling frequency, and δ is the delay parameter; Step 2: Concatenate the feature maps output by the three wavelet kernel convolutions along the channel dimension to obtain the fused feature map p 123 , as shown in the following formula: p 123 = Concat(ψ1(τ)*x(τ), ψ2(τ)*x(τ), ψ3(τ)*x(τ)) In the formula, Concat represents concatenation; Step 3: Use the TCA attention mechanism to perform feature fusion on the outputs of the continuous wavelet kernel convolutions of three different scales; Specifically, through a lightweight convolutional network, the channel weights w are dynamically generated for each time step in the time step dimension. 123 The adaptive weights w 123 are used to perform weighted summation with p 123 , and global spatial information is extracted through cross-channel global average pooling (GAP) to obtain the fused feature map p'. 123 The BP algorithm is used for backpropagation update, and the specific formula is as follows: w 123 = Sigmoid(Conv1d(p 123 )) In the formula, each feature map is regarded as a whole, N k represents the number of channels of the feature map, τ represents the number of time steps, c is the number of channels of the weighted feature map, and the weighted feature map performs average pooling on the value of channel c at the spatial position (N k , τ), and the dimension of the finally generated fused feature map is (N k , τ).

4. A fault diagnosis method for a fan gearbox based on a multi-core wavelet denoising network according to claim 1, characterized in that: The specific steps of the third step include: Step 1: Divide the MSAS module into two branches: Branch_1 and Branch_2; connect the two branches in parallel: Specifically, Branch_1 uses two layers of narrow kernel convolutional layers (kernel size 1×3) to extract local information containing high-frequency features, and Branch_2 uses two layers of wide kernel convolutional layers (kernel size 7×7) to extract global information containing low-frequency features; Step 2: Send the fused feature map p 123 into Branch_1 and Branch_2 respectively, and obtain the feature map F(p) after two-layer convolution operation and ReLU activation; Step 3: Obtain the feature map F'(p) through residual connection. The specific formula is as follows: F'(p) = F(p) + p 123 In the formula, the feature map after two convolutional operations and ReLU activation is F(p), and the output after residual connection is F'(p); Step 4: Introduce the SE attention mechanism to dynamically adjust the soft threshold function: Specifically, perform global average pooling (GAP) on F'(p) to calculate the global feature representation of each channel N, that is, the global average value Z of the absolute value of the features of each channel N , generate the attention weight of each channel through a fully connected layer, and the attention coefficient S of the Nth channel N represents the importance of this channel, and the specific formula is as follows: S N = Sigmoid(W a ReLU(W b Z N )) where W a and W b are the weight matrices of the two fully connected layers respectively, the activation function ReLU is used to introduce non-linear transformation, and the Sigmoid function normalizes the generated weights to between [0, 1], which is used for S N to obtain the Nth group of dynamic soft threshold λ N ; Step 5: Use the soft threshold function to dynamically denoise the features of the two branches to obtain feature maps T1 and T2, and the formula is as follows: λ N = Z N S N T N = sign(F N '(p)) max(|F N '(p)| - λ N , 0) Where, T N represents the feature map after soft threshold denoising, F N '(p) represents the feature of the Nth channel, λ N represents the dynamic soft threshold; Step 6: Linearly sum the two obtained feature maps to obtain a fused feature map T 12 , and its formula is as follows: Wherein, T1 and T2 are the denoised feature maps of branch_1 and branch_2 respectively, and T 12 is the fused feature map.

5. A fault diagnosis method for a fan gearbox based on a multi-core wavelet denoising network according to claim 1, characterized in that: The specific steps of the fourth step include: Step 1: Extract the feature vectors corresponding to each time step and input them into the FD-Transformer encoder; Step 2: Stack three FD-Transformer encoders; Step 3: Reduce the output feature channels to half of the input features; Step 4: Flatten the output feature matrix.

6. The fault diagnosis method of a fan gearbox based on a multi-core wavelet denoising network according to claim 1, characterized in that: The specific steps of the fifth step include: Step 1: Input the flattened feature matrix into the MLP classifier for classification; Step 2: Repeat the iteration until the training loss is the lowest; Step 3: Input the validation set into the trained diagnostic model to fine-tune the parameters; Step 4: Input the test set into the tuned model to obtain the diagnostic result and complete the fault diagnosis.

7. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Analog circuit early fault diagnosis method based on subsequence division and Transform

    CN120951916A

  • Fault detection method and device for high-power pulse source based on artificial intelligence

    CN121540978A

  • Artificial intelligence-based fault detection method and device for high-power pulse sources

    CN121540978B

  • Agricultural machinery wet clutch gear shifting system fault diagnosis method based on closed continuous time unit

    CN121743679A