An intelligent rotating machinery fault diagnosis method
Through the progressive processing framework of feature extraction-channel feature enhancement-time sequence modeling-multi-scale feature fusion and the multi-scale multi-head attention module, combined with the beluga optimization algorithm and the SE-CNN-BiLSTM-MSMHA network, the fault diagnosis problem of rotating machinery in a strong noise environment is solved, and efficient and flexible fault identification is achieved under complex operating conditions.
Patent Information
- Application Number
- CN202510873852.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-06-27
AI Technical Summary
In a highly noise environment, the fault signal of the rotating machinery is easily masked, resulting in difficulty in feature extraction and identification. Traditional diagnostic models have feature mismatch problems under multiple operating conditions such as variable speed and variable load, making it difficult to achieve accurate fault diagnosis.
A progressive processing framework with feature extraction-channel feature enhancement-time sequence modeling-multi-scale feature fusion and a multi-scale multi-head attention module with parameter adaptation is adopted, and a multi-scale multi-head attention module is optimized by white whale optimization algorithm and a SE-CNN-BiLSTM-MSMHA attention enhancement network are used to achieve fault diagnosis through transfer learning.
It improves the accuracy and robustness of fault diagnosis in strong noise environments, and can accurately identify fault types under complex operating conditions, solving the problem of feature mismatch and noise interference in traditional methods.
Smart Images

Figure CN120387056B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mechanical fault diagnosis, and in particular to an intelligent rotating machinery fault diagnosis method. Background Art
[0002] Failures in rotating machinery can impact normal production and reduce product quality, or even lead to significant economic losses and serious safety and casualties. Due to the harsh operating environment, fault signal characteristics are often masked by strong noise, making feature extraction and identification difficult. Furthermore, in real industrial scenarios, rotating machinery often operates under multiple operating conditions, such as variable speeds and loads. Differences in feature distributions across different operating conditions can cause feature mismatches in traditional diagnostic models. Therefore, timely and accurate diagnosis of cross-operating faults in rotating machinery under strong noise is of great practical significance.
[0003] In real-world industrial scenarios, rotating machinery often operates under multiple operating conditions, including variable speeds and loads. Differences in feature distribution across different operating conditions can lead to feature mismatch in traditional diagnostic models. Research has shown that when operating conditions vary by more than 15% from the rated parameters, the false alarm rate of existing fixed-threshold diagnostic methods rises to over 30%. This lack of adaptability across operating conditions, coupled with strong noise interference, further complicates fault feature extraction and pattern recognition. Therefore, research into advanced signal noise reduction techniques is needed to improve fault diagnosis accuracy.
[0004] In industrial environments, rotating machinery vibration signals are often subject to complex noise interference. This is especially true in extreme operating conditions, such as wind turbine gearboxes and mining machinery, where background noise energy can completely mask fault signatures. According to consensus in the field of mechanical vibration diagnosis, when the full-band signal-to-noise ratio (SNR) of a vibration signal falls below 0 dB (i.e., noise power ≥ signal power), traditional time-frequency analysis methods (such as envelope spectrum and fast Fourier transform) fail due to their inability to effectively separate noise from fault signatures. Furthermore, rotating machinery often operates under multiple operating conditions, such as variable speed and load. Differences in characteristic distributions across different operating conditions can lead to insufficient cross-condition adaptability of traditional diagnostic models. This inter-condition adaptability, coupled with strong noise interference, further complicates fault feature extraction and pattern recognition. Therefore, research is needed on advanced signal denoising techniques, fault feature extraction, and intelligent bearing fault diagnosis theories and transfer learning strategies. This will continuously enrich the theoretical and key technologies for bearing health monitoring and fault diagnosis, thereby improving the accuracy and robustness of cross-condition fault diagnosis methods in strong noise environments. Summary of the Invention
[0005] The main purpose of the present invention is to solve the problem that fault signals in rotating machinery vibration data measured in a strong noise environment are easily masked, and fault feature extraction and identification are difficult. An intelligent rotating machinery fault diagnosis method is provided, which adopts a progressive processing framework of "feature extraction-channel feature enhancement-time series modeling-multi-scale feature fusion" and a "parameter-adaptive multi-scale multi-head attention module" to achieve efficient and flexible fault diagnosis through modular design. At the same time, the model architecture proposed in the present invention shows significant advantages in bearing fault diagnosis: the feature extraction-channel feature enhancement mechanism improves feature characterization capability, multi-scale time series modeling enhances adaptability to complex working conditions, and the parameter-adaptive module ensures strong noise resistance. It has excellent engineering practical value in the diagnosis of rotating machinery faults in a strong noise environment, thereby providing a scientific basis for the diagnosis of rotating machinery faults in a strong noise environment and providing reliable support for protecting people's lives and property.
[0006] The technical solution to achieve the purpose of the present invention is: an intelligent rotating machinery fault diagnosis method, comprising the following steps:
[0007] Step 1: Obtain vibration signals of rotating machinery under source domain working conditions and target domain working conditions;
[0008] Step 2: The vibration signals of the source and target domains are subjected to noise reduction processing using a noise reduction processing module. The noise reduction processing module consists of two submodules: a variational mode decomposition module optimized by the Beluga optimization algorithm, and a variational mode decomposition screening and reconstruction module using the ideal solution sorting method.
[0009] Step 3: Perform short-time Fourier transform on the denoised vibration signal to obtain a two-dimensional time-frequency image. The source domain dataset is generated based on the two-dimensional time-frequency image obtained from the source domain working condition, and the target domain dataset is generated based on the two-dimensional time-frequency image obtained from the target domain working condition.
[0010] Step 4: Build an attention enhancement network model based on SE-CNN-BiLSTM-MSMHA, use the source domain dataset to complete the pre-training of the network model, and obtain the source domain working condition model;
[0011] Step 5: Using transfer learning, freeze all network model parameters before the first bidirectional long short-term memory layer in the source domain working condition model. Fine-tune the parameters of the source domain working condition model using a small amount of data samples from the target domain dataset to obtain the target domain working condition model.
[0012] Step 6: Using the methods of steps 2 and 3, perform noise reduction on the vibration signal to be diagnosed, obtain a two-dimensional time-frequency image, and then input the time-frequency image into the fault diagnosis model corresponding to the working condition of the vibration signal to achieve fault diagnosis.
[0013] Furthermore, in step 2, in the noise reduction processing module, the White Whale optimization algorithm optimizes the variational mode decomposition submodule, and the decomposition steps are as follows:
[0014] Step S21, set the penalty factor α and the number of decomposition layers N The search range of these two parameters;
[0015] Step S22: Initialize the position of the beluga whale group, set the minimum envelope entropy as the fitness function, and then calculate the fitness value of each individual whale;
[0016] Step S23, iteratively update the position of the whale group based on the white whale optimization algorithm until the preset termination condition is met, and output the one with the best fitness value. N The optimal parameter combination of and α;
[0017] Step S24: Based on the optimal parameter combination, the vibration signal is subjected to VMD decomposition to obtain N IMF components;
[0018] The ideal solution sorting method selects and reconstructs submodules for variational mode decomposition. The reconstruction steps are as follows:
[0019] Step S25: N eigenmode components, when N When it is an even number, it is removed by the center frequency N / 2 high-frequency components; when N When it is an odd number, it is removed by the center frequency ( N +1) / 2 high frequency components;
[0020] Step S26, calculating three indices of the remaining eigenmode components: envelope entropy, kurtosis index, and correlation coefficient with the input signal;
[0021] Step S27, calculating the comprehensive scores of the three indicators using the ideal solution sorting method, and removing the component with the smallest score;
[0022] Step S28: The remaining eigenmode components are added and reconstructed as the denoised signal.
[0023] Furthermore, in step 4, the attention enhancement network of SE-CNN-BiLSTM-MSMHA has the following structure: first, the construction of the model starts from the input layer, and passes through three consecutive convolutional layers, a squeeze and excitation module, the fourth convolutional layer, a pooling layer, a flattening layer and the first bidirectional long short-term memory layer; secondly, it is reshaped into a one-dimensional sequence through the reshape layer and connected to the multi-scale multi-head attention module; thirdly, it is flattened by connecting a flattening layer, and after flattening, it passes through two fully connected layers with ReLU activation function and a layer with softmaxThe output layer of the activation function performs fault classification and identification;
[0024] The multi-scale multi-head attention module adopts a three-channel parallel architecture, including: first, constructing a layer group containing three sets of parallel bidirectional long short-term memory networks, where each bidirectional long short-term memory layer has a different number of hidden neurons and all enable full sequence output mode; second, each bidirectional long short-term memory layer is connected to a parameter-adaptive multi-head attention module. In the multi-head attention module, the number of attention heads is fixed to 4, and the key vector dimension is dynamically calculated. The key vector dimension = the number of hidden units / the number of attention heads; finally, a deep splicing strategy is used to fuse the features of the three sets of attention outputs to form three parallel processing streams. The mathematical formula of the multi-scale multi-head attention module is as follows:
[0025] First are three bidirectional long short-term memory networks:
[0026] H i =BiLSTM i ( X )
[0027] Among them, BiLSTM i For the module i BiLSTM channels, X is the input sequence, H i is the output sequence;
[0028] Next, we construct the query, key, value matrix and attention score calculation formula for the subsequent attention mechanism;
[0029] Query Matrix Q i,j :
[0030] Q i,j = H i W Q (i,j)
[0031] Bond Matrix K i,j :
[0032] K i,j = H i W K (i,j)
[0033] Value Matrix V i,j :
[0034] V i,j = H i W V (i,j)
[0035] Attention score Attention i,j :
[0036] Attention i,j =V i,j × softmax(Q i,j × K i,j ) / ( d k ( i )) 1 / 2
[0037] Multi-head attention score MultiHead(i) :
[0038] MultiHead(i)=Concat(Attention i,1 ,Attention i,2 ,Attention i,3 ,Attention i,4 ) W o
[0039] in, Concat () is the splicing transformation function; j Representative j An attention head, j =(1,2,3,4); W Q (i,j) , W K (i ,j) , W V (i,j) It is i BiLSTM channels, j Three learnable weight matrices for each attention head; d k ( i ) is an adaptive parameter, indicating that the input i The output dimension / number of attention heads of a BiLSTM is used to scale the dot product to stabilize training; W ois a learnable weight matrix; the attention score is calculated by matrix multiplication of query and key, softmax After the function is normalized, it is multiplied by the value matrix to obtain the j The attention output of the four heads is then concatenated to obtain the i The multi-head attention output of each channel is obtained by splicing the multi-head attention output of the three channels to obtain the final multi-scale multi-head attention output sequence. The mathematical formula is: F fused =Concat(MultiHead(1), MultiHead(2), MultiHead (3)) .
[0040] Furthermore, in step five, freezing all network model parameters before the first bidirectional long short-term memory layer in the source domain working condition model includes: freezing all parameters of the first three convolutional layers, the SE module, the fourth convolutional layer, the pooling layer, and the flattening layer in the source domain working condition model. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0042] Figure 1 This is a rotating machinery fault diagnosis flow chart provided by the present invention;
[0043] Figure 2 This is the structure diagram of the attention enhancement network (SE-CNN-BiLSTM-MSMHA);
[0044] Figure 3 This is the classification confusion matrix diagram of the attention enhancement network model using the transfer learning strategy in working condition 1;
[0045] Figure 4 This is the classification confusion matrix diagram of the attention enhancement network model using the transfer learning strategy in working condition 2;
[0046] Figure 5 This is the classification confusion matrix diagram of the attention enhancement network model using the transfer learning strategy in working condition 3. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in combination with the implementation cases and drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention for which protection is sought, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0048] like Figure 1 , an intelligent rotating machinery fault diagnosis method, comprising the following steps:
[0049] Step 1: Obtain vibration signals of rotating machinery under source domain working conditions and target domain working conditions;
[0050] Step 2: The vibration signals of the source and target domains are subjected to noise reduction processing using a noise reduction processing module. The noise reduction processing module consists of two submodules: a variational mode decomposition module optimized by the Beluga optimization algorithm, and a variational mode decomposition screening and reconstruction module using the ideal solution sorting method.
[0051] Step 3: Perform short-time Fourier transform on the denoised vibration signal to obtain a two-dimensional time-frequency image. The source domain dataset is generated based on the two-dimensional time-frequency image obtained from the source domain working condition, and the target domain dataset is generated based on the two-dimensional time-frequency image obtained from the target domain working condition.
[0052] Step 4: Build an attention enhancement network model based on SE-CNN-BiLSTM-MSMHA, use the source domain dataset to complete the pre-training of the network model, and obtain the source domain working condition model;
[0053] Step 5: Using transfer learning, freeze all network model parameters before the first bidirectional long short-term memory layer in the source domain working condition model. Fine-tune the parameters of the source domain working condition model using a small amount of data samples from the target domain dataset to obtain the target domain working condition model.
[0054] Step 6: Using the methods of steps 2 and 3, perform noise reduction on the vibration signal to be diagnosed, obtain a two-dimensional time-frequency image, and then input the time-frequency image into the fault diagnosis model corresponding to the working condition of the vibration signal to achieve fault diagnosis.
[0055] Furthermore, in step 2, in the noise reduction processing module, the White Whale optimization algorithm optimizes the variational mode decomposition submodule, and the decomposition steps are as follows:
[0056] Step S21, set the penalty factor α and the number of decomposition layers N The search range of these two parameters;
[0057] Step S22: Initialize the position of the beluga whale group, set the minimum envelope entropy as the fitness function, and then calculate the fitness value of each individual whale;
[0058] Step S23, iteratively update the position of the whale group based on the white whale optimization algorithm until the preset termination condition is met, and output the one with the best fitness value. N The optimal parameter combination of and α;
[0059] Step S24: Based on the optimal parameter combination, the vibration signal is subjected to VMD decomposition to obtain N IMF components;
[0060] The ideal solution sorting method selects and reconstructs submodules for variational mode decomposition. The reconstruction steps are as follows:
[0061] Step S25: N eigenmode components, when N When it is an even number, it is removed by the center frequency N / 2 high-frequency components; when N When it is an odd number, it is removed by the center frequency ( N +1) / 2 high frequency components;
[0062] Step S26, calculating three indices of the remaining eigenmode components: envelope entropy, kurtosis index, and correlation coefficient with the input signal;
[0063] Step S27, calculating the comprehensive scores of the three indicators using the ideal solution sorting method, and removing the component with the smallest score;
[0064] Step S28: The remaining eigenmode components are added and reconstructed as the denoised signal.
[0065] Furthermore, in step 4, the attention enhancement network of SE-CNN-BiLSTM-MSMHA has the following structure: first, the construction of the model starts from the input layer, and passes through three consecutive convolutional layers, a squeeze and excitation module, the fourth convolutional layer, a pooling layer, a flattening layer and the first bidirectional long short-term memory layer; secondly, it is reshaped into a one-dimensional sequence through the reshape layer and connected to the multi-scale multi-head attention module; thirdly, it is flattened by connecting a flattening layer, and after flattening, it passes through two fully connected layers with ReLU activation function and a layer with softmax The output layer of the activation function performs fault classification and identification;
[0066] The SE-CNN-BiLSTM-MSMHA attention-enhanced network uses a progressive processing framework of "feature extraction-channel feature enhancement-time series modeling-multi-scale feature fusion" and a "parameter-adaptive multi-scale multi-head attention module" to automatically learn and extract fault information from raw vibration signals, enabling efficient and flexible fault diagnosis. This model framework also significantly improves the model's ability to resist noise interference, maintaining accurate fault diagnosis performance even in high-noise environments.
[0067] The attention enhancement network of SE-CNN-BiLSTM-MSMHA adopts a collaborative strategy of convolutional layers and SE modules. It achieves spatial downsampling of two-dimensional time-frequency images through three-level convolution while maintaining the integrity of local time-frequency features. The SE module is connected after the third convolution layer, which captures channel statistics through global average pooling. Combined with the fully connected layer, it generates channel attention weights and dynamically recalibrates multi-channel features to strengthen key fault characteristics and suppress noise interference. After the SE module, the fourth convolution layer is connected to perform high-order nonlinear transformation on the recalibrated channel features to ensure that the features are fully screened and enhanced before entering the BiLSTM layer in the time series modeling stage. This avoids the problem of invalid features interfering with time series modeling in the traditional CNN-BiLSTM architecture.
[0068] The multi-scale multi-head attention module adopts a three-channel parallel architecture, including: first, constructing a layer group containing three sets of parallel bidirectional long short-term memory networks, where each bidirectional long short-term memory layer has a different number of hidden neurons and all enable full sequence output mode; second, each bidirectional long short-term memory layer is connected to a parameter-adaptive multi-head attention module. In the multi-head attention module, the number of attention heads is fixed to 4, and the key vector dimension is dynamically calculated. The key vector dimension = the number of hidden units / the number of attention heads; finally, a deep splicing strategy is used to fuse the features of the three sets of attention outputs to form three parallel processing streams. The mathematical formula of the multi-scale multi-head attention module is as follows:
[0069] First are three bidirectional long short-term memory networks:
[0070] H i =BiLSTM i ( X )
[0071] Among them, BiLSTM i For the module i BiLSTM channels, X is the input sequence, H i is the output sequence;
[0072] Next, we construct the query, key, value matrix and attention score calculation formula for the subsequent attention mechanism;
[0073] Query Matrix Q i,j :
[0074] Q i,j = H i W Q (i,j)
[0075] Bond Matrix K i,j :
[0076] K i,j = H i W K (i,j)
[0077] Value Matrix V i,j :
[0078] V i,j = H i W V (i,j)
[0079] Attention score Attention i,j :
[0080] Attention i,j =V i,j × softmax(Q i,j × K i,j ) / ( d k ( i )) 1 / 2
[0081] Multi-head attention score MultiHead(i) :
[0082] MultiHead(i)=Concat(Attention i,1 ,Attention i,2 ,Attention i,3 ,Attention i,4 )W o
[0083] in, Concat () is the splicing transformation function; j Representative j An attention head, j =(1,2,3,4); W Q (i,j) , W K (i ,j) , W V (i,j) It is i BiLSTM channels, j Three learnable weight matrices for each attention head; d k ( i ) is an adaptive parameter, indicating that the input i The output dimension / number of attention heads of a BiLSTM is used to scale the dot product to stabilize training; W o is a learnable weight matrix; the attention score is calculated by matrix multiplication of query and key, softmax After the function is normalized, it is multiplied by the value matrix to obtain the j The attention output of the four heads is then concatenated to obtain the i The multi-head attention output of each channel is obtained by splicing the multi-head attention output of the three channels to obtain the final multi-scale multi-head attention output sequence. The mathematical formula is: F fused =Concat(MultiHead(1), MultiHead(2), MultiHead (3)) ;
[0084] The multi-scale multi-head attention module adopts a dynamic adaptive multi-scale architecture strategy, in which the number of hidden units in the three-channel BiLSTM is configured with different numbers of neurons, forming a multi-scale observation system with fine-medium-coarse time granularity. This design can simultaneously capture local state transition characteristics and long-range dependencies; the adaptive Key dimension design, Key vector dimension = number of hidden units / number of attention heads, makes the parameter scale of the multi-head attention module of each channel self-match with the feature complexity, can capture feature details at different scales, and ensure that the model can fully utilize multi-scale information; the combination of the two greatly improves the robustness and generalization ability of the model under complex working conditions, enabling it to accurately diagnose fault types under complex working conditions.
[0085] Furthermore, in step five, all network model parameters before the first bidirectional long short-term memory layer in the source domain working condition model are frozen, including: freezing all parameters of the first three convolutional layers, SE module, fourth convolutional layer, pooling layer, and flattening layer in the source domain working condition model; a layered freezing strategy is adopted to retain the general time-frequency feature extraction capabilities learned in the source domain data, such as basic features such as edges and textures, to avoid the degradation of low-level features caused by insufficient target domain data, and to allow the BiLSTM layer and its subsequent MSMHA module to adjust parameters, so that the network model can adaptively learn the timing dynamic features under different working conditions, especially to capture periodic fault features in vibration signals; through the layered freezing strategy, the model maintains the stability of feature extraction while enhancing the adaptability of timing modeling, and solves the "catastrophic forgetting" problem caused by global parameter adjustment in traditional transfer learning, so that the diagnostic accuracy remains stable when the working conditions change.
[0086] The method of the present invention is described below using a specific rotating machinery failure case as an example.
[0087] Example 1: Taking the CWRU dataset as an example, the data information of the CWRU dataset is shown in Table 1.
[0088]
[0089] The specific implementation process is as follows:
[0090] (1) Obtain the vibration signals of the rotating machinery under the source domain working condition (working condition 0) and the target domain working conditions (working conditions 1, 2, and 3). In order to simulate the strong noise environment in actual work, Gaussian noise with an intensity of 15 dB is added to the obtained vibration signals;
[0091] (2) The vibration signals of the source domain working condition and the target domain working condition are subjected to noise reduction processing through the noise reduction processing module. The noise reduction processing module includes two submodules: the variational mode decomposition module optimized by the Beluga optimization algorithm and the variational mode decomposition screening and reconstruction module using the ideal solution sorting method;
[0092] (3) Perform short-time Fourier transform on the denoised vibration signal to obtain a two-dimensional time-frequency image. The source domain dataset is generated by the two-dimensional time-frequency image obtained from the source domain working condition, and the target domain dataset is generated by the two-dimensional time-frequency image obtained from the target domain working condition.
[0093] (4) Construct an attention enhancement network model based on SE-CNN-BiLSTM-MSMHA, use the source domain dataset to complete the pre-training of the network model, and obtain the source domain working condition model. The network model structure is as follows: Figure 2 ;
[0094] (5) Through the transfer learning method, all network model parameters before the first bidirectional long short-term memory layer in the source domain working condition model are frozen, and the parameters of the source domain working condition model are fine-tuned using a small amount of data samples in the target domain dataset to obtain the target domain working condition model;
[0095] (6) Using the methods of step 2 and step 3, the vibration signal to be diagnosed is subjected to noise reduction processing and a two-dimensional time-frequency image is obtained. The time-frequency image is then input into the fault diagnosis model corresponding to the working condition of the vibration signal to achieve fault diagnosis. The noise reduction method, diagnosis model and its corresponding number information are shown in Table 2. The final diagnosis results are compared in Tables 3 and 4. The classification confusion matrix of the attention enhancement network model under each working condition is shown in Figure 3 、 Figure 4 、 Figure 5 .
[0096]
[0097]
[0098]
[0099] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. An intelligent rotating machinery fault diagnosis method, characterized in that: The following steps are involved: Step 1: Obtain vibration signals of rotating machinery under source domain working conditions and target domain working conditions; Step 2: The vibration signals of the source and target domains are subjected to denoising processing using a noise reduction processing module. This module consists of two submodules: a variational mode decomposition submodule optimized by the Beluga optimization algorithm, and a variational mode decomposition screening and reconstruction submodule using the ideal solution sorting method. Step 3: Perform short-time Fourier transform on the denoised vibration signal to obtain a two-dimensional time-frequency image. The source domain dataset is generated based on the two-dimensional time-frequency image obtained from the source domain working condition, and the target domain dataset is generated based on the two-dimensional time-frequency image obtained from the target domain working condition. Step 4: Build an attention enhancement network model based on SE-CNN-BiLSTM-MSMHA, use the source domain dataset to complete the pre-training of the network model, and obtain the source domain working condition model; The SE-CNN-BiLSTM-MSMHA attention enhancement network has the following structure: first, the model construction starts with the input layer, and passes through three consecutive convolutional layers, a squeeze and excitation module, a fourth convolutional layer, a pooling layer, a flattening layer, and the first bidirectional long short-term memory layer; second, it is reshaped into a one-dimensional sequence through a reshape layer and connected to a multi-scale multi-head attention module; Thirdly, a flattening layer is connected to flatten the network, and after flattening, two fully connected layers with ReLU activation function and an output layer with softmax activation function are used for fault classification and identification. Among them, the multi-scale multi-head attention module adopts a three-channel parallel architecture, including: first, constructing a layer group containing three groups of parallel bidirectional long short-term memory networks, where each bidirectional long short-term memory layer has a different number of hidden neurons and all enable full sequence output mode; secondly, each bidirectional long short-term memory layer is connected to a parameter-adaptive multi-head attention module. In the multi-head attention module, the number of attention heads is fixed to 4, and the key vector dimension is dynamically calculated. The key vector dimension = the number of hidden units / the number of attention heads; finally, a deep splicing strategy is used to fuse the features of the three groups of attention outputs to form three parallel processing streams; the mathematical formula of the multi-scale multi-head attention module is as follows: First are three bidirectional long short-term memory networks: H i =BiLSTM i (X) Among them, BiLSTM i is the i-th BiLSTM channel in the module, X is the input sequence, H i is the output sequence; Next, we construct the query, key, value matrix and attention score calculation formula for the subsequent attention mechanism; Query matrix Q i,j : Q i,j =H i W Q (i,j) Bond matrix K i,j : K i,j =H i W K (i,j) Value matrix V i,j : V i,j =H i W V (i,j) Attention score i,j : Attention i,j =V i,j ×softmax(Q i,j ×K i,j ) / (d k (i)) 1 / 2 MultiHead(i) multi-head attention score: MultiHead(i)=Concat(Attention i,1 ,Attention i,2 ,Attention i,3 ,Attention i,4 )W o Among them, Concat() is the concatenation transformation function; j represents the jth attention head, j = (1, 2, 3, 4); W Q (i,j) , W K (i,j) , W V (i,j) are the three learnable weight matrices of the i-th BiLSTM channel and the j-th attention head; d k (i) is an adaptive parameter that represents the output dimension / number of attention heads of the input i-th BiLSTM, which is used to scale the dot product to stabilize training; W o is a learnable weight matrix; the attention score is calculated by matrix multiplication of the query and key, normalized by the softmax function, and multiplied by the value matrix to obtain the attention output of the j-th head; then the attention outputs of the four heads are spliced to obtain the multi-head attention output of the i-th channel; finally, the multi-head attention outputs of the three channels are spliced to obtain the final multi-scale multi-head attention output sequence. The mathematical formula is: F fused =Concat(MultiHead(1),MultiHead(2),MultiHead(3)) Step 5: Using transfer learning, freeze all network model parameters before the first bidirectional long short-term memory layer in the source domain working condition model. Fine-tune the parameters of the source domain working condition model using a small amount of data samples from the target domain dataset to obtain the target domain working condition model. Step 6: Using the methods of steps 2 and 3, perform noise reduction on the vibration signal to be diagnosed, obtain a two-dimensional time-frequency image, and then input the time-frequency image into the fault diagnosis model corresponding to the working condition of the vibration signal to achieve fault diagnosis.
2. The intelligent rotating machinery fault diagnosis method according to claim 1, characterized in that: In step 2, in the noise reduction processing module, the White Whale optimization algorithm optimizes the variational mode decomposition submodule, and the decomposition steps are as follows: Step S21, setting the search range of the two parameters of penalty factor α and decomposition level number N; Step S22: Initialize the position of the beluga whale group, set the minimum envelope entropy as the fitness function, and then calculate the fitness value of each individual whale; Step S23, iteratively updating the position of the whale group based on the Beluga Whale Optimization Algorithm until the preset termination condition is met, and outputting the optimal parameter combination of N and α with the best fitness value; Step S24: performing VMD decomposition on the vibration signal based on the optimal parameter combination to obtain N IMF components; The ideal solution sorting method selects and reconstructs submodules for variational mode decomposition. The reconstruction steps are as follows: Step S25: for N eigenmode components, when N is an even number, remove N / 2 high-frequency components by the center frequency; when N is an odd number, remove (N+1) / 2 high-frequency components by the center frequency; Step S26, calculating three indices of the remaining eigenmode components: envelope entropy, kurtosis index, and correlation coefficient with the input signal; Step S27, calculating the comprehensive scores of the three indicators using the ideal solution sorting method, and removing the component with the smallest score; Step S28: The remaining eigenmode components are added and reconstructed as the denoised signal.
3. The intelligent rotating machinery fault diagnosis method according to claim 1, characterized in that: In step 5, freezing all network model parameters before the first bidirectional long short-term memory layer in the source domain working condition model includes: freezing all parameters of the first three convolutional layers, the SE module, the fourth convolutional layer, the pooling layer, and the flattening layer in the source domain working condition model.
Citation Information
Patent Citations
Mechanical equipment fault diagnosis method based on parallel network and transfer learning
CN115758212A
Fan bearing fault diagnosis method based on improved multi-head self-attention mechanism-BiLSTM
CN117786558A