Fault diagnosis method for main shaft bearing of steam turbine
Through the combination of noise adaptive random convolution blocks and global attention mechanism, the existing model has solved the problem of insufficient fault feature extraction and complex structure in a noisy environment, and achieved high-precision and lightweight fault diagnosis, which is suitable for the practical application of steam turbine spindle bearings.
Patent Information
- Application Number
- CN202510823936.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
Existing models cannot effectively extract weak key fault features and lack the ability to model nonlinear dynamic characteristics, resulting in low fault diagnosis accuracy, and complex model structures are difficult to deploy and apply in actual industrial scenarios.
The fault diagnosis method combining noise adaptive random convolution blocks and global attention mechanism is adopted to reduce noise sensitivity through noise adaptive random convolution blocks, and the global attention mechanism is used to enhance the attention ability of key fault characteristics, and simplify the model structure.
High-precision fault identification is achieved in high-noise environments, with few model parameters and high calculation efficiency, suitable for actual industrial scenarios, and has good robustness and adaptability.
Smart Images

Figure CN120333834A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of steam turbine fault diagnosis, and particularly relates to a fault diagnosis method for a steam turbine main shaft bearing. Background Art
[0002] As a key device in a thermal power generation unit, a steam turbine is an important link for converting the thermal energy generated by fuel combustion into mechanical energy. The stability of its operating state is directly related to the reliability and safety of the entire power generation system. Since steam turbines usually operate for a long time under complex and harsh working conditions such as high temperature and high pressure, key components including main shaft bearings are extremely vulnerable to wear, fatigue and other damages, resulting in a decline in equipment performance, component deterioration, and even failure shutdown events. Therefore, timely preventive maintenance and fault diagnosis of the main shaft bearing are of great significance for ensuring the safe operation of the equipment and improving power generation efficiency.
[0003] In the field of fault diagnosis, as a typical rotating machinery, the vibration signals generated by a steam turbine main shaft bearing during operation often contain rich fault feature information. Based on this, traditional rotating machinery fault diagnosis technologies mainly rely on vibration signal processing methods and artificial experience knowledge. Such methods usually use sparse pulse sequences as the main basis for fault identification, and extract and analyze the feature information in the signals. A large number of studies have been carried out on feature extraction means such as empirical mode decomposition, wavelet packet decomposition, and spectral kurtosis, and certain results have been achieved. However, these methods generally rely on a full understanding of the equipment degradation mechanism and a large amount of domain knowledge, and it is difficult to meet the processing requirements of multi-source and large-scale data under complex working conditions.
[0004] With the development of intelligent manufacturing and artificial intelligence technologies, data-driven deep learning methods have gradually become a research hotspot in fault diagnosis. These methods can automatically learn complex non-linear feature relationships from large-scale sample data without manual feature design, and show good robustness and adaptability under complex working conditions and big data backgrounds. Among them, structures such as convolutional neural networks (CNNs), graph neural networks (GNNs), and long short-term memory networks (LSTMs) have been widely applied to the fault diagnosis tasks of mechanical equipment, and have shown high diagnostic accuracy and generalization ability in multiple experimental scenarios. On this basis, to further improve the modeling ability of global features in time series signals, researchers have introduced the Transformer structure into the main shaft bearing fault diagnosis task. The Transformer structure can focus on the correlation between different time points in time series signals and has excellent global information extraction ability. Therefore, some studies have fused the CNN and Transformer structures to achieve complementary modeling of local features and global features, enhancing the model's adaptability to different working conditions while improving the fault diagnosis accuracy.
[0005] However, the existing methods still have the following limitations when applied to the fault diagnosis of steam turbine main shaft bearings: (1) The operation data collected in the actual industrial site is often interfered by strong noise, and the noise may mask the key weak fault characteristics, resulting in the model being unable to effectively extract the weak key fault characteristics; (2) The steam turbine main shaft bearing has high nonlinearity. The existing deep learning models have insufficient ability to model the nonlinear dynamic characteristics, are vulnerable to environmental interference, and the complex nonlinear dynamic characteristics will pose challenges to the convergence stability and training efficiency of the model, further affecting the diagnostic accuracy; (3) The structure of the existing deep learning models is complex, making it difficult to deploy and apply in the actual industrial scenario. Summary of the Invention
[0006] The purpose of the present invention is to solve the problems that the existing models cannot effectively extract the weak key fault characteristics and have insufficient ability to model the nonlinear dynamic characteristics, resulting in low fault diagnosis accuracy, and the complex structure of the existing models makes it difficult to deploy and apply in the actual industrial scenario. A fault diagnosis method for steam turbine main shaft bearings is proposed.
[0007] The technical solution adopted by the present invention to solve the above technical problems is: A fault diagnosis method for steam turbine main shaft bearings, the method specifically includes the following steps: Step 1: Collect the original vibration signal of the steam turbine main shaft bearing during operation, and then preprocess the collected original vibration signal to obtain the preprocessed vibration signal of the main shaft bearing; And divide the preprocessed vibration signal of the main shaft bearing into vibration signal segments of equal length according to the set time window length; Step 2: Obtain a vibration signal data matrix based on all the vibration signal segments, and construct a fault diagnosis model; The fault diagnosis model includes an input module, a feature extraction module and an output module; Among them, the input module sequentially includes an average pooling layer, a convolutional layer, a BN layer and a GELU activation function layer; The feature extraction module includes three feature extraction units, and each feature extraction unit includes a noise adaptive random convolution block, a global attention mechanism layer, a first LN layer, a feedforward layer and a second LN layer; The output module includes an average pooling layer and a linear layer; Step 3: Use the vibration signal data matrix as the input of the fault diagnosis model, and output the fault diagnosis result of the steam turbine main shaft bearing through the fault diagnosis model.
[0008] Further, an acceleration sensor is used to collect the original vibration signals of the steam turbine main shaft bearing during operation.
[0009] Further, the preprocessing of the collected original vibration signals is specifically as follows: The collected original vibration signals are sequentially detrended and normalized.
[0010] Further, the working process within the input module is as follows: The vibration signal data matrix is used as the input of the average pooling layer, and the output of the average pooling layer is used as the input of the convolutional layer; The output of the convolutional layer is used as the input of the BN layer, and the output of the BN layer is used as the input of the GELU activation function layer; The output of the GELU activation function layer is used as the output of the input module.
[0011] Further, the working process of the feature extraction module is as follows: The output of the input module is used as the input of the first feature extraction unit, the output of the first feature extraction unit is used as the input of the second feature extraction unit, the output of the second feature extraction unit is used as the input of the third feature extraction unit, and the output of the third feature extraction unit is used as the output of the feature extraction module.
[0012] Further, the working process of the first feature extraction unit is as follows: Within the first feature extraction unit, the input of the first feature extraction unit first passes through a noise-adaptive random convolution block, and the output of the noise-adaptive random convolution block is used as the input of the global attention mechanism layer; The output of the global attention mechanism layer is used as the input of the first LN layer, and the output of the first LN layer is residually connected to the output of the noise-adaptive random convolution block to obtain a residual connection result a; The residual connection result a is used as the input of the feed-forward layer, and the output of the feed-forward layer is used as the input of the second LN layer. The output of the second LN layer is residually connected to a to obtain a residual connection result b; The residual connection result b is used as the output of the first feature extraction unit.
[0013] Further, the noise-adaptive random convolution block includes a first convolutional layer to an Nth convolutional layer; the working process of the noise-adaptive random convolution block is as follows: Step 1: A random number between 0 and 1 is generated for each convolutional layer, and the random number corresponding to the th convolutional layer is denoted as ; The feature matrix input to the noise-adaptive random convolution block is denoted as ; Step 2: Initialization , initialize the convolutional layer count ; Step 3: Compare the random number corresponding to the th convolutional layer with the threshold ; If is greater than , then execute Step 4; If is less than or equal to , then set , and then execute Step 5; Step 4: The th convolutional layer uses a convolutional kernel of size to perform convolutional processing on : where represents the convolution operation, represents the feature matrix output by the th convolutional layer; Then execute Step 5; Step 5: Determine whether is equal to N; If is equal to N, then continue to execute Step 6; If is not equal to N, then set , and return to execute Step 3; Step 6: Process respectively to obtain the processing result : where BN represents the batch normalization layer and GELU represents the activation function layer; Concatenate in the channel dimension, and use the concatenated result as the feature matrix output by the noise adaptive random convolution block.
[0014] Furthermore, the working process of the global attention mechanism layer is as follows: Denote the feature matrix input to the global attention mechanism layer as , and calculate the mean and variance of all elements in the feature matrix : where represents the th element in the feature matrix The value of the element in the row and column represents the number of rows of the feature matrix and the number of columns of the feature matrix; The energy value of the element is: where is the smoothing term; Convert the energy value to the attention weight , and apply the attention weight to to obtain : where represents the Hadamard product, and represents the element in the th row and th column of the enhanced feature matrix output by the global attention mechanism layer .
[0015] Furthermore, the working process of the output module is as follows: The input of the output module is the output of the feature extraction module. Inside the output module, the input feature matrix first passes through the average pooling layer, and then the output of the average pooling layer is used as the input of the linear layer, and the output of the linear layer is used as the output of the output module, that is, the fault diagnosis result is obtained.
[0016] Furthermore, the loss function used during the training of the fault diagnosis model is the cross-entropy loss function.
[0017] The beneficial effects of the present invention are: (1) The noise-adaptive random convolution block designed in the present invention guides the model to form an inductive bias for global information through random perturbation of the local features of the vibration signal, significantly reducing the sensitivity to noise while retaining the key fault features, and at the same time simplifying the model complexity.
[0018] (2) The global attention mechanism designed in the present invention can effectively enhance the model's ability to focus on key fault features and has good noise identification ability.
[0019] (3) The fault diagnosis model of the present invention has good non-linear dynamic characteristic modeling ability, and can still achieve high-precision identification of steam turbine main shaft bearing faults in a high-noise environment. At the same time, the model has few parameters and high calculation efficiency, and has good practicability and promotion value.
[0020] The method of the present invention can effectively extract the key fault feature information hidden in a complex noise environment, and take into account the lightweight of the model structure and the diagnostic accuracy. It is applicable to deployment and application in actual industrial scenarios and can meet the intelligent fault identification requirements in actual industrial scenarios. Description of the Drawings
[0021] Figure 1 is a schematic structural diagram of the fault diagnosis model of the present invention; Figure 2 is a flowchart of the operation of the noise adaptive random convolution layer; Figure 3 is a flowchart of the operation of the global attention mechanism layer. Detailed Embodiments
[0022] Detailed Embodiment 1: Combined with Figure 1 This embodiment is described. A fault diagnosis method for a steam turbine main shaft bearing according to this embodiment specifically includes the following steps: Step 1: Collect the original vibration signal of the steam turbine main shaft bearing during operation, and then preprocess the collected original vibration signal to obtain the preprocessed vibration signal of the main shaft bearing; And divide the preprocessed vibration signal of the main shaft bearing into vibration signal segments of equal length according to the set time window length; The vibration signal of the main shaft bearing includes the values at a continuous plurality of sampling points. For example, when the length of the preprocessed vibration signal of the main shaft bearing is 1000 and the set time window length is 100, the processed vibration signal of the main shaft bearing can be divided into 10 vibration signal segments, and the length of each vibration signal segment is 100; Step 2: Obtain a vibration signal data matrix based on all the vibration signal segments and construct a fault diagnosis model; The fault diagnosis model includes an input module, a feature extraction module, and an output module; Among them, the input module sequentially includes an average pooling layer, a convolutional layer (i.e., a conventional two-dimensional convolutional layer), a BN (batch normalization) layer, and a GELU activation function layer; The feature extraction module includes three feature extraction units, and each feature extraction unit includes a noise adaptive random convolution (NARC) block, a global attention mechanism (LGA) layer, a first LN (Layer Normalization) layer, a feedforward layer, and a second LN layer; The output module includes an average pooling layer and a linear layer, and outputs whether there is a fault and the specific fault type through the linear layer; And the constructed fault diagnosis model is trained based on the currently publicly available bearing dataset. After the fault diagnosis model is trained, step three is then executed; Step three: Use the vibration signal data matrix as the input of the fault diagnosis model, and output the fault diagnosis result of the steam turbine main shaft bearing through the fault diagnosis model.
[0023] The present invention proposes a fault diagnosis framework combining noise adaptive convolution and lightweight global attention mechanism, which can achieve intelligent diagnosis with high robustness, high accuracy and high efficiency under complex working conditions, and provide reliable technical support for the health management of key components such as steam turbine main shaft bearings.
[0024] The method of the present invention has the following advantages: (1) Strong anti-noise performance: By introducing a noise adaptive mechanism, the interference of noise to weak features is effectively suppressed; (2) Superior global modeling ability: Using a lightweight global attention mechanism, the global perception ability of key features is improved; (3) Lightweight and efficient structure: The overall model has low computational resource requirements and is suitable for deployment on edge computing platforms; (4) High diagnostic accuracy: Accurate fault identification can be achieved under different working conditions, with strong robustness and adaptability.
[0025] Specific implementation method two: The difference between this implementation method and the first specific implementation method is that an acceleration sensor is used to collect the original vibration signal of the steam turbine main shaft bearing during operation.
[0026] Other steps and parameters are the same as those in the first specific implementation method.
[0027] The acceleration sensor collects acceleration signals in three axial directions, and then synthesizes the acceleration signals in the three axial directions into a combined acceleration signal according to the Pythagorean theorem. The combined acceleration signal is used to represent the total intensity of the acceleration received by the object, and then the combined acceleration signal is preprocessed.
[0028] Specific implementation method three: The difference between this implementation method and the first or second specific implementation method is that the preprocessing of the collected original vibration signal is specifically as follows: The collected original vibration signal is sequentially subjected to detrending and normalization processing.
[0029] Other steps and parameters are the same as those in the first or second specific implementation method.
[0030] Specific implementation method four: The difference between this implementation method and one of the first to third specific implementation methods is that the working process in the input module is as follows: Use the vibration signal data matrix as the input of the average pooling layer, and then use the output of the average pooling layer as the input of the convolutional layer; Take the output of the convolutional layer as the input of the BN layer, and then take the output of the BN layer as the input of the GELU activation function layer; Take the output of the GELU activation function layer as the output of the input module.
[0031] Other steps and parameters are the same as those in any one of the first to third specific embodiments.
[0032] The present invention uses an input module to perform feature preprocessing on the input vibration signal data matrix. Among them, the average pooling layer is used to perform a mean operation on the local time series window, and the convolutional layer is used to perform time-domain dimensionality reduction processing on the vibration signal. The input module can achieve a compressed representation of time series features while retaining key vibration features.
[0033] The definition of the GELU activation function is: Among them, represents the cumulative distribution function of the standard normal distribution.
[0034] Specific embodiment five: The difference between this embodiment and any one of the first to fourth specific embodiments is that the working process of the feature extraction module is as follows: Take the output of the input module as the input of the first feature extraction unit, take the output of the first feature extraction unit as the input of the second feature extraction unit, take the output of the second feature extraction unit as the input of the third feature extraction unit, and take the output of the third feature extraction unit as the output of the feature extraction module.
[0035] Other steps and parameters are the same as those in any one of the first to fourth specific embodiments.
[0036] In the present invention, the working processes of the first feature extraction unit, the second feature extraction unit, and the third feature extraction unit are the same.
[0037] Specific embodiment six: The difference between this embodiment and any one of the first to fifth specific embodiments is that the working process of the first feature extraction unit is as follows: Inside the first feature extraction unit, the input of the first feature extraction unit first passes through a noise-adaptive random convolution block, and then the output of the noise-adaptive random convolution block is used as the input of the global attention mechanism layer; Take the output of the global attention mechanism layer as the input of the first LN layer, and then perform a residual connection between the output of the first LN layer and the output of the noise-adaptive random convolution block to obtain a residual connection result a; Take the residual connection result a as the input of the feed-forward layer, then take the output of the feed-forward layer as the input of the second LN layer, and perform a residual connection between the output of the second LN layer and a to obtain a residual connection result b; Use the residual connection result b as the output of the first feature extraction unit.
[0038] The other steps and parameters are the same as those in any one of the specific embodiments 1 to 5.
[0039] The present invention optimizes the feature distribution through layer normalization and residual connection. Finally, a feedforward layer is used for non-linear mapping to further output features.
[0040] Specific embodiment 7: Combine Figure 2 To illustrate this embodiment. The difference between this embodiment and any one of the specific embodiments 1 to 6 is that the noise adaptive random convolution block includes a first convolutional layer to an Nth convolutional layer; the working process of the noise adaptive random convolution block is as follows: Step 1: Generate a random number between 0 and 1 for each convolutional layer, and denote the random number corresponding to the th convolutional layer as ; Denote the feature matrix input to the noise adaptive random convolution block as ; Step 2: Initialize , and initialize the convolutional layer count ; Step 3: Compare the random number corresponding to the th convolutional layer with the threshold ; The threshold can be set according to the actual situation and is a number between 0 and 1; If is greater than , then execute Step 4; If is less than or equal to , then let , that is, the output of the th convolutional layer does not need to be processed by the th convolutional layer, and then execute Step 5; Step 4: The th convolutional layer uses a convolutional kernel of size to perform convolutional processing on : Among them, represents the convolutional operation, represents the feature matrix output by the th convolutional layer; , that is is randomly selected between 0 and , is the kernel size. The selection of the convolution kernel sizes of N convolutional layers has a complementary influence and can be different; Then execute step 5; Step 5, judge whether it is equal to N; If is equal to N, then continue to execute step 6; If is not equal to N, then let , and return to execute step 3; Step 6, process respectively to obtain the processing result : Among them, BN represents the batch normalization layer, and GELU represents the activation function layer; For perform splicing in the channel dimension, and use the splicing result as the feature matrix output by the noise adaptive random convolution block.
[0041] Other steps and parameters are the same as those in any one of the specific embodiments one to six.
[0042] The NaRC block designed by the present invention can perturb only local features on the premise of keeping the global dimension of the signal unchanged. Since noise usually affects the local details of the signal, and the NaRC block enhances the model's adaptability to local changes through random convolution operations, avoiding the model's over-reliance on noise-related pseudo-features, while retaining the fault impact segment and periodic features. At the same time, it enhances the model's inductive bias for global features, enabling the model to still capture key fault patterns in a noisy environment, thereby improving the model's ability to identify key fault information in a noisy background and enhancing the robustness to noise interference. In addition, the convolution operation of the NaRC block does not contain a bias term, and the introduced number of parameters is extremely small, greatly reducing the structural complexity of the model.
[0043] Specific embodiment eight: Combine Figure 3 to illustrate this embodiment. The difference between this embodiment and any one of the specific embodiments one to seven is that the working process of the global attention mechanism layer is as follows: Denote the feature matrix input to the global attention mechanism layer as , and calculate the mean and variance of all elements in the feature matrix : Among them, represents the value of the element in the row and column of the feature matrix , Denote the feature matrix The number of rows (corresponding to the number of channels of the input features), Denote the feature matrix The number of columns (corresponding to the length of the input features); The energy value of the element is: For: Among them, is the smoothing term, and the lower the energy value indicates that The higher the importance; Convert the energy value to the attention weight , Denote the activation function, and apply the attention weight to , to obtain : Among them, Denote the Hadamard product, Denote the enhanced feature matrix output by the global attention mechanism layer The element in the row and column of the
[0044] Other steps and parameters are the same as those in any one of the first to seventh specific embodiments.
[0045] In the present invention, the eigenvalue of each input feature is regarded as an energy unit. Since noise is manifested as random discrete values and its energy value is manifested as an abnormal feature, by minimizing the energy difference of the same type of features, important features and noise can be distinguished. The global attention mechanism is used to extract global features and identify noise interference, enhancing the model's global attention ability to key fault features and effectively improving the diagnostic performance of the method of the present invention. Moreover, the global attention mechanism layer can reduce the computational complexity by adopting a sparse attention matrix and a dimensionality reduction strategy, and realize the modeling of the spatial feature dependence relationship.
[0046] Specific embodiment nine: The difference between this embodiment and any one of the first to eighth specific embodiments is that the working process of the output module is as follows: The input of the output module is the output of the feature extraction module. In the output module, the input feature matrix first passes through the average pooling layer, and then the output of the average pooling layer is used as the input of the linear layer, and the output of the linear layer is used as the output of the output module, that is, the fault diagnosis result is obtained.
[0047] Other steps and parameters are the same as those in any one of the first to eighth specific embodiments.
[0048] The output module of the present invention can achieve an end-to-end mapping from time-domain feature abstraction to fault diagnosis determination.
[0049] Embodiment 10: Different from any one of Embodiments 1 to 9, the loss function used during the training of the fault diagnosis model is the cross-entropy loss function.
[0050] Other steps and parameters are the same as any one of Embodiments 1 to 9.
[0051] The fault diagnosis model of the present invention is supervised and trained based on the vibration signals of the steam turbine main shaft bearings under various known fault states, and regularization means such as batch normalization and Dropout are combined to improve the generalization ability and convergence stability of the model. Additionally, a small number of vibration signal samples with noise can be introduced for adversarial training according to the industrial scenario to improve the anti-noise performance.
[0052] The above examples of the present invention are only for explaining in detail the calculation model and calculation process of the present invention, rather than limiting the embodiments of the present invention. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is impossible to list all the embodiments here. Any obvious changes or variations derived from the technical solutions of the present invention still fall within the protection scope of the present invention.
Claims
1. A fault diagnosis method for a steam turbine main shaft bearing, characterized in that, The method specifically includes the following steps: Step 1: Collect the original vibration signals of the steam turbine main shaft bearing during operation, and then preprocess the collected original vibration signals to obtain the vibration signals of the main shaft bearing after preprocessing; And divide the vibration signals of the main shaft bearing after preprocessing into vibration signal segments of equal length according to the set time window length; Step 2: Obtain a vibration signal data matrix based on all vibration signal segments, and construct a fault diagnosis model; The fault diagnosis model includes an input module, a feature extraction module, and an output module; Among them, the input module sequentially includes an average pooling layer, a convolutional layer, a BN layer, and a GELU activation function layer; The feature extraction module includes three feature extraction units, and each feature extraction unit includes a noise-adaptive random convolution block, a global attention mechanism layer, a first LN layer, a feed-forward layer, and a second LN layer; The output module includes an average pooling layer and a linear layer; Step 3: Use the vibration signal data matrix as the input of the fault diagnosis model, and output the fault diagnosis result of the steam turbine main shaft bearing through the fault diagnosis model.
2. The fault diagnosis method of a steam turbine main shaft bearing according to claim 1, characterized in that The acceleration sensor is used to collect the original vibration signals of the steam turbine main shaft bearing during operation.
3. The fault diagnosis method of a steam turbine main shaft bearing according to claim 2, characterized in that, The preprocessing of the collected original vibration signals is specifically as follows: Perform detrending and normalization processing on the collected original vibration signals in sequence.
4. The fault diagnosis method of a steam turbine main shaft bearing according to claim 3, characterized in that, The working process of the input module is as follows: Use the vibration signal data matrix as the input of the average pooling layer, and then use the output of the average pooling layer as the input of the convolutional layer; Use the output of the convolutional layer as the input of the BN layer, and then use the output of the BN layer as the input of the GELU activation function layer; Use the output of the GELU activation function layer as the output of the input module.
5. The fault diagnosis method of a steam turbine main shaft bearing according to claim 4, characterized in that, The working process of the feature extraction module is as follows: Use the output of the input module as the input of the first feature extraction unit, use the output of the first feature extraction unit as the input of the second feature extraction unit, use the output of the second feature extraction unit as the input of the third feature extraction unit, and use the output of the third feature extraction unit as the output of the feature extraction module.
6. The fault diagnosis method for a steam turbine main shaft bearing according to claim 5, characterized in that, The working process of the first feature extraction unit is as follows: In the first feature extraction unit, the input of the first feature extraction unit first passes through the noise-adaptive random convolution block, and then uses the output of the noise-adaptive random convolution block as the input of the global attention mechanism layer; Use the output of the global attention mechanism layer as the input of the first LN layer, and then perform a residual connection between the output of the first LN layer and the output of the noise-adaptive random convolution block to obtain a residual connection result a; Use the residual connection result a as the input of the feed-forward layer, then use the output of the feed-forward layer as the input of the second LN layer, and perform a residual connection between the output of the second LN layer and a to obtain a residual connection result b; Use the residual connection result b as the output of the first feature extraction unit.
7. The fault diagnosis method of a steam turbine main shaft bearing according to claim 6, characterized in that, The noise-adaptive random convolution block includes a first convolutional layer to an Nth convolutional layer; the working process of the noise-adaptive random convolution block is as follows: Step 1: Generate a random number between 0 and 1 for each convolutional layer, and denote the random number corresponding to the th convolutional layer as ; Denote the feature matrix input to the noise adaptive random convolution block as ; Step 2, Initialization , Initialize the convolutional layer count ; Step 3, compare the random number corresponding to the th convolutional layer with the threshold in terms of magnitude; If is greater than , then perform step 4; If less than or equal to , then let , and then execute step 5; Step 4, the th convolutional layer uses a convolutional kernel of size to perform convolutional processing on : Among them, represents a convolution operation, represents the feature matrix output by the -th convolutional layer; Then execute step 5; Step 5, determine whether is equal to N; If is equal to N, then proceed to step 6; If is not equal to N, then let , and return to execute step 3; Step 6. Process respectively to obtain a processing result : Among them, BN represents the batch normalization layer, and GELU represents the activation function layer; Pair Concatenate on the channel dimension, and use the concatenation result as the feature matrix output by the noise adaptive random convolution block.
8. A fault diagnosis method for a steam turbine main shaft bearing according to claim 7, characterized in that The working process of the global attention mechanism layer is as follows: Denote the feature matrix input to the global attention mechanism layer as , and calculate the mean and variance of all elements in the feature matrix : Among them, represents the value of the element in the th row and th column of the feature matrix ; represents the number of rows of the feature matrix ; Element energy value is as follows: Among them, is the smoothing term; Convert the energy value into attention weights and apply the attention weights to , Obtain : Among them, represents the Hadamard product, represents the enhanced feature matrix output by the global attention mechanism layer in the row and column elements.
9. A fault diagnosis method for a steam turbine main shaft bearing according to claim 8, characterized in that, The working process of the output module is as follows: The input of the output module is the output of the feature extraction module. Within the output module, the input feature matrix first passes through an average pooling layer, and then the output of the average pooling layer is used as the input of the linear layer. The output of the linear layer is used as the output of the output module, that is, the fault diagnosis result is obtained.
10. A fault diagnosis method for a steam turbine main shaft bearing according to claim 9, characterized in that, The loss function used during the training of the fault diagnosis model is the cross-entropy loss function.
Citation Information
Patent Citations
Face recognition method and system based on random discard of convolutional data and storage medium
CN109934132A
Image noise estimation method based on deep convolutional neural network
CN110852966A
Small sample rolling bearing fault diagnosis method based on convolutional transformer generative adversarial network
CN115859142A
Deep neural network bearing fault diagnosis method containing attention mechanism
CN115905806A
Attention mechanism-based rolling bearing domain adaptive fault diagnosis method
CN116106012A
Cited By
Turbine set fault diagnosis method based on associated information fusion
CN121412798A
Gearbox composite fault diagnosis method and system
CN121682642A
A method and system for diagnosing complex faults in gearboxes
CN121682642B