Mechanical bearing fault diagnosis method and system based on deformable attention mechanism
By converting vibration signals into two-dimensional representations and combining them with the RFB module and the Swin-Deformable Transformer, the problem of high-precision identification of bearing faults under complex working conditions is solved, achieving an accuracy of 98.3% and an F1 score of 98.1%, reducing computational complexity, and making it suitable for industrial equipment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG JIANZHU UNIV
- Filing Date
- 2026-02-04
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies struggle to achieve high-precision identification of bearing faults under complex operating conditions, especially in the early stages of weak faults and in environments with strong noise. Traditional methods are unable to capture key discriminative information, and high-precision models have high computational complexity, making them difficult to deploy in real time on industrial embedded devices with limited computing resources.
A mechanical bearing fault diagnosis method based on deformable attention mechanism is adopted to transform vibration signals into two-dimensional representations with spatial structural information. The fault texture perception is enhanced by visual coding technology. Combined with hierarchical feature learning strategy, local details and global correlation information are collaboratively mined. A feature extraction module and a multi-scale deep feature mining module are designed. The RFB module and Swin-Deformable Transformer are used for feature extraction and classification.
It achieves high-precision identification of bearing faults under complex working conditions, with an accuracy rate of 98.3% and an F1 score of 98.1%. It reduces computational complexity, is suitable for industrial equipment with limited computing resources, and improves the automation level of fault diagnosis.
Smart Images

Figure CN121962767A_ABST
Abstract
Description
A Method and System for Fault Diagnosis of Mechanical Bearings Based on Deformable Attention Mechanism Technical Field
[0001] This invention relates to the field of bearing fault diagnosis technology, and in particular to a method and system for diagnosing mechanical bearing faults based on deformable attention mechanisms. Background Technology
[0002] Rolling bearings are among the most common and vulnerable critical components in thermal power plants, and their operating condition has a decisive impact on the safety, reliability, and service life of the power generation system. Under complex operating conditions and long-term operation, bearings are susceptible to damage from multiple factors, including load fluctuations, lubrication deterioration, and environmental noise, leading to damage in the inner ring, outer ring, or rolling elements. Since early fault characteristics are often weak and easily masked by noise, accurately sensing and effectively identifying the operating condition of bearings remains a crucial problem that urgently needs to be solved in the field of condition monitoring for thermal power plant equipment.
[0003] Vibration signals, due to their high sensitivity to changes in structural state, are widely used in bearing fault diagnosis research. Traditional methods typically rely on artificially constructed time-domain, frequency-domain, or time-frequency-domain features. For example, Li Shule et al. used Generalized Variational Mode Decomposition (GVMD) and fractional Fourier Transform to extract fault frequencies under varying speeds. However, such methods not only heavily rely on expert prior knowledge but also often struggle to capture crucial discriminative information when facing early, weak faults or extremely noisy environments, resulting in limited recognition accuracy. In recent years, deep learning methods, with their powerful end-to-end feature extraction capabilities, have gradually become the mainstream paradigm for fault diagnosis, significantly improving the automation level of fault diagnosis. Among them, Convolutional Neural Networks (CNNs) excel in local feature extraction, but their fixed receptive field limits effective modeling of cross-scale and long-range features. To overcome this limitation, Guo Junfeng et al. and Sun Junjing et al. respectively constructed multi-scale CNNs by introducing attention mechanisms and dilated convolutions, which expanded the receptive field and improved accuracy to some extent. However, their ability to focus on key features under strong noise interference is still insufficient. To further capture global features, Transformers based on self-attention mechanisms have been introduced into the diagnostic field. The CNN-Transformer fusion model proposed by Yang Xu et al., and the tree-inspired hierarchical decision network proposed by Jiao Weidong et al., have achieved refined diagnosis of fault location and severity through complex network architectures, with significantly better accuracy than traditional models. However, although Transformers based on self-attention mechanisms have advantages in global dependency modeling, this performance improvement often comes at the cost of huge computational overhead and parameter scale. Existing high-precision models are generally structurally complex and difficult to deploy in real time on industrial embedded devices with limited computing resources. How to reduce model complexity while ensuring high-precision diagnosis has become an urgent problem to be solved. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a mechanical bearing fault diagnosis method and system based on a deformable attention mechanism. Unlike traditional one-dimensional signal analysis, this invention first utilizes visual coding technology to transform vibration signals into two-dimensional representations with spatial structural information, enhancing the perception of subtle fault textures. Subsequently, a hierarchical feature learning strategy is designed to collaboratively mine local details and global correlation information. This method, through structured signal representation and a hierarchical feature learning strategy, collaboratively mines key discriminative information in bearing fault signals, providing an effective approach for high-precision identification of bearing faults under complex operating conditions.
[0005] On the one hand, a mechanical bearing fault diagnosis method based on deformable attention mechanism is provided, including: acquiring the vibration signal to be diagnosed of a thermal power plant bearing, performing preprocessing operations on the vibration signal to be diagnosed to obtain the image to be diagnosed; inputting the image to be diagnosed into a trained bearing fault diagnosis model, and outputting the bearing fault diagnosis result; wherein, the trained bearing fault diagnosis model includes a feature extraction module, a multi-scale deep feature mining module, and a fault classification module connected in sequence; the image to be diagnosed is input into the feature extraction module, which extracts features at different scales and splices and fuses the features at different scales to obtain enhanced features; the multi-scale deep feature mining module processes the enhanced features using three cascaded S-DT modules, each S-DT module using a shifted window self-attention mechanism to enhance cross-window information interaction; finally, the fault classification module classifies the extracted features to obtain the bearing fault diagnosis result.
[0006] On the other hand, a mechanical bearing fault diagnosis system based on a deformable attention mechanism is provided, including: an acquisition module configured to acquire the vibration signal to be diagnosed of a thermal power plant bearing, perform preprocessing operations on the vibration signal to be diagnosed, and obtain an image to be diagnosed; and a diagnosis module configured to input the image to be diagnosed into a trained bearing fault diagnosis model and output the bearing fault diagnosis result. The trained bearing fault diagnosis model includes a feature extraction module, a multi-scale deep feature mining module, and a fault classification module connected in sequence. The image to be diagnosed is input into the feature extraction module, which extracts features at different scales and splices and fuses these features to obtain enhanced features. The multi-scale deep feature mining module processes the enhanced features using three cascaded S-DT modules. Each S-DT module uses a shifted window self-attention mechanism to enhance cross-window information interaction. Finally, the fault classification module classifies the extracted features to obtain the bearing fault diagnosis result.
[0007] In another aspect, an electronic device is also provided, comprising: a memory for non-transitory storage of computer-readable instructions; and a processor for executing the computer-readable instructions, wherein the computer-readable instructions, when executed by the processor, perform the method described in the first aspect.
[0008] In another aspect, a storage medium is also provided for non-transitory storage of computer-readable instructions, wherein when the non-transitory computer-readable instructions are executed by a computer, the method described in the first aspect is performed.
[0009] In another aspect, a computer system is also provided, including a computer program that, when run on one or more processors, is used to implement the method described in the first aspect above.
[0010] The above technical solution has the following advantages or beneficial effects: Unlike traditional one-dimensional signal analysis, this invention first utilizes visual coding technology to transform vibration signals into two-dimensional representations with spatial structural information, thereby enhancing the perception of subtle fault textures. Subsequently, feature enhancement and extraction methods are designed for feature extraction and representation. Finally, a multi-level feature learning strategy is designed to collaboratively mine local details and global correlation information. This method, through enhanced structured signal representation and a hierarchical feature learning strategy, collaboratively mines key discriminative information in bearing fault signals, providing an effective approach for high-precision identification of bearing faults under complex working conditions. Attached Figure Description
[0011] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0012] Figure 1 is a schematic diagram of the bearing fault diagnosis system in Example 1; Figure 2 is a schematic diagram of the RFB module architecture in Example 1; Figure 3 is a schematic diagram of the multi-scale deep feature mining module architecture in Example 1; Figure 4 is a schematic diagram of different network model architectures in Example 1; Figure 5 is a training accuracy curve of different network models in Example 1 on the CWRU dataset; Figures 6(a)-6(f) are confusion matrices of different network models in Example 1 on the CWRU dataset; Figures 7(a)-7(f) are t-SNE visualization comparison diagrams of different network models in Example 1 on the CWRU dataset. Detailed Implementation
[0013] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0014] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0015] As a core component of key equipment in thermal power plants, the operating status of rolling bearings directly determines the safety and reliability of the equipment. Under complex working conditions, the vibration signal fault characteristics exhibit nonlinear and multi-scale characteristics. Traditional methods have problems such as insufficient multi-scale feature modeling, difficulty in fine-grained fault identification, and weak faults being easily masked by noise.
[0016] To address this, this study proposes a bearing fault diagnosis method, RFB-S-DT, which combines a Receptive Field Block (RFB) module with a deformable attention mechanism. First, the Gramian Angular Summation Field (GASF) is used to convert one-dimensional vibration signals into two-dimensional image representations, enhancing the time-series correlation and visualizing fault features, highlighting subtle fault information. Then, an RFB module is introduced into the shallow layers of the network, using a multi-branch convolutional structure to achieve adaptive enhancement of the multi-scale receptive field, accurately capturing fault features of different damage levels and locations. Finally, a deep feature mining module is constructed using a Swin-Deformable Transformer, leveraging deformable attention and a hierarchical windowing strategy to improve the global modeling capability of key fault regions while reducing computational complexity.
[0017] Experimental results based on the CWRU bearing dataset show that the proposed method achieves an accuracy of 98.3% and an F1 score of 98.1% across 10 fault diagnosis tasks, significantly outperforming mainstream models such as AlexNet, ViT, and Swin-T. With a reasonable complexity of 26.9M parameters and 4.65G FLOPs, it significantly improves diagnostic performance and robustness, providing an effective solution for fine-grained fault diagnosis of bearings under complex operating conditions and contributing to condition monitoring and intelligent operation and maintenance of mechanical equipment.
[0018] This embodiment provides a mechanical bearing fault diagnosis method based on a deformable attention mechanism. The mechanical bearing fault diagnosis method based on a deformable attention mechanism includes: S101: acquiring the vibration signal to be diagnosed from a bearing in a thermal power plant, performing preprocessing operations on the vibration signal to obtain a image to be diagnosed; S102: inputting the image to be diagnosed into a trained bearing fault diagnosis model, and outputting the bearing fault diagnosis result; wherein, the trained bearing fault diagnosis model includes a feature extraction module, a multi-scale deep feature mining module, and a fault classification module connected in sequence; the image to be diagnosed is input into the feature extraction module, which extracts features at different scales and splices and fuses the features at different scales to obtain enhanced features; the multi-scale deep feature mining module processes the enhanced features using three cascaded S-DT modules, each S-DT module employing a shifted window self-attention mechanism to enhance cross-window information interaction; finally, the fault classification module classifies the extracted features to obtain the bearing fault diagnosis result.
[0019] The fault classification module includes: a downsampling layer, a global average pooling layer, and a fully connected layer connected in sequence.
[0020] Further, in step S101: the vibration signal to be diagnosed from the bearing of the thermal power plant equipment is acquired by using an accelerometer installed on the housing near the outer ring of the bearing. Thermal power plant equipment includes, for example, steam turbines, various water pumps, fans, and generators.
[0021] Further, in step S101: preprocessing the vibration signal to be diagnosed to obtain the image to be diagnosed, the preprocessing operation includes: first, slicing the original one-dimensional vibration signal using a sliding window; second, compressing each slice sample using a segmented aggregation approximation algorithm to obtain time series data; and then, converting the time series data into two-dimensional image data using the Gram angle and field algorithm GASF.
[0022] For example, in the data preprocessing stage, the original one-dimensional vibration signal is first sliced using a sliding window technique, with the window length set to 1024 sampling points. To reduce the data dimensionality and retain key time-series features, the Piecewise Aggregate Approximation (PAA) algorithm is automatically used to compress each sample to 224 dimensions. Subsequently, Gramian Angular Summation Field (GASF) is used to convert the time-series data into two-dimensional image data to meet the input requirements of the neural network. A strict data partitioning strategy is employed, with samples of all categories divided into training, validation, and test sets in a 6:2:2 ratio to ensure the objectivity of the model evaluation. The specific fault types, corresponding labels, and sample distribution of the CWRU bearing dataset are shown in Table 1.
[0023] Table 1. Detailed composition and experimental environment of the CWRU dataset.
[0024] The Gram Angle and Field Algorithm (GASF) can encode one-dimensional time series data into two-dimensional visual images, thereby making certain features of the data more intuitive.
[0025] The specific implementation steps of the Gram angle and field algorithm are as follows: First, obtain... The original fault signal of each sample. Let the signal of each sample in the time series be represented as... ,in For signal acquisition time, This represents the signal value at the corresponding time. The data is then normalized and scaled to the interval [0, 1], denoted as... .
[0026]
[0027] Next, the one-dimensional signal sequence in the Cartesian coordinate system will be transformed to the polar coordinate system using the following formula:
[0028] in Indicates angle, Represents the radius. From formula (2), it can be seen that the inverse cosine function at angles... The wavelength decreases monotonically within the range [0, π]. As the wavelength increases, each... It corresponds to only one value in the polar coordinate system and forms curvature between different angles on the polar coordinate circle.
[0029] Based on the above analysis, the GASF matrix is defined as follows:
[0030] in, As unit vectors, the size of the visually encoded image is determined by the dimension of the original data. .
[0031] Furthermore, the performance metrics for evaluating the trained bearing fault diagnosis model during the training process include accuracy and F1 score.
[0032] To more effectively extract discriminative features from GASF images converted from fault signals, this invention introduces a Receptive Field Block (RBB) module to extract and characterize features.
[0033] Further, the feature extraction module includes: a first branch, a second branch, and a third branch in parallel; the first branch includes: a first convolutional layer and a second convolutional layer connected in sequence; the second branch includes: a third convolutional layer and a fourth convolutional layer connected in sequence; the third branch includes: a fifth convolutional layer, a sixth convolutional layer, and a seventh convolutional layer connected in sequence; the input ends of the first convolutional layer, the second convolutional layer, and the third convolutional layer are all input to two-dimensional image data; the output ends of the second convolutional layer, the fourth convolutional layer, and the seventh convolutional layer are all connected to the input end of the stitching unit; the output end of the stitching unit is connected to the input end of the eighth convolutional layer, and the output end of the eighth convolutional layer is connected to the first input end of the adder; the second input end of the adder inputs the original two-dimensional image data; the output end of the adder outputs the enhanced features.
[0034] Furthermore, the feature extraction module is used to concatenate the features extracted from each branch along the channel dimension, and perform cross-channel information fusion through a convolutional layer, ultimately outputting a feature map with multi-scale information aggregation.
[0035] As shown in Figure 2, the feature extraction module consists of multiple parallel processing streams, each responsible for capturing features at different scales. The features extracted from each branch are concatenated along the channel dimension and then processed through a 1... One convolutional layer performs cross-channel information fusion, ultimately outputting a feature map with multi-scale information aggregation. :
[0036] in, Indicates 1 1. The weight matrix of the fused convolution. This represents a convolution operation; the ReLU activation function is used to enhance the model's nonlinear expressive power.
[0037] The beneficial effects of the feature extraction module are mainly reflected in its ability to expand the receptive field and enhance multi-scale feature extraction capabilities in convolutional neural networks (CNNs). Specifically, the RFB module dynamically adjusts the receptive field by introducing convolutional kernels of different sizes, thereby capturing more diverse and richer feature representations.
[0038] To further improve the accuracy of fault type identification, the effective features extracted by the RFB module are used as input, and the multi-scale deep feature mining module Swin-Deformable Transformer (S-DT) is introduced to carry out multi-scale deep feature mining. The S-DT module is built on the hierarchical architecture of Swin Transformer.
[0039] To alleviate the quadratic computational complexity faced by traditional Vision Transformer (ViT) in global self-attention. The problem is that the multi-scale deep feature mining module first divides the input features into a set of non-overlapping local windows and performs self-attention calculation within each window.
[0040] Compared to global attention, this window-based approach reduces computational cost to (in Indicates the number of windows, (This indicates the number of tokens within the window), while still effectively modeling local dependency structures.
[0041] As shown in Figure 3, the multi-scale deep feature mining module adopts a three-stage hierarchical design. As the stages progress, the spatial resolution of the feature map gradually decreases, while the channel dimension gradually increases, realizing the transition from low-level texture features to high-level semantic features.
[0042] Furthermore, the multi-scale deep feature mining module includes: a first downsampling layer, a first S_DT module, a second downsampling layer, a second S_DT module, a third downsampling layer, and a third S_DT module connected in sequence; the internal structures of the first S_DT module, the second S_DT module, and the third S_DT module are identical; the first S_DT module includes: a first-layer normalization unit, a shift window self-attention mechanism unit, a second-layer normalization unit, a third-layer normalization unit, a feedforward fully connected network layer, and the output of the first S_DT module; the output of the first-layer normalization unit is connected to the output of the first S_DT module, the output of the shift window self-attention mechanism unit is connected to the input of the second-layer normalization unit; the output of the second-layer normalization unit is connected to the input of the third-layer normalization unit, the input of the feedforward fully connected network is also connected to the output of the third-layer normalization unit, and the output of the feedforward fully connected network is connected to the output of the first S_DT module.
[0043] Furthermore, the multi-scale deep feature mining module is used to divide the input features into a set of non-overlapping local windows and perform self-attention calculation within each window.
[0044] The first S_DT module is used to extract local dependencies through a window self-attention mechanism and optimize it through a normalization layer and a feedforward fully connected network. This reduces the computational cost while enabling the model to capture semantic dependencies over longer distances and further extract fine-grained features.
[0045] The second S_DT module is used to implement extended cross-window information interaction, further enhance the expression of global features, and optimize the interactivity of features in the spatial dimension. It further learns complex structural patterns and mid-level semantic relationships between different regions.
[0046] The third S_DT module is used to integrate and aggregate high-level features. It achieves cross-level feature aggregation through a global average pooling layer and finally outputs high-level semantic information, providing effective support for fault classification.
[0047] The beneficial effects of the multi-scale deep feature mining module are: by introducing a multi-stage feature extraction and attention mechanism, it gradually realizes the transition from low-level local features to high-level semantic features, enables fine-grained feature extraction at different scales, enhances the feature expression capability of the model, and thus improves the diagnostic accuracy and robustness of the model under complex fault modes.
[0048] Within each S_DT module, a shifted window self-attention mechanism is used to enhance cross-window information interaction.
[0049] Furthermore, the shift window self-attention mechanism unit includes: unlike traditional self-attention mechanisms, the shift window self-attention mechanism adopts window segmentation and window translation strategies, which reduces computational overhead while achieving effective fusion of local and global information.
[0050] In the shift-window self-attention mechanism, the input feature map is first divided into multiple non-overlapping local windows, and the elements within each window compute self-attention with each other. This localized computation method significantly reduces computational cost while capturing local information within the window. The size of each window is typically a hyperparameter that determines the range of information the model can capture.
[0051] To model the relationships between different windows while maintaining low computational complexity, the shifted window self-attention mechanism achieves this by translating the position of each window. Specifically, during each self-attention calculation, the window is translated by an offset, allowing each window to interact with elements in neighboring windows.
[0052] For a window Its translation form can be expressed as:
[0053] in, This indicates a window that has been panned.
[0054] By using deformable offsets, the model selectively focuses on important spatial regions, which is particularly important for capturing subtle features in the time-frequency plots of fault data.
[0055] Assume the original query, key, and value are respectively , and The offset is Then the attention can be transformed into:
[0056] in The position information is introduced by the offset. It is the dimension of the key vector.
[0057] The deformable attention mechanism proposed in this invention is the shift window self-attention mechanism.
[0058] Furthermore, the training process of the trained bearing fault diagnosis model includes: constructing a dataset and dividing the dataset into a training set, a validation set, and a test set according to a certain ratio; the dataset consists of bearing vibration signals of known bearing fault types; inputting the training set into the bearing fault diagnosis model to train the model; stopping the training when the loss function value of the model no longer decreases, thus obtaining the trained bearing fault diagnosis model.
[0059] To verify the fault diagnosis performance and effectiveness of the proposed method, the Case Western Reserve University (CWRU) bearing dataset, a recognized benchmark dataset, was used for validation. The experimental process flowchart is shown in Figure 1. This dataset introduces single-point faults on the inner ring, outer ring, and rolling elements of the test bearing using electrical discharge machining (EDM). Vibration signals are collected by an accelerometer mounted on the drive end, with a sampling frequency set to 12 kHz. To construct a representative 10-class fault diagnosis task, this experiment selected fault data including normal states and faults of varying severity. Specifically, the fault categories cover three different damage diameters (0.007 inches, 0.014 inches, and 0.021 inches), each containing three different fault locations: the inner ring, outer ring, and rolling elements. Therefore, the experimental dataset contains one normal state and nine fault states.
[0060] In this invention, several representative deep learning models in visual tasks were selected as comparison baselines to ensure that the evaluation results have broad reference value.
[0061] As shown in Figure 4, AlexNet, as a typical representative of early deep convolutional networks, achieved a breakthrough in large-scale image classification performance through a five-layer convolutional structure, laying the foundation for subsequent network designs. The VGG network further validated the role of network depth in enhancing feature representation capabilities through a deeper and more regular convolutional stacking architecture. InceptionV3 introduced an efficient structure through multi-scale convolutional modules, enabling the model to fully capture feature information from different receptive fields. ViT utilizes a pure attention mechanism to model global dependencies between image patches, representing a significant paradigm shift from convolutional architectures to Transformer architectures. Swin-T, within a hierarchical Transformer framework, employs a shifted window strategy, achieving both high efficiency and consideration of local and global feature representation. These models cover the main technical routes in the field of deep vision, providing a systematic and robust benchmark for subsequent method performance evaluation and analysis.
[0062] In deep learning-based classification tasks, choosing appropriate evaluation metrics is crucial for assessing model performance. For evaluating classification accuracy, this study employs the following two metrics:
[0063]
[0064] TP, TN, FP, and FN represent the number of true positives, true negatives, false positives, and false negatives, respectively. Accuracy refers to the proportion of correctly predicted samples out of the total number of samples. The F1 score is the harmonic mean of precision and recall, which comprehensively considers the precision and recall of the classification model. A higher F1 score indicates that the model classifies correctly while having fewer false positives and false negatives. Furthermore, FLOPs are often used to evaluate the time complexity of a model; it refers to the number of floating-point operations required to run the network model once. The number of parameters refers to the number of parameters that need to be learned in the model and is an important indicator of the model's space complexity. The more parameters a model has, the more GPU memory is required to train the model, and the more disk space is required to store the model parameters. The parameter details of the RFB-S-DT neural network model proposed in this invention are shown in Table 2.
[0065] To systematically verify the effectiveness and superiority of the proposed RFB-S-DT network in bearing fault diagnosis, this section compares and analyzes the experimental results of different mainstream network models on the CWRU dataset from multiple perspectives, including overall performance comparison, training convergence characteristics, classification detail analysis, and feature distribution visualization. It also discusses the experimental phenomena in depth in conjunction with the characteristics of the model structure.
[0066] Table 2 RFB-S-DT Model Parameter Configuration
[0067] Table 3 compares the performance of AlexNet, VGG, InceptionV3, ViT, Swin-T, and the proposed RFB-S-DT model on the CWRU dataset's 10-class classification task. The results show that the traditional CNN models AlexNet and VGG achieve accuracies of 92.4% and 93.1%, respectively. While they can extract basic fault features, their performance improvement is limited by their fixed receptive field and single-scale modeling capabilities. InceptionV3 enhances feature representation through multi-scale convolution, but its accuracy is only 90.6%, indicating that relying solely on convolutional scale expansion is insufficient to characterize complex fault patterns. In contrast, the introduction of the Transformer significantly improves model performance, with ViT and Swin-T achieving accuracies of 94.8% and 96.5%, respectively, demonstrating the crucial role of long-range dependency modeling in fault differentiation. The proposed RFB-S-DT achieves the best results in both accuracy and F1 score (98.3% and 98.1%), an improvement of approximately 1.8 percentage points over Swin-T, reflecting the enhanced feature discriminative power and robustness of the RFB module. Meanwhile, the model maintains a reasonable level in terms of FLOPs and parameter count, demonstrating a good balance between performance and complexity.
[0068] Table 3. Performance Comparison of Different Network Models Network Model Accuracy (%) F1 Score (%) FLOPs (G) Number of Parameters (M) Alexnet 92.4 92.2 0.6 8 56.2 VGG 93.1 92.7 7.4 125.6 Inception V3 90.6 90.3 3.1 0 22.4 ViT 94.8 94.4 5.6 30.2 Swin-T 96.5 96.2 4.5 228.1 RFB-S-DT 98.3 98.1 4.6 5 26.9 Table 5 shows the accuracy changes of different models during 200 training rounds. All models converged quickly in the early stages of training, but there were significant differences in convergence speed and stability. Traditional CNN models (AlexNet, VGG) fluctuated significantly in the later stages, resulting in lower final accuracy. The core reason is that the fixed receptive field limits the feature representation ability, and the gradient is prone to local optima during training, making it difficult to continuously optimize to better performance. ViT and Swin-T converged more smoothly, thanks to the global modeling capability of the Transformer architecture. However, ViT's global self-attention calculation led to unstable gradient propagation in the early stages of training, resulting in a slightly slower convergence speed than Swin-T. The proposed RFB-S-DT achieved high accuracy after approximately 50 iterations and maintained a small range of fluctuations in the later stages.
[0069] To analyze the recognition capabilities of different models across various fault categories, Figures 6(a)-6(f) present the confusion matrices of the six networks on the test set. AlexNet, VGG, and InceptionV3 exhibit significant misclassification between different damage levels at the same fault location (e.g., B014 / B021, IR014 / IR021). This is because traditional CNNs rely on local convolution operations, capturing only local features at a single scale, failing to distinguish subtle feature differences arising from fault damage levels, and are susceptible to noise interference leading to feature confusion. In contrast, ViT and Swin-T show significantly enhanced diagonal elements, reducing misclassification and indicating that the self-attention mechanism has advantages in modeling global correlations, although slight confusion still exists in the early damage stage. RFB-S-DT shows the best confusion matrix performance, with recognition accuracy approaching or reaching 0.99 for each category, especially with almost no misclassification between different damage levels in the inner and outer circles. This demonstrates that the combined effect of multi-scale receptive field enhancement and deformable attention effectively improves the model's ability to distinguish fine-grained fault features.
[0070] t-SNE visualization was used to evaluate the effectiveness of different models in feature extraction and classification. Figures 7(a)-7(f) show that the feature distributions of AlexNet and VGG exhibit significant inter-cluster overlap and weak discriminative power. This is because traditional CNN feature extraction relies on local receptive fields, failing to capture the global correlation and multi-scale characteristics of fault signals, resulting in a lack of uniqueness in feature representation. InceptionV3 shows improved inter-cluster spacing, but intra-cluster dispersion still exists, indicating that simply expanding the receptive field through multi-scale convolution without combining global modeling and feature selection mechanisms makes it difficult to form highly aggregated class features. In contrast, ViT and Swin-T form relatively clear class cluster structures, demonstrating the Transformer's advantage in high-level feature abstraction, but there is still room for improvement in the compactness of some classes. RFB-S-DT's feature distribution exhibits the best inter-cluster separation and intra-cluster compactness, with clear class boundaries, validating its effectiveness in multi-scale feature modeling and long-range dependency learning at the feature level. To systematically evaluate the effectiveness of each structural component in the proposed RFB-S-DT network, a series of ablation experiments were conducted on the CWRU bearing fault dataset. All comparative models were trained and tested under the same experimental conditions, including consistent data partitioning, number of training epochs, optimizer parameters, and evaluation metric settings. The relevant quantitative results are summarized in Table 4.
[0071] Table 4. Comparison of Ablation Experiment Results for Different Network Structures and Key Modules | No. | Model Structure | RFB | Multi-scale | RFB | DeformableAttention | Accuracy (%) | F1-Score (%) | FLOPs (G) | Params (M) | A1 | CNN Backbone | ××× | 92.4 | 92.2 | 0.6 | 856.2 | A2 | RFB-CNN | √√ | × | 95.2 | 94.9 | 3.9 | 829.4 | A3 | Swin-T | ××× | 96.5 | 96.2 | 4.5 | 228.1 | A4 | S-DT (No RFB) | ××√ | 97.1 | 96.9 | 4.5 | 827.6 | A5 | RFB-Single+ S-DT | √×√ | 97.6 | 97.2 | 4.6 | 026.5 | A6 | RFB-Dilate (d = 3) + S-DT | √×√ | 97.9 | 97.6 | 4.6 | 226.7 | A7 | RFB-Multi+ S-DT√√√98.398.14.6526.9 The ablation experiments show that each structural module plays a crucial role in improving model performance. The baseline model A1, using a CNN backbone, achieves an accuracy of 92.4% and an F1 score of 92.2%, but exhibits significant limitations in multi-scale feature modeling. A2, after introducing the RFB module, enhances the receptive field across multiple branches through multi-scale convolutions, resulting in a 2.8% performance improvement while significantly reducing the number of parameters, demonstrating high parameter efficiency. A3, using a Swin-T backbone, achieves an accuracy of 96.5% and an F1 score of 96.2%, validating the Transformer's advantages in global dependency modeling. Further introducing a deformable attention mechanism to construct an S-DT structure improves the performance of model A4 to 97.1% and 96.9%, with only a limited increase in computational cost. With the gradual integration of different forms of RFB modules with S-DT, the model performance continues to improve. Among them, the A7 model, which combines multi-scale RFB with deformable attention, achieves the best results under the premise of controllable complexity, with an accuracy of 98.3% and an F1 score of 98.1%, which fully verifies the synergistic enhancement effect between receptive field enhancement and deformable attention mechanism.
[0072] This invention addresses the shortcomings of insufficient multi-scale feature modeling capabilities and difficulties in distinguishing complex fault modes in bearing fault signals. It proposes an RFB-S-DT bearing fault diagnosis network that integrates receptive field enhancement and a deformable attention mechanism. This method converts one-dimensional vibration signals into two-dimensional image representations using GASF (Gas-Assisted Field Synthesis), and combines the RFB module with a hierarchical attention structure of Swin-DeformableTransformer to achieve collaborative modeling of multi-scale local features and global dependencies. Experimental results on 10 types of fault diagnosis based on the CWRU bearing dataset show that the proposed model significantly outperforms various mainstream convolutional networks and Transformer models in terms of classification accuracy and F1 score, achieving an accuracy of 98.3% and an F1 score of 98.1%, while maintaining a good balance between computational complexity and parameter size. Further training process analysis, confusion matrix, and t-SNE feature visualization results demonstrate that the model has faster convergence speed, more stable optimization characteristics, clearer inter-cluster boundaries, and higher intra-cluster compactness, effectively reducing the probability of misclassification between similar fault modes. The ablation experiments further validated the key roles and synergistic enhancement effects of the RFB module and deformable attention mechanism in performance improvement. In summary, the RFB-S-DT model significantly improves bearing fault diagnosis accuracy while ensuring controllable computational efficiency, exhibiting good generalization ability and engineering application potential, and providing an effective technical path for intelligent fault diagnosis of complex rotating machinery.
[0073] Example 2 This example provides a mechanical bearing fault diagnosis system based on deformable attention mechanism. The mechanical bearing fault diagnosis system based on deformable attention mechanism includes: an acquisition module, which is configured to: acquire the vibration signal to be diagnosed of a thermal power plant bearing, perform preprocessing operations on the vibration signal to be diagnosed, and obtain the image to be diagnosed; a diagnosis module, which is configured to: input the image to be diagnosed into a trained bearing fault diagnosis model and output the bearing fault diagnosis result; wherein, the trained bearing fault diagnosis model includes a feature extraction module, a multi-scale deep feature mining module, and a fault classification module connected in sequence; the image to be diagnosed is input into the feature extraction module, which extracts features at different scales and splices and fuses the features at different scales to obtain enhanced features; the multi-scale deep feature mining module processes the enhanced features using three cascaded S-DT modules, each S-DT module using a shifted window self-attention mechanism to enhance cross-window information interaction; finally, the fault classification module classifies the extracted features to obtain the bearing fault diagnosis result.
[0074] It should be noted that the above-mentioned acquisition module and output module correspond to steps S101 to S102 in Embodiment 1. The examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.
[0075] The descriptions of each embodiment in the above embodiments have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0076] The proposed system can be implemented in other ways. For example, the system implementation examples described above are merely illustrative, and the module division described above is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed.
[0077] Example 3 This embodiment also provides an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory to cause the electronic device to perform the method described in Example 1.
[0078] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0079] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.
[0080] In the implementation process, each step of the above method can be completed by the integrated logic circuits in the processor hardware or by software instructions.
[0081] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.
[0082] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in connection with these embodiments can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0083] Example 4 This embodiment also provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, complete the method described in Example 1.
[0084] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A mechanical bearing fault diagnosis method based on deformable attention mechanism, characterized in that, include: The process involves acquiring the vibration signal to be diagnosed from the bearing of a thermal power plant, performing preprocessing operations on the vibration signal to obtain the image to be diagnosed, inputting the image to be diagnosed into the trained bearing fault diagnosis model, and outputting the bearing fault diagnosis result. The trained bearing fault diagnosis model includes a feature extraction module, a multi-scale deep feature mining module, and a fault classification module connected in sequence. The image to be diagnosed is input into the feature extraction module, which extracts features at different scales and splices and fuses these features to obtain enhanced features. The multi-scale deep feature mining module processes the enhanced features using three cascaded S-DT modules. Each S-DT module uses a shift window self-attention mechanism to enhance cross-window information interaction. Finally, the fault classification module classifies the extracted features to obtain the bearing fault diagnosis result.
2. The mechanical bearing fault diagnosis method based on deformable attention mechanism as described in claim 1, characterized in that, The vibration signal to be diagnosed is preprocessed to obtain the image to be diagnosed. The preprocessing operation includes: first, slicing the original one-dimensional vibration signal using a sliding window; second, compressing each slice sample using a segmented aggregation approximation algorithm to obtain time series data; and then, converting the time series data into two-dimensional image data using Gram angle and field algorithms.
3. The mechanical bearing fault diagnosis method based on deformable attention mechanism as described in claim 1, characterized in that, The feature extraction module includes three parallel branches: a first branch, a second branch, and a third branch. The first branch includes a first convolutional layer and a second convolutional layer connected in sequence. The second branch includes a third convolutional layer and a fourth convolutional layer connected in sequence. The third branch includes a fifth convolutional layer, a sixth convolutional layer, and a seventh convolutional layer connected in sequence. The inputs of the first, second, and third convolutional layers are all two-dimensional image data. The outputs of the second, fourth, and seventh convolutional layers are all connected to the input of a stitching unit. The output of the stitching unit is connected to the input of an eighth convolutional layer, and the output of the eighth convolutional layer is connected to the first input of an adder. The second input of the adder receives the original two-dimensional image data. The output of the adder outputs enhanced features. The feature extraction module is used to stitch the features extracted from each branch along the channel dimension and perform cross-channel information fusion through a convolutional layer, ultimately outputting a multi-scale information aggregation feature map.
4. The mechanical bearing fault diagnosis method based on deformable attention mechanism as described in claim 1, characterized in that, The feature extraction module consists of multiple parallel processing streams, each responsible for capturing features at different scales; the features extracted from each branch are concatenated along the channel dimension and then processed through a 1... One convolutional layer performs cross-channel information fusion, ultimately outputting a feature map with multi-scale information aggregation. : in, Indicates 1 1. The weight matrix of the fused convolution. This represents a convolution operation; the ReLU activation function is used to enhance the model's nonlinear expressive power.
5. The mechanical bearing fault diagnosis method based on deformable attention mechanism as described in claim 1, characterized in that, The multi-scale deep feature mining module includes: a first downsampling layer, a first S_DT module, a second downsampling layer, a second S_DT module, a third downsampling layer, and a third S_DT module connected in sequence; the internal structures of the first S_DT module, the second S_DT module, and the third S_DT module are identical; the first S_DT module includes: a first-layer normalization unit, a shift window self-attention mechanism unit, a second-layer normalization unit, a third-layer normalization unit, a feedforward fully connected network, and the output of the first S_DT module; the output of the first-layer normalization unit is connected to the output of the first S_DT module, the output of the shift window self-attention mechanism unit is connected to the input of the second-layer normalization unit; the output of the second-layer normalization unit is connected to the input of the third-layer normalization unit, the input of the feedforward fully connected network is also connected to the output of the third-layer normalization unit, and the output of the feedforward fully connected network is connected to the output of the first S_DT module.
6. The mechanical bearing fault diagnosis method based on deformable attention mechanism as described in claim 1, characterized in that, The multi-scale deep feature mining module is used to divide the input features into a set of non-overlapping local windows and perform self-attention computation within each window. The first S_DT module is used to extract local dependencies through the window self-attention mechanism and optimize it through normalization layers and feedforward fully connected networks. This reduces the computational cost while enabling the model to capture semantic dependencies over longer distances and further extract fine-grained features. The second S_DT module is used to expand the interaction of cross-window information, further enhance the expression of global features, and optimize the interactivity of features in the spatial dimension. It also learns complex structural patterns and mid-level semantic associations between different regions. The third S_DT module is used to integrate and aggregate high-level features. It achieves cross-level feature aggregation through a global average pooling layer and finally outputs high-level semantic information, providing effective support for fault classification.
7. The mechanical bearing fault diagnosis method based on deformable attention mechanism as described in claim 5, characterized in that, The shifted window self-attention mechanism unit includes: the input feature map is first divided into multiple non-overlapping local windows, and the elements within each window calculate self-attention with each other; during each self-attention calculation, the window is shifted according to an offset, so that each window can interact with the elements of neighboring windows; for a window Its translation form can be expressed as: in, This represents the window after translation; assuming the original query, key, and value are respectively... , and The offset is Then the attention can be transformed into: in The position information is introduced by the offset. It is the dimension of the key vector.
8. A mechanical bearing fault diagnosis system based on deformable attention mechanism, characterized in that, include: The acquisition module is configured to: acquire the vibration signal to be diagnosed of the bearing in the thermal power plant, perform preprocessing operations on the vibration signal to be diagnosed, and obtain the image to be diagnosed; the diagnosis module is configured to: input the image to be diagnosed into the trained bearing fault diagnosis model and output the bearing fault diagnosis result; wherein, the trained bearing fault diagnosis model includes a feature extraction module, a multi-scale deep feature mining module and a fault classification module connected in sequence. The image to be diagnosed is input into the feature extraction module, which extracts features at different scales and splices and fuses these features to obtain enhanced features. The multi-scale deep feature mining module processes the enhanced features using three cascaded S-DT modules. Each S-DT module uses a shift window self-attention mechanism to enhance cross-window information interaction. Finally, the fault classification module classifies the extracted features to obtain the bearing fault diagnosis result.
9. An electronic device, characterized in that, include: A memory for non-transitory storage of computer-readable instructions; and a processor for executing the computer-readable instructions, wherein the computer-readable instructions, when executed by the processor, perform the method described in any one of claims 1-7.
10. A storage medium, characterized in that, Non-transitory storage of computer-readable instructions, wherein when the non-transitory computer-readable instructions are executed by a computer, the method of any one of claims 1-7 is performed.