A bearing fault diagnosis method based on sample adaptive fractional order attention transformer

CN122673601APending Publication Date: 2026-09-01NANJING TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610839820.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0006]针对现有技术存在的不足,本发明提出一种基于样本自适应分数阶注意力Transformer的轴承故障诊断方法,旨在解决传统故障诊断方法对复杂振动模式建模不足、普通注意力机制自适应能力弱以及复杂噪声环境下特征提取不稳定的问题,提高轴承故障诊断的准确率、稳定性和抗噪性能

Benefits of technology

1、本发明采用一层卷积Token特征提取模块,能够在保持结构简洁的同时有效提取原始振动信号的局部时序特征,实现由原始序列到Token表示的高效映射;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122673601A_ABST
    Figure CN122673601A_ABST
Patent Text Reader

Abstract

The application discloses a bearing fault diagnosis method based on a sample adaptive fractional order attention Transformer, and comprises the following steps: acquiring a bearing vibration sample; performing local feature extraction and Tokenization representation on the bearing vibration sample to obtain an initial Token feature representation; performing gate weighting on the initial Token feature representation to obtain a gate enhanced feature; introducing position information to obtain an input feature sequence; designing a sample adaptive fractional order parameter according to a statistical quantity feature of the input feature sequence, combining a multi-head attention mechanism and a feedforward residual network to perform feature coding; performing layer normalization and Token attention pooling on the coded sequence feature to extract a global fault representation; and outputting a bearing fault diagnosis result through a feature projection layer and a linear classifier. The application can improve the accuracy, stability and anti-noise performance of bearing fault diagnosis, and is suitable for bearing fault recognition tasks under complex working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of rotating machinery fault diagnosis technology, and in particular to a bearing fault diagnosis method based on a sample adaptive fractional attention Transformer. Background Technology

[0002] Bearings are critical fundamental components in rotating machinery systems, widely used in aero-engines, wind power generation equipment, machine tools, rail transportation, and various industrial transmission systems. During long-term operation, bearings are susceptible to wear, fatigue, impact loads, and complex operating conditions, leading to inner ring failures, outer ring failures, rolling element failures, and combined failures. Since bearing failures directly impact the safety and reliability of the entire mechanical system, researching accurate, stable, and efficient bearing fault diagnosis methods is of great significance.

[0003] Traditional bearing fault diagnosis methods often rely on manually constructed time-domain, frequency-domain, or time-frequency-domain features, such as root mean square, kurtosis, envelope spectrum, and wavelet energy, which are then combined with classifiers such as support vector machines, random forests, and K-nearest neighbors for fault identification. While these methods are effective under simple operating conditions, they suffer from problems such as strong reliance on human experience, limited feature generalization ability, and insufficient robustness in complex noisy environments and variable operating conditions.

[0004] In recent years, deep learning methods have been widely used in bearing fault diagnosis. Convolutional neural networks can extract local features, and Transformer structures can model long-range dependencies. However, existing methods still have the following shortcomings: First, when using the original sequence directly for modeling, the effective expressive power of the input features is limited and they are easily affected by noise. Second, traditional attention mechanisms typically employ fixed forms of similarity metrics, making it difficult to adaptively adjust the attention distribution based on sample characteristics; Third, ordinary position encoding and simple pooling methods are not sufficient to represent critical fault segments; Fourth, when bearing failure samples are complex and the boundaries between categories are blurred, traditional models have limited ability to characterize intra-class compactness and inter-class separability.

[0005] Therefore, a bearing fault diagnosis method that can take into account local feature extraction, global dependency modeling, sample adaptive attention adjustment, and the ability to focus key fault information is needed. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention proposes a bearing fault diagnosis method based on a sample adaptive fractional-order attention Transformer. This method aims to solve the problems of insufficient modeling of complex vibration modes, weak adaptive capability of ordinary attention mechanisms, and unstable feature extraction under complex noise environments in traditional fault diagnosis methods, thereby improving the accuracy, stability, and noise resistance of bearing fault diagnosis.

[0007] To achieve the above-mentioned technical objectives, the present invention provides the following technical solution: A bearing fault diagnosis method based on sample adaptive fractional attention Transformer specifically includes the following steps: S1. Collect vibration signal data during bearing operation and preprocess it; based on the preprocessed vibration signal data, construct bearing vibration samples to obtain a bearing fault diagnosis dataset, and divide the bearing fault diagnosis dataset into training set, validation set and test set. S2. Construct a convolutional token feature extraction module, taking the bearing vibration sample as input, to extract the initial local temporal features of the bearing vibration signal and form the initial token feature representation; S3. Construct a Token gating module to perform gating weighting on the initial Token feature representation to obtain gating enhanced features; S4. Construct a one-dimensional learnable position encoding module, add learnable position embedding to the gated enhancement features, and obtain input sequence features containing position information; S5. Construct a sample adaptive fractional-order Transformer encoder, input the input sequence features containing positional information into the sample adaptive fractional-order multi-head attention module and feedforward network module for feature encoding, and obtain deep temporal semantic features. S6. Input the deep temporal semantic features into the normalization layer and the Token attention pooling module to obtain global fault representation features; S7. The global fault characterization features are input into the feature projection layer for mapping, and then input into the linear classifier to output the bearing fault category prediction result. S8 consists of a bearing fault diagnosis model composed of a convolutional token feature extraction module, a token gating module, a one-dimensional learnable position encoding module, a sample adaptive fractional-order Transformer encoder, a layer normalization and token attention pooling module, a feature projection layer, and a linear classifier. The model is trained using a training set, the final model is determined using a validation set, and the bearing fault diagnosis performance of the final model is evaluated using a test set.

[0008] Furthermore, step S1 specifically includes: Vibration acceleration signals of the bearing under different operating conditions are collected by an accelerometer installed on the bearing housing to form a raw vibration dataset; the operating conditions include the normal state and the fault states corresponding to different fault locations and different fault degrees. Standardize or normalize the original vibration dataset; The preprocessed vibration signal is truncated or randomly cut according to a fixed sample length to generate vibration signal samples of a fixed length. Each sample is then assigned a corresponding fault category label to obtain the bearing fault diagnosis dataset. The bearing fault diagnosis dataset is divided into a training set, a validation set, and a test set. Gaussian noise is added to the training set samples.

[0009] Furthermore, step S2 specifically includes: One-dimensional convolution is used to extract the initial local temporal features of the bearing vibration signal from the bearing vibration sample; then batch normalization, GELU activation and dropout are performed on the initial local temporal features in sequence to obtain the initial token feature representation.

[0010] Furthermore, step S3 specifically includes: The initial token feature representation is input into the gating network to generate a gating coefficient matrix; the initial token features are then recalibrated using the gating coefficient matrix to obtain the gated enhancement features, as expressed by the formula: ; in, This represents the initial token characteristics. The gating coefficient matrix, For element-wise multiplication, This is a gating enhancement feature.

[0011] Furthermore, step S4 specifically includes: Construct a learnable position embedding matrix , where N represents the token sequence length and D represents the embedding dimension; The learnable location embedding matrix is ​​added to the gated enhancement features to obtain the input sequence features containing location information.

[0012] Furthermore, step S5 specifically includes: Linear mapping is performed on the input sequence features containing location information to generate a query matrix, a key matrix, and a value matrix; The query matrix, key matrix, and value matrix are split head-to-head, and the query matrix and key matrix are L2 normalized. Calculate the pairwise Euclidean distance between the L2-normalized query vector and the key vector; Extract the mean, standard deviation, and maximum value of the input sequence features to construct a statistical feature vector. ; The adaptive fractional-order parameter α for the sample is generated based on the statistical feature vector, and the formula is expressed as follows: ; in, This represents the lower bound of the sample adaptive fractional-order parameter. This represents the upper limit of the sample adaptive fractional-order parameter. This represents the Sigmoid activation function. AlphaNet represents a fractional-order parameter mapping function; Introducing learnable temperature parameters Based on the sample adaptive fractional-order parameter α and the learnable temperature parameter The attention score is constructed using the following formula: ; in, This indicates that after L2 normalization, the th The query vector corresponding to the query position is the first query position and the second query vector. Euclidean distance between the key vectors corresponding to each key position; This represents the attention score from position i to position j. The attention scores are then normalized using Softmax to obtain the attention weights. The value matrix is ​​weighted and summed using attention weights, and then the attention output features are obtained through multi-head concatenation and linear mapping. The attention output features are residually connected with the input sequence features to obtain intermediate features, which are then input into a feedforward network. The output of the feedforward network is then residually connected with the intermediate features to obtain deep temporal semantic features.

[0013] Furthermore, step S6 specifically includes: Calculate the importance score of each token in the deep temporal semantic feature H; The importance score is normalized to obtain the token pooling weight; By using the token pooling weights, the feature vectors corresponding to each token in HHH are weighted and summed to obtain the global fault representation features.

[0014] Furthermore, step S7 specifically includes: The global fault characterization features are input into the feature projection layer, and then passed through the fully connected layer, layer normalization, GELU activation and Dropout in sequence to obtain the final fault diagnosis features. The final fault diagnosis features are input into a linear classifier to obtain the predicted probability of each fault category; the category with the highest predicted probability is selected as the bearing fault diagnosis result.

[0015] Furthermore, step S8 specifically includes: On the training set, the bearing fault diagnosis model is trained using a multi-class cross-entropy loss function; the learnable parameters in the convolutional token feature extraction module, token gating module, one-dimensional learnable position encoding module, sample adaptive fractional-order Transformer encoder, layer normalization and token attention pooling module, feature projection layer and linear classifier are updated by backpropagation algorithm. Monitor the model classification accuracy on the validation set and save the model parameters corresponding to the highest classification accuracy on the validation set to obtain the final model; The fault diagnosis accuracy, loss value, classification report, confusion matrix, and comparison results of different models are calculated on the test set to evaluate the bearing fault diagnosis performance of the final model.

[0016] Based on the above technical solution, the present invention has at least the following beneficial effects: 1. This invention employs a single-layer convolutional token feature extraction module, which can effectively extract local temporal features of the original vibration signal while maintaining structural simplicity, thereby achieving efficient mapping from the original sequence to the token representation; 2. This invention introduces a Token gating module to adaptively recalibrate the features after convolutional embedding, which can enhance important features related to fault diagnosis, suppress redundant and noisy features, and improve the quality of feature representation. 3. This invention uses one-dimensional learnable positional encoding, which enables the model to explicitly perceive the relative and absolute positions of each token in the sequence, thus improving the ability to model temporal structures. 4. This invention proposes a sample-adaptive fractional-order multi-head attention mechanism, which dynamically generates fractional-order parameters by inputting feature statistics and adjusts the attention distribution by combining learnable temperature parameters. Compared with traditional fixed dot product attention, it has stronger sample adaptability and expression flexibility. 5. This invention replaces the traditional dot product similarity calculation with a distance-based attention modeling method, which can more meticulously characterize the differences between tokens and improve the modeling ability of local patterns and global correlations of bearing faults; 6. This invention uses Token attention pooling instead of simple average pooling, which can automatically focus on key time segments that contribute more to fault identification, thereby improving diagnostic performance and noise resistance. 7. The present invention has a simple overall structure, moderate parameters, and stable training, making it suitable for bearing condition monitoring and online fault diagnosis scenarios. Attached Figure Description

[0017] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 The overall flowchart of a bearing fault diagnosis method based on sample adaptive fractional attention Transformer provided by the present invention; Figure 2 This is a schematic diagram of the sample adaptive fractional-order multi-head attention module structure in the embodiment; Figure 3 This is a graph showing the accuracy of bearing fault classification using the method proposed in this invention. Figure 4 The graph shows the bearing fault classification loss versus training loss curves for the method proposed in this invention. Figure 5 This is a schematic diagram of the confusion matrix of bearing fault diagnosis results using the method proposed in this invention; Figure 6 This is a comparison chart of the accuracy of the method proposed in this invention and existing technologies. Detailed Implementation

[0018] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. This will allow for a full understanding of how the present application uses technical means to solve technical problems and achieve technical effects, and to facilitate its implementation.

[0019] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0020] Please refer to Figures 1-6 This illustration shows a specific implementation of this embodiment. This embodiment proposes a bearing fault diagnosis method based on a sample-adaptive fractional-order attention Transformer. This method takes the original bearing vibration sequence as input and achieves accurate identification of different bearing operating states and fault types through steps such as one-layer convolutional token feature extraction, token gating enhancement, one-dimensional learnable position encoding, improved Transformer deep modeling, token attention pooling, and linear classification.

[0021] Please refer to Figure 1 This embodiment proposes a bearing fault diagnosis method based on a sample adaptive fractional attention Transformer, which includes the following steps: S1. Collect vibration signal data during bearing operation and preprocess it; based on the preprocessed vibration signal data, construct bearing vibration samples to obtain a bearing fault diagnosis dataset, and divide the bearing fault diagnosis dataset into training set, validation set and test set. In a preferred embodiment, step S1 specifically includes: Vibration acceleration signals of the bearing under different operating conditions are collected by an accelerometer installed on the bearing housing to form a raw vibration dataset; the operating conditions include the normal state and the fault states corresponding to different fault locations and different fault degrees. Standardize or normalize the original vibration dataset to reduce the dimensional differences between different samples. The preprocessed vibration signal is truncated or randomly cut according to a fixed sample length to generate vibration signal samples of a fixed length. Each sample is then assigned a corresponding fault category label to obtain the bearing fault diagnosis dataset. The bearing fault diagnosis dataset is divided into a training set, a validation set, and a test set. Gaussian noise is added to the training set samples to enhance the model's noise resistance and generalization ability. In this embodiment, the vibration signal sample length is preferably set to 1024, so a single input sample can be represented as a 1024×1 one-dimensional vibration sequence.

[0022] S2. Construct a convolutional token feature extraction module, taking the bearing vibration sample as input, to extract the initial local temporal features of the bearing vibration signal and form the initial token feature representation; like Figure 1 As shown, the convolutional token feature extraction module is used to convert the original one-dimensional vibration sequence into token sequence features suitable for processing by the Transformer encoder. Compared with directly inputting the original sequence into the Transformer, this module can extract local temporal patterns first, while reducing the sequence length and subsequent attention computation complexity.

[0023] In a preferred embodiment, step S2 specifically comprises: The preprocessed bearing vibration sample is represented as follows Where B represents the batch size, L represents the sample length, and C represents the number of vibration signal channels; The initial local temporal features of the bearing vibration signal are extracted from the bearing vibration samples using one-dimensional convolution, as expressed by the formula: ; in, This represents the output of the d-th convolutional channel at time step t, where K represents the kernel size. This represents the weight of the d-th convolutional kernel at position r. This represents the value of the input signal at time step tr. This represents the bias term of the d-th convolutional kernel; The initial local temporal features are then subjected to batch normalization, GELU activation, and Dropout processing in sequence to obtain the initial token feature representation. Where N represents the number of tokens and D represents the embedding dimension.

[0024] In this embodiment, a one-dimensional convolution layer is preferably used as the tokenization feature extraction module. The kernel size can be set to 15, the stride can be set to 4, and the embedding dimension can be set to 64. Through this processing, the original one-dimensional vibration sequence with a length of 1024 is mapped to a shorter token sequence, thereby reducing the computational burden of the subsequent Transformer encoder while preserving the local impact and short-term fluctuation features.

[0025] S3. Construct a Token gating module to perform gating weighting on the initial Token feature representation to obtain gating enhanced features; Please refer to Figure 1 The Token gating module is located after the convolutional Token feature extraction module and is used to adaptively recalibrate the initial Token features. Since the bearing vibration signal contains both fault impact information and noise disturbance and irrelevant background fluctuations, it is necessary to dynamically weight the Token features so that the model pays more attention to the key segments and key feature dimensions related to fault diagnosis.

[0026] In a preferred embodiment, step S3 specifically comprises: Inputting the initial token feature representation into the gating network generates the gating coefficient matrix, as expressed by the formula: ; in, This represents a gated mapping function composed of fully connected layers. This represents the Sigmoid activation function. This is the gating coefficient matrix; The initial token features are recalibrated using the gating coefficient matrix to obtain the gated enhanced features, expressed by the formula: ; in, This represents the initial token characteristics. For element-wise multiplication, This is a gating enhancement feature.

[0027] In this embodiment, the following is adopted: An enhanced form, rather than directly adopting Features are compressed. This amplifies key token features while preserving original token information, preventing important fault information from being excessively suppressed during gating. For tokens containing significant shocks, abrupt changes, or local energy enhancements, the gating module can increase their response strength; for noisy segments or irrelevant fluctuation segments, the token gating module can reduce their impact on subsequent Transformer encoders.

[0028] S4. Construct a one-dimensional learnable position encoding module, add learnable position embedding to the gated enhancement features, and obtain input sequence features containing position information; The Transformer structure itself does not possess explicit position-aware capabilities, while the fault impacts, periodic fluctuations, and local abnormal responses in the bearing vibration signal are all closely related to time and position. Therefore, this embodiment introduces a one-dimensional learnable position code after the Token gating module, enabling the model to perceive the relative position and temporal order of each Token in the original vibration sequence.

[0029] In a preferred embodiment, step S4 specifically comprises: Construct a learnable position embedding matrix , where N represents the token sequence length (i.e. the number of tokens) and D represents the embedding dimension (i.e. the feature dimension of each token). The learnable location embedding matrix is ​​added to the gated augmented features to obtain the input sequence features containing location information. .

[0030] In this embodiment, the position encoding matrix P is a learnable parameter that can be automatically updated through backpropagation during model training. Compared with fixed sinusoidal position encoding, learnable position encoding can automatically adapt the importance of different token positions according to the actual distribution of bearing vibration signals, which helps to improve the model's ability to express local fault segments and global temporal structures.

[0031] S5. Construct a sample adaptive fractional-order Transformer encoder, input the input sequence features containing positional information into the sample adaptive fractional-order multi-head attention module and feedforward network module for feature encoding, and obtain deep temporal semantic features. Please refer to Figure 2The sample-adaptive fractional-order Transformer encoder is the core module of this embodiment. Unlike the dot-product similarity-based attention mechanism in traditional Transformers, this module constructs attention scores based on pairwise distances between query and key vectors, and introduces a sample-adaptive fractional-order parameter α and a learnable temperature parameter. This allows the attention distribution to be dynamically adjusted based on the statistical characteristics of different bearing vibration samples.

[0032] In a preferred embodiment, step S5 specifically includes: A linear mapping is performed on the input sequence features containing location information to generate a query matrix Q, a key matrix K, and a value matrix V; the formula is expressed as: ; in, , , These represent the learnable parameter matrices for the query, key, and value, respectively. The query matrix, key matrix, and value matrix are split head-to-head, and the query matrix and key matrix are L2 normalized. Calculate the pairwise Euclidean distance between the L2-normalized query vector and the key vector; In this embodiment, distance-based attention calculation is used instead of traditional dot product similarity calculation. For bearing fault diagnosis tasks, the distance between different tokens can more directly reflect the differences between local vibration modes, impact responses and fault characteristics, which is beneficial to improving the model's ability to distinguish complex fault modes.

[0033] Extract the mean (Mean), standard deviation (Std), and maximum value (Max) of the input sequence to form a statistical feature vector. ; The adaptive fractional-order parameter α for the sample is generated based on the statistical feature vector, and the formula is expressed as follows: ; in, This represents the lower bound of the sample adaptive fractional-order parameter. This represents the upper limit of the sample adaptive fractional-order parameter. This represents the Sigmoid activation function. AlphaNet represents a fractional-order parameter mapping function; , The preset value is used in this embodiment. It is 0.9. It is 1.3; In this embodiment, the sample adaptive fractional-order parameter α is not a fixed hyperparameter, but is dynamically generated by the statistical features of the current input sample. When the fault impact in the sample is significant and the local differences are large, the model can generate a larger α to enhance the sensitivity to token distance differences; when the sample noise is strong or the feature changes are relatively smooth, the model can generate a smaller α to make the attention distribution smoother and reduce the interference of noise on the model's judgment.

[0034] Introducing learnable temperature parameters In this embodiment, the formula for calculating the learnable temperature parameter is expressed as follows: ; Where ρ is a learnable parameter in the calculation process of the learnable temperature parameter. To prevent numerically unstable minimal constants, This is the softplus activation function; In this embodiment, the temperature parameter Used to adjust the smoothness of attention distribution; When the size is smaller, the attention distribution is more concentrated, and the model focuses more on a few key tokens; When the size is large, the attention distribution is smoother, and the model is able to take into account more token information.

[0035] Based on sample-adaptive fractional-order parameter α and learnable temperature parameter The attention score is constructed using the following formula: ; in, , Indicates the first The query vector corresponding to the position and the first position Euclidean distance between the key vectors corresponding to each position; and These represent the query vector and key vector at the corresponding positions, respectively. This represents the attention score from position i to position j. The attention scores are then subjected to Softmax normalization to obtain the attention weights, expressed by the following formula: ; in, This represents the attention weight of the i-th token to the j-th token; The value matrix is ​​weighted and summed using attention weights, and then the attention output features are obtained through multi-head concatenation and linear mapping. The attention output features are residually connected with the input sequence features to obtain intermediate features, which are then input into a feedforward network. The output of the feedforward network is then residually connected with the intermediate features to obtain deep temporal semantic features.

[0036] In this embodiment, the sample-adaptive fractional-order Transformer encoder can simultaneously model long-range dependencies and local differences between different tokens. The sample-adaptive fractional-order parameter α and the temperature parameter... The combined effect of these factors enables the model to dynamically adjust the attention distribution for different noise intensities, different fault categories, and different vibration modes, thereby improving the accuracy and robustness of fault diagnosis.

[0037] S6. Input the deep temporal semantic features into the normalization layer and the Token attention pooling module to obtain global fault representation features; Please refer to Figure 1 The token sequence output by the Transformer encoder still contains features at multiple time points. In order to obtain a global fault representation for final classification, this embodiment adopts the token attention pooling method to assign different importance weights to different tokens.

[0038] In a preferred embodiment, step S6 specifically includes: The importance score of each token in the deep temporal semantic feature H is calculated using the following formula: ; in, This represents the importance score of the t-th token; This represents the feature vector corresponding to the t-th token. and This represents the learnable parameters in the token attention pooling process. for Transpose of; This represents the GELU activation function; The importance score is normalized to obtain the token pooling weight, expressed by the formula: ; in, The attention weight of the t-th token; By using the token pooling weights to perform a weighted summation of the feature vectors corresponding to each token in H, the global fault representation features are obtained. .

[0039] In this embodiment, Token attention pooling is more selective than average pooling. For tokens containing fault impacts, obvious energy mutations, or discriminative vibration modes, the Token attention pooling module can assign higher weights. For tokens with irrelevant background vibrations or noise, lower weights are assigned, thereby improving the discriminativeness of the final fault characteristics.

[0040] S7. The global fault characterization features are input into the feature projection layer for mapping, and then input into the linear classifier to output the bearing fault category prediction result. In a preferred embodiment, step S7 specifically comprises: The global fault characterization features are input into the feature projection layer, and then sequentially passed through a fully connected layer, layer normalization, GELU activation, and Dropout to obtain the final fault diagnosis features. , This represents the projection mapping function consisting of a fully connected layer, layer normalization, GELU activation, and Dropout. The final fault diagnosis features are input into a linear classifier to obtain the predicted probability of each fault category, expressed by the formula: ; in, and These represent the weights and biases of the classifier, respectively; c represents the c-th fault category. The function is Softmax; the category with the highest predicted probability is selected as the bearing fault diagnosis result.

[0041] S8 consists of a convolutional token feature extraction module, a token gating module, a one-dimensional learnable position encoding module, a sample adaptive fractional-order Transformer encoder, a layer normalization and token attention pooling module, a feature projection layer, and a linear classifier, forming a bearing fault diagnosis model. The model is trained using a training set, the final model is determined using a validation set, and the bearing fault diagnosis performance of the final model is evaluated using a test set. In a preferred embodiment, step S8 specifically includes: The bearing fault diagnosis model is trained on the training set using a multi-class cross-entropy loss function; the formula for the multi-class cross-entropy loss function is as follows: ; Where M represents the number of training samples. This indicates the total number of bearing failure categories. This represents the true label of the i-th sample in class c. This represents the probability that the model predicts the i-th sample belongs to the c-th type of fault; The backpropagation algorithm is used to update the learnable parameters in the convolutional token feature extraction module, token gating module, one-dimensional learnable position encoding module, sample adaptive fractional-order Transformer encoder, layer normalization and token attention pooling module, feature projection layer and linear classifier. Monitor the model classification accuracy on the validation set and save the model parameters corresponding to the highest classification accuracy on the validation set to obtain the final model; The fault diagnosis accuracy, loss value, classification report, confusion matrix, and comparison results of different models are calculated on the test set to evaluate the bearing fault diagnosis performance of the final model.

[0042] This concludes the description of all steps in the method proposed in this invention. Furthermore, to verify the effectiveness, reliability, and superiority of the bearing fault diagnosis method based on a sample adaptive fractional attention Transformer proposed in this invention, the following experimental examples are provided: This method was implemented on the publicly available bearing failure dataset provided by CWRU. The CWRU dataset contains vibration acceleration signals of the bearings acquired by placing an accelerometer above the bearing housing at the motor drive end, with a sampling frequency of 12 kHz. All failed bearings in the CWRU dataset were subjected to single-point damage machining using electrical discharge machining. The CWRU publicly available bearing failure dataset includes three types of failed bearings and one type of normal bearing. The failure locations of the failed bearings are the outer ring, rolling elements, and inner ring, with failure diameters of 0.18, 0.36, and 0.54 mm. Specifically, Table 1 below shows all bearing failure types and their corresponding labels: Table 1. Fault types and labels in the CWRU and PT datasets

[0043] In this embodiment, the length of each bearing vibration sample is set to 1024, the number of samples is set to 250, and the training set, validation set, and test set are divided in a 6:2:2 ratio. The model training iterations are set to 50 times, the batch size is set to 8, the token embedding dimension can be set to 64, the feature projection dimension can be set to 128, the number of attention heads can be set to 4, and the Transformer encoder depth can be set to 1. The AdamW optimizer is used during model training, and the initial learning rate is set to 3×10. -4 The weight decay factor is set to 1×10. -4 To demonstrate the robustness and versatility of the model in complex industrial environments, Gaussian noise with a signal-to-noise ratio of 0dB was added to the training, validation, and test sets in the experiments to simulate actual working conditions such as sensor noise, electromagnetic interference, and mechanical disturbances.

[0044] As an explanation, Figure 3 and Figure 4 The accuracy and loss curves of the model of this invention during the training process are shown respectively. Figure 3 It can be seen that as the number of training iterations increases, the accuracy of both the training set and the test set gradually rises and remains at a high level, indicating that the model of this invention can effectively learn the fault discrimination features in bearing vibration signals. Figure 4 It can be seen that the training set loss and the test set loss decrease significantly and gradually stabilize as training progresses, indicating that the model training process is stable and the convergence effect is good.

[0045] Figure 5 The confusion matrix results of the model of this invention on the test set are shown. Figure 5 As can be seen, the model of the present invention has a high recognition rate in most fault categories, indicating that the model of the present invention can effectively distinguish different bearing fault categories.

[0046] Figure 6 The results show a comparison of the fault diagnosis accuracy of the model of this invention with other existing models. Figure 6 It can be seen that the model of this invention has a higher fault diagnosis accuracy than the comparative model, indicating that the proposed one-layer convolutional token feature extraction, token gating enhancement, one-dimensional learnable position encoding, and sample adaptive fractional attention mechanism can effectively improve the bearing fault diagnosis performance.

[0047] In summary, this embodiment performs sample segmentation, noise enhancement, and standardization on bearing vibration signals. It then utilizes a single-layer convolutional token feature extraction module, a token gating module, a one-dimensional learnable position encoding module, a sample adaptive fractional-order Transformer encoder, a layer normalization and token attention pooling module, a feature projection layer, and a linear classifier to construct a complete end-to-end bearing fault diagnosis process. This invention effectively enhances local impact information, global token dependencies, and sample adaptive feature representation capabilities in bearing vibration signals, exhibiting high classification accuracy, good noise resistance, and significant engineering application value.

[0048] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0049] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).

[0050] The above embodiments provide a detailed description of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A bearing fault diagnosis method based on sample adaptive fractional attention Transformer, characterized in that, Specifically, the following steps are included: S1. Collect vibration signal data during bearing operation and preprocess it; based on the preprocessed vibration signal data, construct bearing vibration samples to obtain a bearing fault diagnosis dataset, and divide the bearing fault diagnosis dataset into training set, validation set and test set. S2. Construct a convolutional token feature extraction module, taking the bearing vibration sample as input, to extract the initial local temporal features of the bearing vibration signal and form the initial token feature representation; S3. Construct a Token gating module to perform gating weighting on the initial Token feature representation to obtain gating enhanced features; S4. Construct a one-dimensional learnable position encoding module, add learnable position embedding to the gated enhancement features, and obtain input sequence features containing position information; S5. Construct a sample adaptive fractional-order Transformer encoder. Design sample adaptive fractional-order parameters based on the statistical features of the input feature sequence. Then, combine a multi-head attention mechanism and a feedforward residual network to encode the features of the input feature sequence to obtain deep temporal semantic features. S6. Input the deep temporal semantic features into the normalization layer and the Token attention pooling module to obtain global fault representation features; S7. The global fault characterization features are input into the feature projection layer for mapping, and then input into the linear classifier to output the bearing fault category prediction result. S8 consists of a bearing fault diagnosis model composed of a convolutional token feature extraction module, a token gating module, a one-dimensional learnable position encoding module, a sample adaptive fractional-order Transformer encoder, a layer normalization and token attention pooling module, a feature projection layer, and a linear classifier. The model is trained using a training set, the final model is determined using a validation set, and the bearing fault diagnosis performance of the final model is evaluated using a test set.

2. The bearing fault diagnosis method based on sample adaptive fractional attention Transformer according to claim 1, characterized in that, Step S1 specifically includes: Vibration acceleration signals of the bearing under different operating conditions are collected by an accelerometer installed on the bearing housing to form a raw vibration dataset; the operating conditions include the normal state and the fault states corresponding to different fault locations and different fault degrees. Standardize or normalize the original vibration dataset; The preprocessed vibration signal is truncated or randomly cut according to a fixed sample length to generate vibration signal samples of a fixed length. Each sample is then assigned a corresponding fault category label to obtain the bearing fault diagnosis dataset. The bearing fault diagnosis dataset is divided into a training set, a validation set, and a test set. Gaussian noise is added to the training set samples.

3. The bearing fault diagnosis method based on sample adaptive fractional attention Transformer according to claim 1, characterized in that, Step S2 is as follows: One-dimensional convolution is used to extract the initial local temporal features of the bearing vibration signal from the bearing vibration sample; then batch normalization, GELU activation and dropout are performed on the initial local temporal features in sequence to obtain the initial token feature representation.

4. The bearing fault diagnosis method based on sample adaptive fractional attention Transformer according to claim 1, characterized in that, Step S3 is as follows: The initial token feature representation is input into the gating network to generate a gating coefficient matrix; the initial token features are then recalibrated using the gating coefficient matrix to obtain the gated enhancement features, as expressed by the formula: ; in, This represents the initial token characteristics. The gating coefficient matrix, For element-wise multiplication, This is a gating enhancement feature.

5. The bearing fault diagnosis method based on sample adaptive fractional attention Transformer according to claim 1, characterized in that, Step S4 is as follows: Construct a learnable position embedding matrix , where N represents the token sequence length and D represents the embedding dimension; The learnable location embedding matrix is ​​added to the gated enhancement features to obtain the input sequence features containing location information.

6. The bearing fault diagnosis method based on sample adaptive fractional attention Transformer according to claim 1, characterized in that, Step S5 specifically includes: Linear mapping is performed on the input sequence features containing location information to generate a query matrix, a key matrix, and a value matrix; The query matrix, key matrix, and value matrix are split head-to-head, and the query matrix and key matrix are L2 normalized. Calculate the pairwise Euclidean distance between the L2-normalized query vector and the key vector; Extract the mean, standard deviation, and maximum value of the input sequence features to construct a statistical feature vector. ; The adaptive fractional-order parameter α for the sample is generated based on the statistical feature vector, and the formula is expressed as follows: ; in, This represents the lower bound of the sample adaptive fractional-order parameter. This represents the upper limit of the sample adaptive fractional-order parameter. This represents the Sigmoid activation function. AlphaNet represents a fractional-order parameter mapping function; Introducing learnable temperature parameters Based on the sample adaptive fractional-order parameter α and the learnable temperature parameter The attention score is constructed using the following formula: ; in, This indicates that after L2 normalization, the th The query vector corresponding to the query position is the first query position and the second query vector. Euclidean distance between the key vectors corresponding to each key position; This represents the attention score from position i to position j. The attention scores are then normalized using Softmax to obtain the attention weights. The value matrix is ​​weighted and summed using attention weights, and then the attention output features are obtained through multi-head concatenation and linear mapping. The attention output features are residually connected with the input sequence features to obtain intermediate features, which are then input into a feedforward network. The output of the feedforward network is then residually connected with the intermediate features to obtain deep temporal semantic features.

7. The bearing fault diagnosis method based on sample adaptive fractional attention Transformer according to claim 1, characterized in that, Step S6 specifically includes: Calculate the importance score of each token in the deep temporal semantic feature H; The importance score is normalized to obtain the token pooling weight; By using the token pooling weights, the feature vectors corresponding to each token in H are weighted and summed to obtain the global fault representation features.

8. The bearing fault diagnosis method based on sample adaptive fractional attention Transformer according to claim 1, characterized in that, Step S7 is as follows: The global fault characterization features are input into the feature projection layer, and then passed through the fully connected layer, layer normalization, GELU activation and Dropout in sequence to obtain the final fault diagnosis features. The final fault diagnosis features are input into a linear classifier to obtain the predicted probability of each fault category; the category with the highest predicted probability is selected as the bearing fault diagnosis result.

9. The bearing fault diagnosis method based on sample adaptive fractional attention Transformer according to claim 1, characterized in that, Step S8 specifically includes: On the training set, the bearing fault diagnosis model is trained using a multi-class cross-entropy loss function; the learnable parameters in the convolutional token feature extraction module, token gating module, one-dimensional learnable position encoding module, sample adaptive fractional-order Transformer encoder, layer normalization and token attention pooling module, feature projection layer and linear classifier are updated by backpropagation algorithm. Monitor the model classification accuracy on the validation set and save the model parameters corresponding to the highest classification accuracy on the validation set to obtain the final model; The fault diagnosis accuracy, loss value, classification report, confusion matrix, and comparison results of different models are calculated on the test set to evaluate the bearing fault diagnosis performance of the final model.