Bearing fault diagnosis method and device based on multi-scale space-time fusion and medium

By adopting a multi-scale spatiotemporal fusion method in bearing fault diagnosis, combining fast Fourier transform, variational modal decomposition, swintransformer and global context attention mechanism TCN network, extracting and fusing the local and global features of bearing signals, the problem of insufficient diagnostic accuracy and robustness in the existing technology is solved, and efficient and accurate bearing fault diagnosis is achieved.

CN120011747APending Publication Date: 2025-05-16SHANGHAI UNIV OF ENG SCI
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510054655.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing bearing fault diagnosis technology is difficult to effectively extract the characteristics of fault signals in low signal-to-noise ratio and noise environments, and the commonly used two-dimensional processing methods will lose important timing information, resulting in insufficient diagnostic accuracy and robustness.

Method used

The bearing fault diagnosis method based on multi-scale space-time fusion is adopted, and comprehensive features are extracted through fast Fourier transform and variational modal decomposition, combined with swintransformer and TCN network based on global context attention mechanism, local and global features are extracted, and fusion is carried out through adaptive average pooling to finally output the failure probability of the bearing.

Benefits of technology

It improves the accuracy and robustness of bearing fault diagnosis, can efficiently identify fault signals in low signal-to-noise ratio and noise environments, and enhances the ability to capture long sequence information and generalize models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011747A_ABST
    Figure CN120011747A_ABST
Patent Text Reader

Abstract

The invention relates to a bearing fault diagnosis method and device based on multi-scale space-time fusion and a medium, and the method comprises the following steps: obtaining bearing state data, carrying out the preprocessing and comprehensive feature extraction, inputting a trained bearing fault diagnosis model based on multi-scale space-time fusion, and obtaining a corresponding fault diagnosis result; the comprehensive feature extraction is carried out based on fast Fourier transform and variational mode decomposition; in the bearing fault diagnosis model based on multi-scale space-time fusion, firstly, comprehensive features are input into a swingtransform, local features are extracted through a window attention mechanism, meanwhile, the comprehensive features are input into a TCN network based on a global context attention mechanism, and global features are extracted; then, the local features and the global features are fused through adaptive average pooling; and finally, outputting the fault probability of the bearing according to the fusion features. Compared with the prior art, an accurate diagnosis result can be efficiently obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine fault diagnosis, and in particular relates to a bearing fault diagnosis method, device and medium based on multi-scale time-space fusion. Background Art

[0002] The normal operation of machinery and equipment is the basis for ensuring production efficiency and safety, while the failure of key machine components may cause the equipment to fail to operate normally, and in severe cases may even cause shutdown accidents, which not only cause physical damage to the equipment, but may also lead to production stagnation, economic losses, and even endanger personal safety. Therefore, ensuring the stability and reliability of equipment operation has become a core issue in modern industrial management and operation and maintenance. Among various types of machinery and equipment, bearings are widely used in many fields such as industry, transportation, energy, aerospace, etc., and their quality and performance directly affect the stability and efficiency of rotating machinery. Since bearings play a key role in equipment, their damage will have a chain reaction on the overall operation of the machine. Therefore, how to achieve accurate and timely bearing fault detection to ensure the reliability and safety of its operation is a crucial technical issue in the current industrial field.

[0003] The machine learning-based fault diagnosis model can assist equipment maintenance personnel in bearing fault detection. Although this technology has significant potential and advantages, it currently has the following drawbacks:

[0004] 1. Traditional signal preprocessing methods usually rely on a single frequency or time analysis technique, such as Fourier transform or wavelet transform, which makes it difficult to effectively extract rich features from fault signals at different scales and accurately separate time-frequency features in complex environments, especially in the case of low signal-to-noise ratio and the presence of noise, which affects the accuracy and robustness of fault identification.

[0005] 2. Common fault detection models convert one-dimensional signals into two-dimensional images before processing, which increases feature complexity and may lose some important timing information during the conversion process. This two-dimensional processing method makes it difficult for the model to accurately capture the details of the original signal, resulting in increased errors during model training and difficulty in obtaining efficient feature expression.

[0006] 3. Existing fault diagnosis models are mostly built based on RNN or LSTM. Although they can model sequence information, they may face the problem of gradient vanishing or exploding when processing long sequence data and cannot fully capture patterns on different time scales.

[0007] Therefore, it is necessary to design a new bearing fault diagnosis method to efficiently obtain accurate diagnostic results. Summary of the invention

[0008] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and to provide a bearing fault diagnosis method, equipment and medium based on multi-scale time-space fusion, which can efficiently obtain accurate diagnosis results.

[0009] The purpose of the present invention can be achieved by the following technical solutions:

[0010] The present invention provides a bearing fault diagnosis method based on multi-scale time-space fusion, comprising the following steps:

[0011] Obtain bearing status data, perform preprocessing and comprehensive feature extraction, and then input the trained bearing fault diagnosis model based on multi-scale spatiotemporal fusion to obtain the corresponding fault diagnosis results;

[0012] The comprehensive feature extraction is performed based on fast Fourier transform and variational mode decomposition; in the bearing fault diagnosis model based on multi-scale spatiotemporal fusion, the comprehensive features are first input into swintransformer, and local features are extracted through the window attention mechanism. At the same time, the comprehensive features are input into the TCN network based on the global context attention mechanism to extract global features; then, the local features and the global features are fused through adaptive average pooling; finally, the failure probability of the bearing is output according to the fused features.

[0013] Furthermore, the specific process of the comprehensive feature extraction is as follows:

[0014] S11. Perform fast Fourier transform on the input signal to obtain the overall frequency domain characteristics of the input signal. The calculation formula is:

[0015]

[0016] Among them, X(f) is the frequency domain signal, x(n) is the filtered time domain signal, and N is the number of sampling points;

[0017] S12, performing variational mode decomposition on the input signal, extracting the signal components of each frequency band, and obtaining the time-frequency characteristics of each frequency band of the input signal. The optimization goal of the variational mode decomposition is specifically:

[0018]

[0019] Where K is the number of modes, u k is the kth modal signal, ω k for u k The center frequency of , α is the regularization parameter;

[0020] S13, combining the overall frequency domain features of the input signal and the time-frequency features of each frequency band of the input signal to generate a comprehensive feature set:

[0021] Xstacked =[X FFT ,u1,u2,…,u K ]

[0022] Among them, X stacked is the stacked feature matrix.

[0023] Furthermore, the data processing process of the TCN network based on the global context attention mechanism is as follows:

[0024] S21. Process input data through TCN network:

[0025]

[0026] Among them, x t-l and t are input and output respectively, t represents the time, L is the length of the convolution kernel, and w l is the convolution kernel;

[0027] S22. Aggregate global information using global contextual attention mechanism:

[0028]

[0029] Among them, α j is the weight of the global attention pool, y i represents the feature of the i-th position of the input sequence, that is, the i-th position feature of the TCN output, z i represents the feature of the i-th position of the output sequence, δ(·) represents the bottleneck transformation, and N p is the length of the input sequence.

[0030] Furthermore, the specific expression of the bottleneck transformation δ(·) is as follows:

[0031] δ(·)=W v2 ReLU(LN(W v1 (·)))

[0032] Among them, W v2 is a linear transformation weight matrix used to map the input to a hidden space, W v1 is a linear transformation weight matrix used to map the input mapped to the hidden layer space back to the original space or another target space, LN is layer normalization, and ReLU is the activation function.

[0033] Furthermore, the weight α of the global attention pool j The specific expression is as follows:

[0034]

[0035] Among them, Wk is the linear transformation weight in the global attention mechanism, which is used to adjust the input feature x j Project it and calculate its importance in the global context.

[0036] Furthermore, the bearing fault diagnosis model based on multi-scale spatiotemporal fusion is trained by minimizing the loss function.

[0037] Furthermore, the specific formula of the loss function is as follows:

[0038]

[0039] Among them, N is the number of samples, K is the number of categories, and y ic is the one-hot encoding of the sample target value, h θ (x i ) c is the observed sample x i The predicted probability of belonging to class c.

[0040] Furthermore, when training the bearing fault diagnosis model based on multi-scale spatiotemporal fusion, the specific process of the preprocessing is as follows:

[0041] The bearing status data set is obtained and normalized, and then the sample points are divided according to the set time step and overlap rate, and divided into training set, validation set, and test set according to the set ratio.

[0042] The present invention also provides an electronic device, comprising a memory, a processor, and a program stored in the memory, wherein the processor implements the above method when executing the program.

[0043] The present invention also provides a computer-readable storage medium on which a computer program is stored, and the program implements the above method when executed by a processor.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] 1. The present invention proposes a bearing fault diagnosis method based on multi-scale time-space fusion, which pre-processes and extracts comprehensive features of the acquired bearing status data, and then inputs the trained bearing fault diagnosis model based on multi-scale time-space fusion to obtain the corresponding fault diagnosis results; the comprehensive feature extraction is based on fast Fourier transform and variational mode decomposition, wherein the overall frequency domain characteristics of the input signal can be obtained by fast Fourier transform, and the time-frequency characteristics of each frequency band of the input signal can be obtained by variational mode decomposition. Combining the above two methods can mine the multi-scale features in the input signal; in the bearing fault diagnosis model based on multi-scale time-space fusion, first The comprehensive features are input into swintransformer, and local features are extracted through the window attention mechanism. At the same time, the comprehensive features are input into the TCN network based on the global context attention mechanism to extract global features. The above design can improve the model's ability to capture long sequence information; then, the local features and global features are fused through adaptive average pooling, so that the model can better integrate feature representations at different levels, improve model performance and generalization ability; finally, the failure probability of the bearing is output according to the fused features; the above method can efficiently and accurately identify fault signals, further improving the accuracy and robustness of bearing fault diagnosis.

[0046] 2. In the bearing fault diagnosis model based on multi-scale spatiotemporal fusion proposed in the present invention, a TCN network based on the global contextual attention mechanism is introduced. TCN adopts convolution operation, which can perform efficient parallel calculations and effectively capture local dependencies in sequence data. Compared with traditional RNN, the gradient propagation of TCN in the training process is more stable and is not prone to gradient vanishing and gradient explosion problems. The global contextual attention mechanism is applied at the output of TCN to dynamically adjust the influence of different features by aggregating global information, ensuring that the model can take into account both local and global information, further improving the accuracy of the detection results. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is a flow chart of the method of the present invention;

[0048] Figure 2 It is a schematic diagram of the structure of the global context attention mechanism.

[0049] Among them, r is the bottleneck ratio, C / r ​​represents the hidden representation dimension of the bottleneck, and W v1 and W v2 is the weight matrix of the linear transformation, W k is the linear transformation weight in the global attention mechanism. DETAILED DESCRIPTION

[0050] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0051] Example:

[0052] This embodiment provides a bearing fault diagnosis method based on multi-scale time-space fusion. After preprocessing and comprehensive feature extraction of bearing status data, the bearing fault diagnosis model based on multi-scale time-space fusion is input into the trained bearing fault diagnosis model to obtain the corresponding fault diagnosis result. Figure 1 As shown, the specific process is as follows:

[0053] S1. Comprehensive feature extraction based on fast Fourier transform and variational mode decomposition. The specific process is as follows:

[0054] S11. Perform fast Fourier transform (FFT) on the input signal to obtain the overall frequency domain characteristics of the input signal. The specific calculation formula is:

[0055]

[0056] Among them, X(f) is the frequency domain signal, x(n) is the filtered time domain signal, and N is the number of sampling points.

[0057] S12. Perform variational mode decomposition (VMD) on the input signal, extract the signal components of each frequency band, and obtain the time-frequency characteristics of each frequency band of the input signal. The optimization objectives of variational mode decomposition are:

[0058]

[0059] Where K is the number of modes, u k is the kth modal signal, ω k for u k is the center frequency of , and α is the regularization parameter.

[0060] S13, combining the overall frequency domain features of the input signal and the time-frequency features of each frequency band of the input signal to generate a comprehensive feature set:

[0061] X stacked =[X FFT ,u1,u2,…,u K ]

[0062] Among them, X stacked is the stacked feature matrix.

[0063] S2. Input the comprehensive features into swintransformer and extract local features through the window attention mechanism; at the same time, input the comprehensive features into the TCN network based on the global context attention mechanism to extract global features.

[0064] The calculation method of the swintransformer window attention mechanism is as follows:

[0065]

[0066] Among them, Q represents the query matrix used to query the importance of each feature in the local window, K represents the key matrix, which is used to calculate the attention weight with the query matrix Q, and V represents the value matrix, which is used to generate the final local features under the action of the attention weight. k The dimension of the key.

[0067] In this embodiment, Q, K, and V are all derived from the comprehensive feature X. stacked (Simplified as X in the following formula) is obtained by linear transformation:

[0068] Q=XW Q ,K=XW K ,V=XW V

[0069] Among them, W Q , W K , W V is a learnable weight matrix

[0070] The data processing process of the TCN network based on the global context attention mechanism is as follows:

[0071] S21. Process the input data through the TCN network (time domain convolution network) and use a series of time convolutions to extract time series features:

[0072]

[0073] Among them, x t-l and t are input (i.e., comprehensive features) and output, t represents the time, L is the length of the convolution kernel, and w l is the convolution kernel.

[0074] S22. The global context attention mechanism (GCNet) is used to aggregate global information on the output of the TCN network. The global context attention mechanism is as follows Figure 2 As shown, the formula expression is as follows:

[0075]

[0076] δ(·)=W v2ReLU(LN(W v1 (·)))

[0077]

[0078] Among them, α j is the weight of the global attention pool, y i represents the feature of the i-th position of the input sequence, that is, the i-th position feature of the TCN output, z i Represents the feature of the i-th position of the output sequence, N p is the length of the input sequence, δ(·) represents the bottleneck transformation, W v2 is a linear transformation weight matrix, which maps the input to a hidden space, W v1 is a linear transformation weight matrix, which is used to map the input mapped to the hidden space back to the original space or another target space. LN is layer normalization, ReLU is the activation function, and W k is the linear transformation weight in the global attention mechanism, which is used to adjust the input feature x j Project it and calculate its importance in the global context. Through the global context attention mechanism, the influence of different features can be dynamically adjusted to ensure that the model can take into account both local and global information.

[0079] S3, local features and global features are fused through adaptive average pooling to integrate feature representations at different levels.

[0080] S4. Output the failure probability of the bearing based on the fused features.

[0081] The bearing fault diagnosis model based on multi-scale spatiotemporal fusion is trained by minimizing the loss function. The specific formula of the loss function is as follows:

[0082]

[0083] Among them, N is the number of samples, K is the number of categories, and y ic is the one-hot encoding of the sample target value, h θ (x i ) c is the observed sample x i The predicted probability of belonging to class c.

[0084] Compared with the prior art, the above method has the following advantages:

[0085] (1) Combining fast Fourier transform and variational mode decomposition to extract time-frequency and domain features of signals can mine multi-scale features in fault signals.

[0086] (2) The window attention mechanism of swintransformer is used to extract local features of fault signals, and the temporal convolutional network based on global context attention is used to extract global features to improve the ability to capture long sequence information.

[0087] (3) The local features and global spatial features are fused through adaptive average pooling respectively, so that the model can better integrate feature representations at different levels and improve model performance and generalization ability.

[0088] To verify the effectiveness of the above method, this embodiment uses a bearing dataset from Case Western Reserve University for experiments. This dataset is collected through experiments under different load and speed conditions, contains rich bearing vibration signals, and provides bearing fault vibration data under various working conditions to simulate different bearing states (normal, inner ring fault, outer ring fault, and rolling element fault).

[0089] When the bearing fault diagnosis model based on multi-scale spatiotemporal fusion is trained with the above data set, the specific preprocessing process is as follows: normal baseline data and outer ring fault data with fault diameters of 0.007, 0.014, and 0.021 are selected under loads of 0, 1, and 2. After normalization, the sample points are divided according to a time step of 1024 and an overlap rate of 0.5, and the training set, validation set, and test set are divided according to a ratio of 7:2:1.

[0090] The trained bearing fault diagnosis model based on multi-scale spatiotemporal fusion is analyzed and compared with some mainstream methods currently used in fault diagnosis, and the results are shown in Table 1. It can be found from Table 1 that LSTM, as a classic network for processing time series data, has good time-dependent modeling capabilities, but its accuracy is only 84.71%, which is average in this task; CNN, with its advantages in image processing and feature extraction, has improved the accuracy to 96.46%; TCN, based on the temporal convolutional network, captures long sequence relationships through causal convolution, and the accuracy reaches 95.73%; TICNN combines the advantages of temporal convolution and convolutional network to improve the accuracy to 99.01%. However, the method proposed in this invention (OURS) has achieved an accuracy of 100%, which shows that this method has significant advantages in extracting and fusing local and global features, can accurately identify bearing fault signals, and further improves the accuracy and robustness of diagnosis.

[0091] Table 1 Experimental results of Case Western Reserve University bearing dataset

[0092] LSTM CNN TCN TICNN OURS Accuracy (%) 84.71 96.46 95.73 99.01 100.00

[0093] If the above method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0094] The above description of the embodiments is to facilitate the understanding and use of the invention by those skilled in the art. It is obvious that those skilled in the art can easily make various modifications to these embodiments and apply the general principles described herein to other embodiments without creative work. Therefore, the present invention is not limited to the above embodiments, and improvements and modifications made by those skilled in the art based on the disclosure of the present invention without departing from the scope of the present invention should be within the scope of protection of the present invention.

Claims

1. A bearing fault diagnosis method based on multi-scale spatiotemporal fusion, characterized in that: The following steps are involved: Obtain bearing status data, perform preprocessing and comprehensive feature extraction, and then input the trained bearing fault diagnosis model based on multi-scale spatiotemporal fusion to obtain the corresponding fault diagnosis results; The comprehensive feature extraction is performed based on fast Fourier transform and variational mode decomposition; in the bearing fault diagnosis model based on multi-scale spatiotemporal fusion, the comprehensive features are first input into swintransformer, local features are extracted through the window attention mechanism, and the comprehensive features are simultaneously input into the TCN network based on the global context attention mechanism to extract global features; Then, the local features and the global features are fused through adaptive average pooling; finally, the failure probability of the bearing is output according to the fused features.

2. A bearing fault diagnosis method based on multi-scale spatiotemporal fusion according to claim 1, characterized in that: The specific process of comprehensive feature extraction is as follows: S11. Perform fast Fourier transform on the input signal to obtain the overall frequency domain characteristics of the input signal. The calculation formula is: Among them, X(f) is the frequency domain signal, x(n) is the filtered time domain signal, and N is the number of sampling points; S12, performing variational mode decomposition on the input signal, extracting the signal components of each frequency band, and obtaining the time-frequency characteristics of each frequency band of the input signal. The optimization goal of the variational mode decomposition is specifically: Where k is the number of modes, u k is the kth modal signal, ω k for u k The center frequency of , α is the regularization parameter; S13, combining the overall frequency domain features of the input signal and the time-frequency features of each frequency band of the input signal to generate a comprehensive feature set: X stacked =[X FFT ,u1,u2,…,u K ] Among them, X stacked is the stacked feature matrix.

3. The bearing fault diagnosis method based on multi-scale spatiotemporal fusion according to claim 1 is characterized in that: The data processing process of the TCN network based on the global context attention mechanism is as follows: S21. Process input data through TCN network: Among them, x t-l and t are input and output respectively, t represents the time, L is the length of the convolution kernel, and w l is the convolution kernel; S22. Aggregate global information using global contextual attention mechanism: Among them, α j is the weight of the global attention pool, y i represents the feature of the i-th position of the input sequence, that is, the i-th position feature of the TCN output, z i represents the feature of the i-th position of the output sequence, δ(·) represents the bottleneck transformation, and N p is the length of the input sequence.

4. The bearing fault diagnosis method based on multi-scale spatiotemporal fusion according to claim 3 is characterized in that: The specific expression of the bottleneck transformation δ(·) is as follows: δ( )=W v2 ReLU(LN(W v1 (·))) Among them, W v2 is a linear transformation weight matrix used to map the input to a hidden space, W v1 is a linear transformation weight matrix used to map the input mapped to the hidden layer space back to the original space or another target space, LN is layer normalization, and ReLU is the activation function.

5. The bearing fault diagnosis method based on multi-scale spatiotemporal fusion according to claim 3 is characterized in that: The weight α of the global attention pool j The specific expression is as follows: Among them, W k is the linear transformation weight in the global attention mechanism, which is used to adjust the input feature x j Project it and calculate its importance in the global context.

6. The bearing fault diagnosis method based on multi-scale spatiotemporal fusion according to claim 1 is characterized in that: The bearing fault diagnosis model based on multi-scale spatiotemporal fusion is trained by minimizing the loss function.

7. The bearing fault diagnosis method based on multi-scale spatiotemporal fusion according to claim 6 is characterized in that: The specific formula of the loss function is as follows: Among them, N is the number of samples, K is the number of categories, and y ic is the one-hot encoding of the sample target value, h θ (x i ) c is the observed sample x i The predicted probability of belonging to class c.

8. The bearing fault diagnosis method based on multi-scale spatiotemporal fusion according to claim 1 is characterized in that: When training the bearing fault diagnosis model based on multi-scale spatiotemporal fusion, the specific process of the preprocessing is as follows: The bearing status data set is obtained and normalized, and then the sample points are divided according to the set time step and overlap rate, and divided into training set, validation set, and test set according to the set ratio.

9. An electronic device comprising a memory, a processor, and a program stored in the memory, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Converter fault diagnosis method and system

    CN120744714A

  • Rolling bearing lightweight fault diagnosis method and system based on joint learning strategy

    CN121071769A