Variable working condition bearing fault diagnosis method based on window global mixed attention mechanism and improved convolutional neural network
By adopting a window global hybrid attention mechanism and an improved convolutional neural network in bearing fault diagnosis, combined with transfer learning strategy, the problem that the single attention mechanism is difficult to capture fault characteristics and complex and changeable working conditions is solved, and efficient and accurate bearing fault diagnosis is achieved.
Patent Information
- Application Number
- CN202510207361.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
In the existing bearing fault diagnosis technology, it is difficult to comprehensively and meticulously capture the complete characteristic information of the fault mode. In the actual industrial environment, the bearing operating conditions are complex and changeable, and the number of fault samples is limited, making it difficult to achieve accurate and efficient diagnosis.
The variable-condition bearing fault diagnosis method based on the window global hybrid attention mechanism and an improved convolutional neural network is adopted. By constructing a diagnostic model, combining the window attention mechanism and the global attention mechanism, the local and global characteristics of the bearing signal are extracted, and precise diagnosis across different working conditions is achieved through transfer learning strategies.
It improves the accuracy and efficiency of bearing fault diagnosis, and can effectively identify and diagnose bearing faults under small sample conditions, reduces the amount of calculation parameters and improves the accuracy of model fault identification.
Smart Images

Figure CN120145141A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of bearing fault diagnosis, and particularly relates to a variable-condition bearing fault diagnosis method based on a window global hybrid attention mechanism and an improved convolutional neural network. Background Art
[0002] In the existing bearing fault diagnosis, there are limitations of the single attention mechanism in the process of fault feature extraction. Its separate application often makes it difficult to comprehensively and finely capture the complete feature information of the fault mode, thus affecting the accuracy and efficiency of fault diagnosis. At the same time, due to the high complexity and variability of the bearing operating conditions in the actual industrial environment, and the relatively limited number of available fault samples, this situation poses a significant challenge to accurately and efficiently identifying and diagnosing bearing faults. Summary of the Invention
[0003] To solve the above technical problems, the present invention proposes a variable-condition bearing fault diagnosis method based on a window global hybrid attention mechanism and an improved convolutional neural network to solve the problems existing in the above prior art.
[0004] To achieve the above object, the present invention provides a variable-condition bearing fault diagnosis method based on a window global hybrid attention mechanism and an improved convolutional neural network, including:
[0005] Obtain bearing signals under different working conditions and the target working condition, preprocess the bearing signals to generate signal image data under different working conditions and the target working condition;
[0006] Construct a diagnosis model, in which a window attention mechanism and a global attention mechanism are set to extract the local features and global features corresponding to the signal image data respectively for subsequent recognition;
[0007] Pre-train the diagnosis model with the signal image data under different working conditions to obtain a pre-trained model, and train the pre-trained model with the signal image data under the target working condition to obtain an optimal diagnosis model;
[0008] Obtain the bearing signal to be measured under the target working condition, and diagnose and identify the bearing signal to be measured through the optimal diagnosis model to obtain the fault diagnosis result of the bearing signal to be measured.
[0009] Optionally, the preprocessing method uses wavelet transform conversion.
[0010] Optionally, the diagnosis model includes a depthwise separable convolutional layer module, an attention mechanism module, a fusion module, and an identification module connected in sequence, where the attention mechanism module includes a parallel window attention mechanism and a global attention mechanism, which are respectively connected to the convolutional module and the fusion module in parallel.
[0011] Optionally, the window attention mechanism includes a window structure at different time steps, wherein the window structure includes a first LN layer, a W-MSA layer, a second LN layer, and an MLP layer connected in sequence, wherein the sum of the input data of the first LN layer and the output data of the W-MSA layer is used as the input data of the second LN layer, and the sum of the input data of the second LN layer and the output data of the MLP layer is used as the input data of the window structure of the next time step.
[0012] Optionally, the global attention mechanism includes a channel attention mechanism and a spatial attention mechanism connected in sequence, wherein the sum of the input data and the output data of the channel attention mechanism serves as the input data of the spatial attention mechanism, and the sum of the input data and the output data of the spatial attention mechanism serves as the output data of the global attention mechanism.
[0013] Optionally, the recognition module includes a fully connected layer and an output layer connected in sequence, wherein the output layer adopts a softmax function.
[0014] Compared with the prior art, the present invention has the following advantages and technical effects:
[0015] (1) In view of the limitations of the single attention mechanism in the process of fault feature extraction, its single application is often difficult to comprehensively and finely capture the complete feature information of the fault mode, thus affecting the accuracy and efficiency of fault diagnosis. The present invention constructs a convolutional neural network model and inserts a global attention mechanism (GAM) and a window attention mechanism based on SwinTransformer to form a bearing fault diagnosis model with a hybrid attention mechanism.
[0016] (2) Due to the high complexity and variability of bearing operating conditions in actual industrial environments, coupled with the relatively limited number of available fault samples, this situation poses a significant challenge to the accurate and efficient identification and diagnosis of bearing faults. Therefore, this paper introduces a model-based deep transfer learning strategy to achieve accurate diagnosis transfer across different operating conditions, thereby effectively achieving the goal of efficient identification and diagnosis of bearing faults under small sample conditions.
[0017] (3) Based on the CNN model, the traditional convolutional layer is replaced by the deep separable convolution to construct a deep separable convolutional network (DSCNN). This method also maintains efficient feature extraction capabilities, which makes DSCNN well integrated into the model, reducing the amount of calculation parameters and improving the model fault recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0019] Figure 1 Schematic diagram of the global attention mechanism structure of the embodiment of the present invention;
[0020] Figure 2 Window attention mechanism of the embodiment of the present invention;
[0021] Figure 3 Schematic diagram of the principle of the depthwise separable convolution layer of the embodiment of the present invention;
[0022] Figure 4 Schematic diagram of the principle of transfer learning of the embodiment of the present invention;
[0023] Figure 5 Results of the model test diagnosis accuracy under three working conditions in the experiment on the publicly available dataset of Case Western Reserve University of the embodiment of the present invention;
[0024] Figure 6 Results of the model test diagnosis accuracy of the cross-working condition transfer model in the experiment on the publicly available dataset of Case Western Reserve University of the embodiment of the present invention;
[0025] Figure 7 Bearing fault detection experimental platform of the embodiment of the present invention;
[0026] Figure 8 Results of the model test diagnosis accuracy under three working conditions on the bearing fault detection test bench of the embodiment of the present invention;
[0027] Figure 9 Results of the model test diagnosis accuracy of the cross-working condition transfer model under the bearing fault detection test bench of the embodiment of the present invention;
[0028] Figure 10 Schematic diagram of the model structure of the embodiment of the present invention;
[0029] Figure 11 Schematic diagram of the entire fault diagnosis process of the present invention. Detailed implementation manners
[0030] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.
[0031] It should be noted that the steps shown in the flowchart of the drawings may be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0032] In this embodiment, a variable-condition bearing fault diagnosis method based on a window global hybrid attention mechanism and an improved convolutional neural network is provided. The bearing fault is mainly diagnosed through a diagnostic model after transfer learning. The main contents related to the bearing fault diagnosis include the following:
[0033] (1) The Depthwise Separable Convolutional Neural Networks (DSCNN) is an effective improvement based on the Convolutional Neural Network (CNN). This method inherits the basic structure and principle of CNN, such as the combination of convolutional layer, pooling layer, fully connected layer, etc. By introducing Depthwise Separable Convolution, the complexity and computational requirements of the model are significantly reduced, while maintaining efficient feature extraction capabilities. This enables DSCNN to be well integrated into this model, reducing the number of computational parameters and improving the model's fault recognition accuracy.
[0034] (2) The Global Attention Mechanism is composed of a Channel Attention Mechanism (CAM) and a Spatial Attention Mechanism (SAM). It improves the performance of deep neural networks by reducing information reduction and amplifying global interaction representations. It is connected relatedly through a residual structure. Specifically, the global attention mechanism includes a sequentially connected channel attention mechanism and a spatial attention mechanism. The sum of the input data and output data of the channel attention mechanism is used as the input data of the spatial attention mechanism, and the sum of the input data and output data of the spatial attention mechanism is used as the output data of the global attention mechanism. Its structural schematic diagram is as Figure 1 shown.
[0035] (3) Swin Transformer uses a hierarchical construction method similar to that in convolutional neural networks, which helps to construct tasks such as object detection and instance segmentation. In Swin Transformer, the concept of Windows Multi-Head Self-Attention (W-MSA) is used to reduce the computational amount. At the same time, the concept of Shifted Windows Multi-Head Self-Attention (SW-MSA) is used. Through this method, information can be transmitted in adjacent windows. The present invention uses the moving window in Swin Transformer to construct a window attention mechanism to extract local features of image data, such as Figure 2As shown, the window attention mechanism includes window structures at different time steps, where the window structure includes a first LN layer, a W-MSA layer, a second LN layer, and an MLP layer connected in sequence. The sum of the input data of the first LN layer and the output data of the W-MSA layer is used as the input data of the second LN layer, and the sum of the input data of the second LN layer and the output data of the MLP layer is used as the input data of the window structure at the next time step.
[0036] (4) Transfer learning refers to a learning process that utilizes the similarity between data, tasks, or models to apply a model learned in an old domain to a new domain.
[0037] The present invention uses the method of model transfer to share the parameters of the training model and achieve the transfer classification task between different working conditions.
[0038] Combined with the above related content, the present invention provides the following technical solutions. The specific steps include:
[0039] (1) First, the present invention preprocesses the collected bearing signals and converts them into two-dimensional image data through wavelet transform;
[0040] (2) Taking the CNN convolutional neural network as the basic framework, the feature is initially screened by a convolutional block with a convolution kernel size of 3×3. The convolutional block includes a convolutional layer and a max pooling layer. As Figure 3 shown, the convolutional layer is replaced by a depthwise separable convolution (DSC). The DSC layer replaces the traditional convolutional layer, reducing the number of model parameters and improving the model fault recognition accuracy;
[0041] (3) Secondly, a bearing fault diagnosis model with a hybrid attention mechanism is built, enabling the preprocessed image data to extract local features of the fault image through the window attention mechanism based on Swin Transformer, and at the same time, the fault image data extracts its global features through a convolutional neural network based on the global attention mechanism (GAMAttention).
[0042] The detailed model operation content and specific parameters are as follows:
[0043] (1) After convolution processing by the depthwise separable convolutional layer, a first feature map is generated. Different from the standard convolution, the depthwise separable convolution divides the convolution process into two steps: depth convolution and point convolution, which can not only effectively extract the features of image data but also reduce the number of convolution parameters and the amount of computation. The convolution calculation processes of various types are by Figure 1As shown: In standard convolution, the number of channels of the convolution kernel is the same as that of the input data, and the number of convolution kernels is equal to the number of output channels. It is necessary to multiply and accumulate the data with the same two-dimensional coordinates in each channel to obtain the result of one output channel. In depthwise separable convolution, the input data is first subjected to depthwise convolution. Each convolution kernel has only one channel, and the number of convolution kernels and the number of output channels are the same as the number of input channels. Each channel only needs to perform multiplication operations without accumulation and summation; then pointwise convolution is used for linear connection. Pointwise convolution is equivalent to a standard convolution with a fixed convolution kernel size of 1×1.
[0044] The specific formula for the depthwise separable convolution layer is as follows:
[0045] Depthwise separable convolution can greatly reduce the number of model parameters and the amount of computation. Assume that the size of the input data is C in ×K×K, the convolution kernel size of standard convolution is C in ×K×K, and when the stride is 1, the number of parameters P SC and the amount of computation C SC are:
[0046] P SC = C in ×K×K×C out
[0047] C SC = C in ×N×N×K×K×C out
[0048] Among them, in A*B*C of the data and the convolution kernel size, A, B, and C represent the length, width, and height of the three dimensions of the data in sequence, and the above parameters represent the corresponding numerical values.
[0049] The number of parameters P DSC and the amount of computation C DSC required for depthwise separable convolution to generate the same-sized output are:
[0050] P DSC = C in ×K×K + C in ×C out
[0051] C DSC = C in ×N×N×K×K + C in ×N×N×C out
[0052] The ratio of the number of parameters R P and the amount of computation R C of standard convolution and depthwise separable convolution can be calculated as follows:
[0053]
[0054] As can be seen from the formula, the computational cost of depthwise separable convolution is much smaller than that of standard convolution, and the difference in computational cost between the two can be reflected in complex models.
[0055] (2) Process the first feature map through the global attention mechanism, where the global attention mechanism (Global Attention Mechanism) consists of a channel attention mechanism (CAM) and a spatial attention mechanism (SAM).
[0056] The calculation formula for channel attention is as follows:
[0057]
[0058] The calculation formula for spatial attention is as follows:
[0059]
[0060] Among them, F represents the feature to be processed, σ represents the activation function, f 7×7 represents a 7×7 convolution process, MLP represents a perceptron layer process, W represents the weight, where the subscripts 1 and 0 represent the weights of different layers. In the F feature, the superscripts c and s respectively represent the corresponding channel attention feature and spatial attention feature, and the subscripts avg and max represent the corresponding average pooling and max pooling processes respectively.
[0061] The running steps of the channel attention mechanism and the spatial attention mechanism are represented by Figure 1 unified representation.
[0062] Improve the performance of the deep neural network by reducing information reduction and amplifying global interaction representations. Given the input first feature map F 1 , in the above channel attention mechanism and spatial attention mechanism, the corresponding output channel feature map F 2 and the output feature map F containing spatial features 3 are defined as:
[0063]
[0064]
[0065] Among them, M c and M s represent the processing contents of the channel attention and spatial attention mechanisms respectively; represents an element-wise multiplication operation.
[0066] The channel attention sub-module uses a three-dimensional arrangement to preserve three-dimensional information. For the input feature map, it first performs dimensional conversion and then inputs it into a two-layer MLP (Multi-Layer Perceptron) to amplify the cross-dimensional channel-spatial dependencies, and finally outputs after Sigmoid activation processing. The channel attention sub-module is as shown in Figure 1 the following figure.
[0067] In the spatial attention sub-module, in order to focus on spatial information, two convolutional layers are used for spatial information fusion. First, the number of channels is reduced through a 7×7 convolution to reduce the computational amount, then the number of channels is increased through another 7×7 convolution to keep the number of channels before and after consistent, and finally it goes through Sigmoid processing.
[0068] (3) Swin Transformer uses a hierarchical construction method similar to that in convolutional neural networks, which helps to build tasks such as object detection and instance segmentation. The concept of Windows Multi-Head Self-Attention (W-MSA) is used in Swin Transformer to calculate self-attention in non-overlapping local windows, replacing the global self-attention in the standard Transformer. Assuming that each window contains M×M vectors of fixed dimensions (i.e., patch tokens), the computational complexities of the global MSA module and the window attention based on an h×w block image are respectively:
[0069] Ω(MSA) = 4hwC 2 + 2(hw) 2 C
[0070] Ω(WMSA) = 4hwC 2 + 2M 2 hwC
[0071] where Ω(MSA) represents the complexity of the global MSA module, C represents the number of channels, h represents the height, w represents the width, and M represents the size of each window patch. MSA has a quadratic complexity with respect to the number of patch tokens h×w (a total of h×w patch tokens are calculated h×w times globally). W-MSA has a linear complexity when M is fixed (default set to 7) (a total of h×w patch tokens are calculated M 2 times) within their respective local windows. The huge h×w is unbearable for global self-attention calculation, while the window-based self-attention (W-MSA) has good scalability.
[0072] Although the window-based self-attention module (W-MSA) reduces the computational complexity from quadratic to linear, the lack of communication and connection between windows will limit its modeling and representation ability. The concept of Shifted Windows Multi-Head Self-Attention (SW-MSA) is introduced, and the shifted window partitioning method is adopted, through which information can be transmitted between adjacent windows.
[0073]
[0074] Among them and z l represent the output features of the SW-MSA module and the MLP module respectively; W-MSA and SW-MSA represent the use of conventional window multi-head attention and shifted window multi-head attention respectively, MLP represents a multi-layer perceptron, and LN represents normalization.
[0075] (4) Then, the extracted global spatial features and local features are adaptively average pooled after fusion, enabling the model to better fuse feature representations at different levels and improve the model's performance and generalization ability; the extracted global spatial feature and local feature tensors are converted to the same dimension and then concatenated and fused, and the fused features are then adaptively average pooled so that the model can complete the classification task.
[0076] (5) Finally, the constructed model is pre-trained using the data of one working condition, and then fine-tuned using a small amount of data of another working condition. The current bearing vibration data is identified by the fine-tuned diagnostic model to achieve small-sample bearing fault diagnosis between different working conditions.
[0077] Specifically, data preprocessing: The collected one-dimensional vibration signal is converted into two-dimensional image data using wavelet transform. At the same time, according to different fault states of the bearing (normal, inner ring fault, outer ring fault, rolling element fault) and different damage diameters (0.007 inches, 0.014 inches, and 0.021 inches), a ten-class dataset is constructed according to the ratio of training set: validation set: test set = 7:2:1, and three original datasets are constructed according to three different working conditions (0HP, 1HP, 2HP) for subsequent model training and testing.
[0078] Specifically, model construction: The model is built through the torch library in PyCharm. The window attention mechanism of Swin Transformer is used to extract the local features of the image, and at the same time, the global spatial features of the image are extracted through a convolutional neural network based on the global attention mechanism. Then, the two extracted features are fused and subjected to adaptive averaging processing, and finally, the ten-class detection task is completed.
[0079] Specifically, variable operating condition migration: In complex and variable industrial or environmental recognition tasks, to improve the generalization ability and rapid adaptability of the model, the present invention adopts a carefully designed strategy that combines single-operating-condition deep learning and cross-operating-condition transfer fine-tuning technology. Specifically, first, the present invention trains three independent deep learning models for three representative single-operating-condition data sets respectively. This process aims to capture the unique feature patterns and weight parameters from each specific operating condition to ensure that the model can accurately represent the data characteristics under this operating condition. Subsequently, to address the operating condition changes and small sample challenges commonly encountered in practical applications, a transfer learning mechanism is introduced. As Figure 4 shown, this mechanism selects 10% of the training set in the target operating condition data set to fine-tune the model previously trained under a single operating condition. The core of this step is to use a small amount of target operating condition data to finely adjust the model parameters so that it can quickly adapt to the data distribution and characteristics of the new operating condition, thereby significantly improving the recognition and classification performance of the model under the new operating condition without significantly increasing data requirements and training costs. This strategy makes full use of the advantages of transfer learning, effectively solves the problems of poor model adaptability and small sample learning under variable operating conditions, and at the same time realizes a smooth transition from single-operating-condition learning to cross-operating-condition transfer.
[0080] Combined with the relevant drawings and relevant experimental data, the above technical solutions are described in detail:
[0081] Figure 5 It shows the parameter results of each model after testing under three operating conditions in the CWRU public data set. Among them, the fault test accuracies of the proposed model under three different motor loads of 0HP, 1HP, and 2HP are 100%, 100%, and 100%. By comparing with the diagnostic results of other models, it can be seen that the proposed model has better effects than other models.
[0082] As Figure 6 shown is the schematic diagram of the parameter results of each model after cross-operating-condition transfer of each model. Among them, 0HP-1HP and 0HP-2HP respectively indicate that the model is pre-trained under the 0HP operating condition, and then the trained model parameters are fine-tuned using the data under 1HP and 2HP. 1HP-0HP and 1HP-2HP respectively indicate that the model is pre-trained under the 1HP operating condition, and then the trained model parameters are fine-tuned using the data under 0HP and 2HP to complete the classification task under the target operating condition. Finally, the fault test accuracies under six tasks of 0HP-1HP, 0HP-2HP, 1HP-0HP, 1HP-2HP, 2HP-0HP, and 2HP-1HP are 100%, 100%, 100%, 99.58%, and 100% respectively. By comparing with the diagnostic results of other models, it can be seen that the proposed model has better effects than other models in the transfer task.
[0083] Figure 7 This is the experimental bench for bearing fault detection in the embodiments of the present invention. In this experiment, the data signals collected by the fault diagnosis test bench are used to test the diagnosis model. Taking the SKF-NU-1006 cylindrical roller bearing produced by SKF as an example. The fault sampling time is about 30 s, the rotational speed is 2700 r / min, and it is divided into three different load conditions A, B, and C (0 MPa, 0.25 MPa, and 0.5 MPa). The faults are classified into 7 types including normal, outer ring, inner ring, and rolling element under 1.4 mm and 1.8 mm pitting damages. As Figure 8 shown, the fault test accuracies of the three types A, B, and C are 98.86%, 100%, and 100%. As Figure 9 shown, it is a schematic diagram of the results of each parameter of the model test after cross-condition migration. Among them, A-B and A-C respectively indicate that the model is pre-trained under condition A, and then the trained model parameters are fine-tuned using the data under B and C. B-A and B-C respectively indicate that the model is pre-trained under condition B, and then the trained model parameters are fine-tuned using the data under A and C to complete the classification task under the target condition. Finally, the fault test accuracies of the six tasks of A-B, A-C, B-A, B-C, C-A, and C-B are 100%, 100%, 98.96%, 100%, 98.96%, and 100% respectively.
[0084] The schematic diagram of the diagnostic model structure of the present invention is as Figure 10 shown. By integrating continuous wavelet transform, depthwise separable convolution, Swin window attention mechanism, and GAM global attention mechanism, a DSCNN-SwinGAM network model is constructed. This model improves the convolutional layer of the CNN model, combines the window-global attention mechanism, and introduces the multi-scale feature extraction of Swin and the global attention of the GAM mechanism, effectively capturing time-frequency and local spatial features, significantly enhancing the recognition of fault signals, reducing the number of parameters required for model calculation, greatly improving the fault feature discrimination ability of the model, and accelerating its convergence speed.
[0085] The overall process of the diagnostic method of the present invention is as Figure 11 shown. The overall diagnostic process will be introduced in the following four steps.
[0086] In the first step, the model first receives the two-dimensional time-frequency diagram processed by continuous wavelet transform and uses it as the input of the feature extraction module. With the CNN convolutional neural network as the basic framework, the convolutional block with a convolution kernel size of 3×3 is used for initial feature screening. The convolutional block includes a convolutional layer and a max-pooling layer. The convolutional layer is replaced by depthwise separable convolution (DSC). The DSC layer replaces the traditional convolutional layer, reducing the number of model parameters and improving the fault recognition accuracy of the model.
[0087] In the second step, the image feature data (labeled as capital C in the figure) preprocessed in the first step is used to extract the local features of the fault image through the window attention mechanism based on Swin, and at the same time, the global features are extracted through a convolutional neural network based on the global attention mechanism (GAM). The two parts jointly complete the refinement and classification of the fault features. After being processed by depthwise separable convolution, the first feature map is generated through GAM, and then passes through the channel attention sub-module (CAM) and the spatial attention sub-module (SAM) respectively. The channel attention sub-module uses three-dimensional permutation to retain three-dimensional information. For the input feature map, first perform dimensional transformation, then input it into a two-layer MLP (multi-layer perceptron) to amplify the cross-dimensional channel-spatial dependence relationship, and finally output after Sigmoid activation processing. In the spatial attention sub-module, in order to focus on spatial information, two convolutional layers are used for spatial information fusion. First, reduce the number of channels through a 7×7 convolution to reduce the computational amount, then increase the number of channels through another 7×7 convolution to keep the number of channels before and after the same, and finally perform Sigmoid processing. At the same time, after being processed by depthwise separable convolution, the first feature map is generated through Swin. The sliding window attention mechanism of Swin extracts local features from the fault image. By calculating the attention within each sliding window, this mechanism effectively captures the relationship between features. In addition, the sliding window method allows the model to process each region of the image and extract local features. The overlap between adjacent windows ensures the continuity of local information, enabling the model to more accurately capture the feature dependencies in the time-frequency diagram and extract key global information from the signal.
[0088] In the third step, the global spatial features and local feature tensors extracted through GAM and Swin are converted to the same dimension and then concatenated and fused. Then, the fused features are subjected to adaptive average pooling to enable the model to better fuse feature representations at different levels, improve the model performance and generalization ability so that the model can complete the classification task.
[0089] In the fourth step, the constructed model is pre-trained using the data of one working condition, and then fine-tuned using a small amount of data of another working condition. The current bearing vibration data is identified through the fine-tuned diagnostic model to achieve small-sample bearing fault diagnosis between different working conditions.
[0090] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A variable-condition bearing fault diagnosis method based on a window global mixed attention mechanism and an improved convolutional neural network is characterized in that: include: Acquire bearing signals under different working conditions and target working conditions, preprocess the bearing signals, and generate signal image data under different working conditions and target working conditions; Constructing a diagnostic model, wherein a window attention mechanism and a global attention mechanism are provided in the diagnostic model for respectively extracting local features and global features corresponding to the signal image data for subsequent recognition; Pre-training the diagnostic model using signal image data under different working conditions to obtain a pre-trained model, and training the pre-trained model using signal image data under target working conditions to obtain an optimal diagnostic model; The bearing signal to be tested under the target working condition is obtained, and the bearing signal to be tested is diagnosed and identified through the optimal diagnosis model to obtain the fault diagnosis result of the bearing signal to be tested.
2. The method according to claim 1, characterized in that The preprocessing method adopts wavelet transform.
3. The method according to claim 1, characterized in that The diagnostic model includes a depth-separable convolutional layer module, an attention mechanism module, a fusion module and a recognition module connected in sequence, wherein the attention mechanism module includes a parallel window attention mechanism and a global attention mechanism, which are respectively connected in parallel with the convolution module and the fusion module.
4. The method according to claim 1, characterized in that: The window attention mechanism includes a window structure at different time steps, wherein the window structure includes a first LN layer, a W-MSA layer, a second LN layer, and an MLP layer connected in sequence, wherein the sum of the input data of the first LN layer and the output data of the W-MSA layer is used as the input data of the second LN layer, and the sum of the input data of the second LN layer and the output data of the MLP layer is used as the input data of the window structure of the next time step.
5. According to the method described in claim 1, the global attention mechanism includes a channel attention mechanism and a spatial attention mechanism connected in sequence, wherein the sum of the input data and the output data of the channel attention mechanism serves as the input data of the spatial attention mechanism, and the sum of the input data and the output data of the spatial attention mechanism serves as the output data of the global attention mechanism.
6. The method according to claim 3, characterized in that The recognition module includes a fully connected layer and an output layer connected in sequence, wherein the output layer adopts a softmax function.
Citation Information
Cited By
Zero-sequence current and vibration signal graph domain fused transmission chain composite fault diagnosis method
CN120873559A