A Binary Tree Filter Transformer Model and Its Application in Bearing Fault Diagnosis
By introducing a binary tree filter group into the Transformer model, the characteristic information in the vibration signal of the rolling bearing is extracted, and the problem of low fault diagnosis efficiency in the prior art is solved, achieving higher diagnostic accuracy and fault information characterization capabilities.
Patent Information
- Application Number
- CN202210156171.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-21
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-02-21
AI Technical Summary
The prior art is difficult to effectively extract key feature information in the vibration signal of rolling bearings, resulting in low bearing fault diagnosis efficiency.
The binary tree filter Transformer model is used to extract vibration signals through the binary tree filter group in the word segmenter, and the Transformer block in the encoder performs feature extraction and classification to achieve accurate diagnosis of bearing failures.
This model can effectively extract multiple frequency band information and sort it according to frequency, improve the characterization ability of fault information, achieve higher diagnostic accuracy, and surpass the performance of traditional convolutional neural networks and recurrent neural networks.
Smart Images

Figure CN114662529B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning technology and mechanical fault diagnosis technology, and specifically relates to a binary tree filter Transformer model and a method for using the model to diagnose bearing faults. Background Art
[0002] As one of the most widely used parts of rotating machinery, rolling bearings often operate under high speed and high load conditions. Once a failure occurs, it will bring unpredictable disasters. Timely fault diagnosis of bearings and formulation of appropriate maintenance strategies to effectively reduce ineffective expenses and improve maintenance efficiency are the key points of bearing maintenance. Therefore, how to effectively diagnose rolling bearing faults is of great practical significance.
[0003] The type and severity of rolling bearing faults are reflected in their vibration signals; therefore, some research has used vibration signals to perform fault diagnosis to identify the fault status of bearings (mainly rollers, inner rings and outer rings). However, with the increasing complexity of mechanical equipment and the non-stationary and nonlinear characteristics of vibration signals, it has become more difficult to extract representative key feature information from vibration signals.
[0004] With the development of machine deep learning technology, using deep learning models to extract feature information has become a more effective means. However, in the field of fault diagnosis based on vibration signals, there is currently a lack of a more effective deep learning model. Summary of the invention
[0005] In view of the deficiencies in the prior art, the present invention provides a binary tree filter Transformer model and a method for using the model for bearing fault diagnosis, so as to solve the problem of accurate diagnosis of rolling bearing faults.
[0006] The present invention achieves the above technical objectives through the following technical means.
[0007] A binary tree filter Transformer model, characterized in that it includes a word segmenter, an encoder and a classifier, wherein the vibration signal is extracted by the word segmenter to obtain a sequence z 1 , then the sequence z 1 The output result is obtained by processing the encoder and classifier in sequence;
[0008] The word segmenter is provided with a binary tree filter bank:
[0009]
[0010] Where h(n) is a standard low-pass filter with a cutoff frequency of ε≥0; j is an imaginary unit, h L The frequency band of (n) is [0, 1 / 4F s ],h H The frequency band of (n) is [1 / 4F s , 1 / 2F s ],F s is the sampling frequency;
[0011] For the input vibration signal x(n), the word segmenter includes the following processing steps:
[0012] S1, using the binary tree filter bank to decompose the signal x(n) into k layers, where k is the integer of log2(n);
[0013] S2, take the real part of all sub-signals in the kth layer, and then calculate the RMS
[0014] S3, according to Get the fragment information S and position information K, and find the sequence z 1 :
[0015]
[0016] E A = embed(S,W S )
[0017] E K = embed(K,W P )
[0018] in d model is the embedding dimension, Y=embed(X,W) means expanding the dimension of Y according to X and W.
[0019] Furthermore, the encoder includes N (or N layers) Transformer blocks, each of which contains two submodules, namely a multi-head self-attention submodule and a forward network submodule. For the sequence z, there are:
[0020]
[0021] Among them, the sequence z l It is represented as the input of the l-th layer Transformer block, and the corresponding input of the first layer is the sequence z 1 , Attention(·) represents the multi-head self-attention submodule, z l ′ is the output of the multi-head self-attention submodule and also the input of FFn(·), where FFn(·) represents the feed-forward network submodule. l+1As the output of the l-th layer forward network submodule, it is also the input of the l+1-th layer multi-head self-attention module. Layernorm takes the residual and normalizes it.
[0022] Furthermore, the multi-head self-attention submodule is specifically composed of the query matrix Q s = zW j q , the bond matrix K s = zW j k , value matrix Composition, the corresponding formula is:
[0023]
[0024] Where h is the number of attention heads, j∈[1,h], W j q , W j k and They represent: the jth linear mapping applied to the sequence z to obtain different versions of the query matrix, key matrix, and value matrix, respectively. d k =d model / h, softmax(·) represents the generalization of logistic regression, concat(·) is to connect the matrices, head j is the jth A(Q,K,V), W s For j W j q , W j k , series connection.
[0025] Furthermore, the forward network submodule is represented as:
[0026] FFn(x)=GeLU(0,w 1 x+b 1 ) 2 +b 2
[0027] Where GeLU(·) is the activation function, w 1 and w 2 Represents the weight b that undergoes two linear changes 1 and b 2 Indicates the bias that changes linearly twice.
[0028] Furthermore, the expression of the classifier is as follows:
[0029]
[0030] Where Classifier(·) represents the classifier function, w c1 and w c2 is the weight, b c1 and b c2 is the bias, d cate represents the type in the sample, d ff is the hidden layer dimension.
[0031] A bearing fault diagnosis method based on the above binary tree filter Transformer model: a data set is prepared to train the binary tree filter Transformer model, and then the trained model is used to perform fault diagnosis based on the rolling bearing vibration signal.
[0032] The beneficial effects of the present invention are:
[0033] The present invention provides a binary tree filter Transformer model and its bearing fault diagnosis application, wherein the Transformer model used is a deep learning model proposed in recent years for solving the problem of natural language recognition and processing, and the present invention makes corresponding modifications and adjustments to the word segmenter part of the Transformer model, that is, by introducing a binary tree filter, the improved Transformer model can extract vibration signals (the corresponding original Transformer model is only suitable for natural language extraction); and then using the binary tree filter Transformer model of the present invention, for the input vibration signal, multiple frequency bands can be effectively extracted and sorted from small to large according to frequency, so as to effectively characterize the information in different frequency bands, and finally achieve the beneficial effect of better measuring fault information. In addition, the Transformer model can give good weights to different frequency bands, so as to obtain a more accurate recognition rate, which is impossible for traditional convolutional neural networks and recurrent neural networks to achieve. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 This is a structural diagram of the binary tree filter Transformer model of the present invention;
[0035] Figure 2 This is a flow chart of bearing fault diagnosis of the present invention;
[0036] Figure 3 is the error of each round during model training;
[0037] Figure 4 The accuracy of each round during model training;
[0038] Figure 5 is the worst classification result;
[0039] Figure 6 is the optimal classification result;
[0040] Figure 7(a) is a characteristic trend diagram of each bearing fault type 1 extracted;
[0041] Figure 7(b) is a characteristic trend diagram of each bearing fault type 2 extracted;
[0042] Figure 7(c) is a characteristic trend diagram of each bearing fault type 3 extracted;
[0043] Figure 7(d) is a characteristic trend diagram of each bearing fault type 4 extracted;
[0044] Figure 7(e) is a characteristic trend diagram of each bearing fault type 5 extracted;
[0045] Figure 7(f) is a characteristic trend diagram of each bearing fault type 6 extracted;
[0046] Figure 7(g) is a characteristic trend diagram of each bearing fault type 7 extracted;
[0047] Figure 7(h) is a characteristic trend diagram of each bearing fault type 8 extracted;
[0048] FIG7(i) is a characteristic trend diagram of each bearing fault type 9 extracted;
[0049] FIG7( j ) is a characteristic trend diagram of each bearing fault type 10 extracted. DETAILED DESCRIPTION
[0050] Embodiments of the present invention are described in detail below, examples of the illustrated embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0051] 1. Binary Tree Filter Transformer Model
[0052] like Figure 1 The figure shows the Transformer model of the binary tree filter of the present invention, which mainly includes three parts: a word segmenter, an encoder and a classifier, wherein the word segmenter is used to extract the vibration signal of the rolling bearing. Note: The Transformer model is an existing model, which is literally translated into "Transformer model" in Chinese, but there is no unified or commonly accepted Chinese translation in this field, so this article directly uses its English name for description, that is, "Transformer model", "Transformer block", etc.
[0053] 1. Tokenizer
[0054] The word segmenter is provided with a binary tree filter bank:
[0055]
[0056] Where h(n) is a standard low-pass filter with a cutoff frequency of ε≥0; j is the imaginary unit (equivalent to i); h L (n) is a quasi-analytic low-pass filter, whose frequency band range is [0, 1 / 4F s ],h H (n) is a quasi-analytic high-pass filter, and its frequency band is [1 / 4F s , 1 / 2F s ], F s is the sampling frequency.
[0057] For the input signal x(n), the word segmenter includes the following processing steps:
[0058] 1) Repeated decomposition is performed using the above binary tree filter bank, and finally k layers are decomposed, where k is the integer of log2(n). Suppose the i-th sub-signal of the k-th layer is i=0,1,2…2 k-1 , The center frequency f ic =(i+2 -1 )*2 -k-1 *F s , The bandwidth Δf k =2 -k-1 *F s , The two next-level sub-signals adjacent to it and There exists between:
[0059]
[0060] 2) First take the real part of all sub-signals of the kth layer, that is, to Take the real part; then calculate the RMS of all sub-signals in layer k
[0061] 3) According to Get the segment information S and position information K, where the segment information S is used to distinguish the segments in the same sentence in the original Transformer model. Since this situation is not involved in the present invention, for the convenience of calculation, The location information K is used to record The order of the corresponding center frequencies, where K = 1, 2, ... 2 k-1 . Define the following parameters E A 、E K :
[0062]
[0063] E A = embed(S,W S )
[0064] E K = embed(K,W P )
[0065] in d model is the embedding dimension, Represents a dimension d model ×1 real matrix, and so on Represents a dimension d model ×2 real matrix, Represents a dimension d model ×2 k The real matrix, E A =embed(S,W S ) indicates that according to S and W s , expand E A The dimension of E K = embed(K,W P ) indicates that according to K and W P Expand E K (Reference J.Devlin, M.-W.Chang, K.Lee, and K.Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” arXiv:1810.04805[cs], May 2019)
[0066] Afterwards E A 、E K Adding the three together, we get the sequence z 1 :
[0067]
[0068] 2. Encoder
[0069] The sequence z obtained from the previous word segmenter 1Input to the encoder. The encoder includes N (or N layers) Transformer blocks, each of which contains two submodules, namely the multi-head self-attention submodule and the forward network submodule. For the sequence z, there is:
[0070]
[0071] Among them, the sequence z l It is represented as the input of the l-th layer Transformer block, and the corresponding input of the first layer is the sequence z 1 ; Attention(·) represents the multi-head self-attention submodule; z l ′ is the output of the multi-head self-attention submodule and also the input of FFn(·), where FFn(·) represents the feed-forward network submodule. l+1 It is the output of the l-th layer forward network submodule and the input of the l+1-th layer multi-head self-attention module; layernorm takes the residual and normalizes it.
[0072] The multi-head self-attention submodule is specifically composed of a query matrix Q s = zW j q , key matrix K s = zW j k Sum Matrix Composition, the corresponding formula is as follows:
[0073]
[0074] Where h is the number of attention heads, j∈[1,h], W j q , W j k and They represent: the jth linear mapping applied to the sequence z to obtain different versions of the query matrix, key matrix, and value matrix, respectively. d k =d model / h, softmax(·) represents the generalization of logistic regression, concat(·) is to connect the matrices, head j is the jth A(Q,K,V), W s For j W j q , W j k , series connection.
[0075] The forward network submodule is represented as:
[0076] FFn(x)=GeLU(0,w 1 x+b 1 ) 2 +b 2
[0077] Where GeLU(·) is the activation function, w 1 and w 2 Represents the weight b that undergoes two linear changes 1 and b 2 Indicates the bias that changes linearly twice.
[0078] 3. Classifier
[0079] The sequence z obtained in the previous encoder N+1 Input to the classifier to get the final output The expression of the classifier is as follows:
[0080]
[0081] Where Classifier(·) represents the classifier function, w c1 and w c2 is the weight, b c1 and b c2 is the bias, d cate represents the type in the sample, d ff is the hidden layer dimension.
[0082] 2. Bearing fault diagnosis test
[0083] like Figure 2 The flowchart of the bearing fault diagnosis test based on the binary tree filter Transformer model of the present invention is shown, which specifically includes the following steps:
[0084] S1, collecting rolling bearing vibration signals, and dividing the vibration signals into training samples and test samples;
[0085] S2, input the training samples into the binary tree filter Transformer model to train the model;
[0086] S3, input the test sample into the trained binary tree filter Transformer model to perform fault diagnosis test.
[0087] The specific test examples are as follows:
[0088] Taking SKF-6205 rolling bearing as the experimental object, simulation tests of various fault types were carried out.
[0089] Step 1: Use a vibration sensor to measure the vibration data of the bearing, and collect the data in a PC through a data acquisition card for processing.
[0090] The specific data set is described as follows: There are 10 fault types: normal (1), slight inner ring fault (2), moderate inner ring fault (3), severe inner ring fault (4), slight outer ring fault (5), moderate outer ring fault (6), severe outer ring fault (7), slight roller fault (8), moderate roller fault (9), severe roller fault (10). The numbers in brackets represent labels, and the corresponding final model output results are That is, it is a value from the above labels 1 to 10. The vibration data with a sampling frequency of 10kHz is divided, and there are 300 groups of data for each fault type, with a length of 20,000. In addition, 80% of all the data is used as the training data set and 20% as the test data set. The data sets for each training and test are randomly divided to ensure the performance of the model.
[0091] Step 2: Construct a binary tree filter Transformer model, where the model parameters are detailed in Table 1.
[0092] Table 1: Network structure parameter selection
[0093] Parameter name Numeric Input size [32,1] Batch size 32 Maximum number of training rounds 30 Learning Rate 5e-6 Optimizer Adam Label smoothness 0.1 Number of encoders N 16 <![CDATA[Embedding dimension d model > 64 <![CDATA[Hidden layer dimension d ff > 768 Number of attention heads h 12 Dropout probability 0.1 Positional encoding One-dimensional learnable
[0094] The training set is input into the constructed binary tree filter Transformer model for training. The network training is based on the gradient descent algorithm and the error back propagation algorithm, and the Adam optimizer is used. The error and accuracy of each round are as follows: Figure 3 and Figure 4 shown.
[0095] Step 3: Input the test sample into the trained binary tree filter Transformer model and perform fault diagnosis test. Repeat 20 times. The worst and best classification results are as follows: Figure 5 and Figure 6 shown.
[0096] In order to measure the superiority of the present invention, two existing models were used for comparative testing, wherein Comparison 1: Long Short-Term Memory Network (LSTM); Comparison 2: Gated Recurrent Unit Network (GRU). The comparative models were implemented 20 times respectively to compare the final test results. The results are shown in Table 2, which shows that the accuracy of the present invention is the highest.
[0097] Table 2: Comparison of the method of the present invention with other methods
[0098] method Average accuracy The present invention 99.92% Comparison 1 98.33% Comparison 2 96.84%
[0099] Finally, in the bearing testing method of the present invention, the characteristic trends of the first to tenth bearing fault types extracted are respectively as follows: Figures 7(a) to 7(j) As shown, it can be seen intuitively from the figure that the features extracted by the present invention are simple and easy to characterize, which is beneficial to the subsequent analysis and processing by technicians.
[0100] In the description of the present invention, it should be understood that the terms "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside" and "outside" etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.
[0101] The present invention is not limited to the above-mentioned embodiments. Any obvious improvement, substitution or deformation that can be made by those skilled in the art without departing from the essential content of the present invention belongs to the protection scope of the present invention.
Claims
1. A binary tree filter Transformer model, Features: It includes a word segmenter, an encoder and a classifier, wherein the vibration signal is extracted by the word segmenter to obtain a sequence z 1 , then the sequence z 1 The output result is obtained by processing the encoder and classifier in sequence; The word segmenter is provided with a binary tree filter bank: Where h(n) is a standard low-pass filter, n is the signal time, and the cutoff frequency ε≥0, ε is the parameter in the cutoff frequency of the standard low-pass filter; j is the imaginary unit, h L The frequency band of (n) is [0, 1 / 4F s ],h H The frequency band of (n) is [1 / 4F s , 1 / 2F s ], F s is the sampling frequency; For the input vibration signal x(n), the word segmenter includes the following processing steps: S1, using the binary tree filter bank to decompose the signal x(n) into k layers, where k is the integer of log2(n); S2, take the real part of all sub-signals in the kth layer, and then calculate the RMS S3, according to Get the fragment information S and position information K, and find the sequence z 1 : E A =embed(S,W S ) E K =embed(K,W P ) in Represents a dimension d model ×1 real matrix, Represents a dimension d model ×2 real matrix, Represents a dimension d model ×2 k The real matrix of d model is the embedding dimension, E A =embed(S,W S ) indicates that according to S and W S Expanded dimension, E K = embed(K,W P ) indicates that according to K and W P Expanded dimension.
2. The binary tree filter Transformer model according to claim 1, Features: The encoder includes N Transformer blocks, each of which contains two submodules, namely a multi-head self-attention submodule and a forward network submodule. For sequence z, we have: Among them, the sequence z l It is represented as the input of the l-th layer Transformer block, and the corresponding input of the first layer is the sequence z 1 , Attention(·) represents the multi-head self-attention submodule, z′ l is the output of the multi-head self-attention submodule and also the input of FFn(·), where FFn(·) represents the feed-forward network submodule. l+1 As the output of the l-th layer forward network submodule, it is also the input of the l+1-th layer multi-head self-attention module. Layernorm takes the residual and normalizes it.
3. The binary tree filter Transformer model according to claim 2, Features: The multi-head self-attention submodule is specifically composed of a query matrix Q s =ZW j q , the bond matrix K s = zW j k , value matrix V s = zW j v Composition, the corresponding formula is: Where h is the number of attention heads, j∈[1,h], W j q , W j k and W j v They represent: the jth linear mapping applied to the sequence z to obtain different versions of the query matrix, key matrix, and value matrix, respectively. d k =d model / h, softmax(·) represents the generalization of logistic regression, concat(·) is to connect the matrices, head j is the jth A(Q,K,V), W s For j W j q , W j k , W j v series connection.
4. The binary tree filter Transformer model according to claim 2, Features: The forward network submodule is represented as: FFn(x)=GeLU(0,w 1 x+b 1 ) 2 +b 2 Where GeLU(·) is the activation function, w 1 and w 2 Represents the weight b that undergoes two linear changes 1 and b 2 Indicates the bias that changes linearly twice.
5. The binary tree filter Transformer model according to claim 2, Features: The expression of the classifier is as follows: Where Classifier(·) represents the classifier function, w c1 and w c2 is the weight, b c1 and b c2 is the bias, d cate represents the type in the sample, d ff is the hidden layer dimension.
6. A bearing fault diagnosis method based on the binary tree filter Transformer model according to any one of claims 1 to 5, Features: A data set is created to train the binary tree filter Transformer model, and then the trained model is used to perform fault diagnosis based on the vibration signal of the rolling bearing.
Citation Information
Patent Citations
Fractionally spaced decision feedback Rayleigh Renyi entropy wavelet blind equalization method
CN102164106A
Method and apparatus for filtering streaming data
CN103081430A