Multi-layer Feature Fusion Transformer Neural Network and Method for Detecting Electric Motors
By adopting multi-layer feature fusion technology in the Transformer neural network, the MFCC feature matrix with the number of cepspectral coefficients of different Mel frequencies and the feature fusion between different layers is solved, the problem of low classification accuracy when detecting motor sound time domain signals is achieved, and higher classification accuracy is achieved.
Patent Information
- Application Number
- CN202510422442.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-07
AI Technical Summary
When the Transformer neural network detects motor sound time domain signals, the classification accuracy is low, resulting in poor detection results.
A multi-layer feature fusion Transformer neural network is used to fuse the MFCC feature matrix of the number of cepspectral coefficients of different Mel frequencies through the feature extraction module and the transformation fusion module, and feature fusion is performed between different layers of Transformer.
The classification accuracy of detecting motor sound time domain signals is improved, and the model's expression ability between different layers of features is enhanced.
Smart Images

Figure CN119940417B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of motor testing, and particularly to a multi-layer feature fusion Transformer neural network and method for detecting motors. Background Art
[0002] Common machine learning models for motor sound classification include RNN, CNN, and Transformer.
[0003] Compared with traditional neural networks such as RNN and CNN, Transformer has significant advantages in processing sequence data such as sound and text, mainly in the following aspects: 1. The self-attention mechanism Self-Attention enables it to focus on all positions of the entire sequence, overcoming the disadvantages of RNN and CNN in long-distance dependencies. 2. Strong parallel computing ability and faster training speed. 3. Efficient inference and lower computational resource requirements.
[0004] Due to these advantages, Transformer has been widely used in speech recognition, sound classification, and large language models for dialogue such as ChatGPT and DeepSeek.
[0005] When using Transformer for motor sound classification, relying only on the output of the last layer may lead to the loss of shallow local information, thereby reducing the classification accuracy.
[0006] Therefore, when using a Transformer neural network to detect the time-domain signal of motor sound, the low classification accuracy of the Transformer neural network has become a technical problem to be solved urgently. Summary of the Invention
[0007] The present invention provides a multi-layer feature fusion Transformer neural network and method for detecting motors, which solves the technical problem of low classification accuracy of the Transformer neural network in detecting the time-domain signal of motor sound.
[0008] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:
[0009] A multi-layer feature fusion Transformer neural network for detecting motors, including a feature extraction module, a conversion and fusion module, and a fusion classification module. The feature extraction module includes a first MFCC module and a second MFCC module, and the number of mel-frequency cepstral coefficients of the second MFCC module is greater than that of the first MFCC module;
[0010] The conversion and fusion module includes a first conversion and fusion module and a second conversion and fusion module. The first conversion and fusion module includes a first normalization module, a first fully connected layer, a first encoder, a first mean module, and a first splicing module connected in sequence. The first MFCC module is connected to the first normalization module. The second conversion and fusion module includes a second normalization module, a second fully connected layer, a second encoder, a second mean module, and a second splicing module connected in sequence. The second MFCC module is connected to the second normalization module. Both the first encoder and the second encoder are encoders containing multiple encoding layers, and both the first mean module and the second mean module are mean modules containing multiple mean layers. The number of encoding layers in each encoder is the same as the number of mean layers in the mean module. All the encoding layers in an encoder are connected in sequence, each encoding layer in the encoder is connected to a corresponding mean layer in the adjacent mean module, and all the mean layers in the mean module are connected to the adjacent splicing module.
[0011] Both the first splicing module and the second splicing module are connected to the fusion classification module.
[0012] A further technical solution is that the number of encoding layers in the encoder and the number of mean layers in the mean module are both represented by n. The n-layer encoder includes a first encoding layer to an nth encoding layer, and the n-layer mean module includes a first mean layer to an nth mean layer.
[0013] The first fully connected layer, the first encoding layer to the nth encoding layer of the first encoder are connected in sequence. One encoding layer of the first encoder is connected to one mean layer of the first mean module. The first encoding layer of the first encoder is connected to the first mean layer of the first mean module, and the nth encoding layer of the first encoder is connected to the nth mean layer of the first mean module. All the mean layers of the first mean module are connected to the first splicing module.
[0014] The second fully connected layer, the first encoding layer to the nth encoding layer of the second encoder are connected in sequence. One encoding layer of the second encoder is connected to one mean layer of the second mean module. The first encoding layer of the second encoder is connected to the first mean layer of the second mean module, and the nth encoding layer of the second encoder is connected to the nth mean layer of the second mean module. All the mean layers of the second mean module are connected to the second splicing module.
[0015] A further technical solution is that n ranges from 4 to 6.
[0016] A further technical solution lies in that: the fusion classification module includes a third splicing module, a third fully connected layer, a conversion module, and a binary classification fully connected layer that are connected in sequence. The first splicing module and the second splicing module are both connected to the third splicing module. The conversion module includes a third encoder and a third mean module. The third fully connected layer, the third encoder, the third mean module, and the binary classification fully connected layer are connected in sequence.
[0017] A further technical solution lies in that: the third encoder includes a first coding layer and a second coding layer. The third fully connected layer, the first coding layer of the third encoder, the second coding layer of the third encoder, the third mean module, and the binary classification fully connected layer are connected in sequence.
[0018] A further technical solution lies in that: the number of Mel-frequency cepstral coefficients of the second MFCC module is twice that of the first MFCC module; the number of Mel-frequency cepstral coefficients of the first MFCC module ranges from 32 to 64, and the number of Mel-frequency cepstral coefficients of the second MFCC module ranges from 64 to 128.
[0019] A method for detecting a motor, based on the above-mentioned multi-layer feature fusion Transformer neural network for detecting a motor, includes the following steps
[0020] Step S1: Obtain the time-domain signal of the motor sound, extract the first spectral feature matrix from the time-domain signal of the motor sound through the first MFCC module, and extract the second spectral feature matrix from the time-domain signal of the motor sound through the second MFCC module;
[0021] Step S2: Input the first spectral feature matrix into the first normalization module, and the first splicing module outputs the first spliced feature matrix; input the second spectral feature matrix into the second normalization module, and the second splicing module outputs the second spliced feature matrix;
[0022] Step S3: Input the first spliced feature matrix and the second spliced feature matrix into the fusion classification module to obtain a binary classification result, and the binary classification result is normal or abnormal.
[0023] A further technical solution lies in that: in the step S2,
[0024] After inputting the first spectral feature matrix into the first normalization module, the following steps are further included: the first spectral feature matrix is mean-normalized by the first normalization module to obtain the normalized first normalization matrix; the first normalization matrix passes through the first fully-connected layer to obtain the first fully-connected matrix; the first fully-connected matrix is input into each encoding layer of the first encoder in the first conversion and fusion module, and one encoding layer obtains one encoded layer feature matrix; each encoded layer feature matrix of the first encoder is input into a corresponding mean layer of the first mean module, and each mean layer of the first mean module obtains an averaged feature vector; all the averaged feature vectors of the first mean module are input into the first splicing module for splicing to obtain the first spliced feature matrix.
[0025] After inputting the second spectral feature matrix into the second normalization module, the following steps are further included: the second spectral feature matrix is mean-normalized by the second normalization module to obtain the normalized second normalization matrix; the second normalization matrix passes through the second fully-connected layer to obtain the second fully-connected matrix; the second fully-connected matrix is input into each encoding layer of the second encoder in the second conversion and fusion module, and one encoding layer obtains one encoded layer feature matrix; each encoded layer feature matrix of the second encoder is input into a corresponding mean layer of the second mean module, and each mean layer of the second mean module obtains an averaged feature vector; all the averaged feature vectors of the second mean module are input into the second splicing module for splicing to obtain the second spliced feature matrix.
[0026] A further technical solution lies in: in the step S1, the shape of the first spectral feature matrix is (n_frame, n_mfcc1), and the shape of the second spectral feature matrix is (n_frame, n_mfcc2), where n_frame is the number of frames of the Fourier transform in MFCC, n_mfcc1 is the number of Mel frequency cepstral coefficients of the first MFCC module, and n_mfcc2 is the number of Mel frequency cepstral coefficients of the second MFCC module;
[0027] In the step S2, the shape of the first normalization matrix is (n_frame, n_mfcc1), the shape of the first fully-connected matrix is (n_frame, n_embedding1), the shape of the encoded layer feature matrix of the first encoder is (n_frame, n_embedding1), the shape of the averaged feature vector of the first mean module is (1, n_embedding1), and the shape of the first concatenated feature matrix is (n, n_embedding1); the shape of the second normalization matrix is (n_frame, n_mfcc2), the shape of the second fully-connected matrix is (n_frame, n_embedding2), the shape of the encoded layer feature matrix of the second encoder is (n_frame, n_embedding2), the shape of the averaged feature vector of the second mean module is (1, n_embedding2), and the shape of the second concatenated feature matrix is (n, n_embedding2); n_embedding1 is the number of features embedded by the first fully-connected layer, and n_embedding2 is the number of features embedded by the second fully-connected layer.
[0028] A further technical solution lies in: in the step S3, the steps of inputting the first concatenated feature matrix and the second concatenated feature matrix into the fusion classification module to obtain a binary classification result include the following steps.
[0029] Input the first concatenated feature matrix and the second concatenated feature matrix into the third concatenation module, and the first concatenated feature matrix and the second concatenated feature matrix are concatenated by the third concatenation module to obtain a feature fusion matrix with a shape of (n, n_embedding1 + n_embedding2); input the feature fusion matrix with a shape of (n, n_embedding1 + n_embedding2) into the third fully-connected layer to obtain a third fully-connected matrix with a shape of (n, n_embedding3); input the third fully-connected matrix with a shape of (n, n_embedding3) into the third encoder of the conversion module, and then take the average value through the third mean module to obtain a fused Transformer feature with a shape of (1, n_embedding3); input the fused Transformer feature with a shape of (1, n_embedding3) into the binary classification fully-connected layer to obtain a binary classification result; n_embedding3 is the number of features embedded by the third fully-connected layer.
[0030] The beneficial effects produced by adopting the above technical solution are as follows:
[0031] A multi - layer feature fusion Transformer neural network for detecting motors, including a feature extraction module, a conversion and fusion module, and a fusion classification module. The feature extraction module includes a first MFCC module and a second MFCC module, and the number of mel - frequency cepstral coefficients of the second MFCC module is greater than that of the first MFCC module. The conversion and fusion module includes a first conversion and fusion module and a second conversion and fusion module. The first conversion and fusion module includes a first normalization module, a first fully - connected layer, a first encoder, a first mean module, and a first splicing module connected in sequence, and the first MFCC module is connected to the first normalization module. The second conversion and fusion module includes a second normalization module, a second fully - connected layer, a second encoder, a second mean module, and a second splicing module connected in sequence, and the second MFCC module is connected to the second normalization module. Both the first splicing module and the second splicing module are connected to the fusion classification module. Through the first MFCC module, the second MFCC module, the first conversion and fusion module, the second conversion and fusion module, etc., the Transformer neural network model with multi - layer feature fusion fuses MFCC feature matrices with different numbers of mel coefficients, performs feature fusion between different layers of the Transformer, increases the expression ability of the model between different - layer features, and improves the classification accuracy of detecting the time - domain signal of the motor sound.
[0032] A method for detecting motors, based on the above - mentioned multi - layer feature fusion Transformer neural network for detecting motors, includes the following steps. Step S1: Obtain the time - domain signal of the motor sound, extract the first spectral feature matrix from the time - domain signal of the motor sound through the first MFCC module, and extract the second spectral feature matrix from the time - domain signal of the motor sound through the second MFCC module. Step S2: Input the first spectral feature matrix into the first normalization module, and the first splicing module outputs the first spliced feature matrix. Input the second spectral feature matrix into the second normalization module, and the second splicing module outputs the second spliced feature matrix. Step S3: Input the first spliced feature matrix and the second spliced feature matrix into the fusion classification module to obtain a binary classification result, and the binary classification result is normal or abnormal. Through steps S1 to S3, the classification accuracy of detecting the time - domain signal of the motor sound is improved. Description of the Drawings
[0033] Figure 1 It is the principle block diagram of Embodiment 1 of the present invention;
[0034] Figure 2 It is the data - flow diagram of Embodiment 2 of the present invention. Detailed Embodiments
[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and in no way limits the present application and its application or use. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0036] Many specific details are set forth in the following description in order to provide a thorough understanding of the present application, but the present application may also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.
[0037] Embodiment 1:
[0038] As Figure 1 shown, the present invention discloses a multi-layer feature fusion Transformer neural network for detecting motors, including a feature extraction module, a conversion fusion module, and a fusion classification module. The feature extraction module includes a first MFCC module and a second MFCC module. The number of Mel frequency cepstral coefficients n_mfcc1 of the first MFCC module ranges from 32 to 64, and the number of Mel frequency cepstral coefficients n_mfcc2 of the second MFCC module ranges from 64 to 128. The number of Mel frequency cepstral coefficients n_mfcc2 of the second MFCC module is greater than the number of Mel frequency cepstral coefficients n_mfcc1 of the first MFCC module.
[0039] The conversion fusion module includes a normalization module, a fully connected layer embedding, an encoder Encoder, a mean module, and a splicing module connected in sequence. The encoder Encoder is an n-layer encoder, and the n-layer encoder includes a first coding layer to an nth coding layer. The mean module is an n-layer mean module, and the n-layer mean module includes a first mean layer to an nth mean layer. The normalization module, the fully connected layer embedding, the first coding layer to the nth coding layer are connected in sequence. One coding layer is connected to one mean layer, the first coding layer is connected to the first mean layer, and all the mean layers are connected to the splicing module.
[0040] The value range of n is 4 to 6.
[0041] The number of the conversion fusion modules is two, namely a first conversion fusion module and a second conversion fusion module.
[0042] The first conversion and fusion module includes a first normalization module, a first fully-connected layer embedding1, a first encoder Encoder1, a first mean module, and a first splicing module connected in sequence. The second conversion and fusion module includes a second normalization module, a second fully-connected layer embedding2, a second encoder Encoder2, a second mean module, and a second splicing module connected in sequence. Both the first encoder Encoder1 and the second encoder Encoder2 are n-layer encoders. The n-layer encoder includes a first coding layer to an nth coding layer. Both the first mean module and the second mean module are n-layer mean modules. The n-layer mean module includes a first mean layer to an nth mean layer.
[0043] The first MFCC module, the first normalization module, the first fully-connected layer embedding1, and the first coding layer to the nth coding layer of the first encoder Encoder1 are connected in sequence. One coding layer of the first encoder Encoder1 is connected to one mean layer of the first mean module. The first coding layer of the first encoder Encoder1 is connected to the first mean layer of the first mean module, and so on, until the nth coding layer of the first encoder Encoder1 is connected to the nth mean layer of the first mean module. All mean layers of the first mean module are connected to the first splicing module.
[0044] The second MFCC module, the second normalization module, the second fully-connected layer embedding2, and the first coding layer to the nth coding layer of the second encoder Encoder2 are connected in sequence. One coding layer of the second encoder Encoder2 is connected to one mean layer of the second mean module. The first coding layer of the second encoder Encoder2 is connected to the first mean layer of the second mean module, and so on, until the nth coding layer of the second encoder Encoder2 is connected to the nth mean layer of the second mean module. All mean layers of the second mean module are connected to the second splicing module.
[0045] The fusion classification module includes a third splicing module, a third fully-connected layer embedding3, a conversion module, and a binary classification fully-connected layer connected in sequence. The first splicing module and the second splicing module are both connected to the third splicing module. The conversion module includes a third encoder and a third mean module. The third encoder includes a first coding layer and a second coding layer. The third fully-connected layer embedding3, the first coding layer of the third encoder, the second coding layer of the third encoder, the third mean module, and the binary classification fully-connected layer are connected in sequence.
[0046] The feature extraction module, the conversion and fusion module, and the fusion classification module are a Transformer model, that is, a multi-layer feature fusion Transformer neural network.
[0047] Embodiment 2:
[0048] The present invention discloses a method for detecting an electric motor based on a multi-layer feature fusion Transformer neural network, including step S1 feature extraction, step S2 obtaining each layer of features to be fused by the Transformer, and step S3 feature fusion and obtaining a classification result.
[0049] As Figure 2 shown, it is the data flow diagram of the present application, which is described in detail as follows.
[0050] Step S1 Feature Extraction:
[0051] As Figure 2 shown in the upper part, a time-domain signal of the motor sound is obtained. The time-domain signal of the motor sound is subjected to MFCC feature extraction with the number of the first Mel-frequency cepstral coefficients n_mfcc1 to obtain a first spectral feature matrix with a shape of (n_frame, n_mfcc1). The time-domain signal of the motor sound is subjected to MFCC feature extraction with the number of the second Mel-frequency cepstral coefficients n_mfcc2 to obtain a second spectral feature matrix with a shape of (n_frame, n_mfcc2). n_frame is the number of frames of the Fourier transform in MFCC. The value range of the number of the first Mel-frequency cepstral coefficients n_mfcc1 is 32 to 64. The value range of the number of the second Mel-frequency cepstral coefficients n_mfcc2 is 64 to 128. The number of the second Mel-frequency cepstral coefficients n_mfcc2 is greater than the number of the first Mel-frequency cepstral coefficients n_mfcc1. The number of the second Mel-frequency cepstral coefficients n_mfcc2 is about twice the number of the first Mel-frequency cepstral coefficients n_mfcc1. The first spectral feature matrix is a low-dimensional feature matrix, and the second spectral feature matrix is a high-dimensional feature matrix.
[0052] Step S2 Obtaining Each Layer of Features to be Fused by the Transformer:
[0053] As Figure 2 shown in the left part, the first spectral feature matrix with a shape of (n_frame, n_mfcc1) is subjected to mean normalization through the first normalization module to obtain a first normalized matrix with a shape of (n_frame, n_mfcc1) after normalization. The first normalized matrix is the low-dimensional feature matrix. The first normalized matrix with a shape of (n_frame, n_mfcc1) is passed through the first fully connected layer embedding1 to obtain a first fully connected matrix embedding1 with a shape of (n_frame, n_embedding1).
[0054] The first fully-connected matrix embedding1 with the shape of (n_frame, n_embedding1) is input into each encoding layer of the first encoder Encoder1 in the first conversion and fusion module. One encoding layer obtains an encoding layer feature matrix with the shape of (n_frame, n_embedding1), and a total of n encoding layer feature matrices with the shape of (n_frame, n_embedding1) are obtained. The description is as follows.
[0055] In the first encoder Encoder1 of the first conversion and fusion module, since the Encoder of the Transformer is a sequence-to-sequence model, it does not change the dimension of the input. The first fully-connected matrix embedding1 with the shape of (n_frame, n_embedding1) passes through n encoding layers Encoder and obtains n encoding layer feature matrices with the shape of (n_frame, n_embedding1). The first encoding layer of the first encoder Encoder1 obtains the first encoding layer feature matrix of the first encoder Encoder1, the second encoding layer of the first encoder Encoder1 obtains the second encoding layer feature matrix of the first encoder Encoder1, and so on. The nth encoding layer of the first encoder Encoder1 obtains the nth encoding layer feature matrix of the first encoder Encoder1. The shape of each encoding layer feature matrix of the first encoder Encoder1 is (n_frame, n_embedding1).
[0056] Each encoding layer feature matrix with the shape of (n_frame, n_embedding1) of the first encoder Encoder1 is input into a corresponding mean layer of the first mean module. Each mean layer of the first mean module obtains an averaged feature vector with the shape of (1, n_embedding1), and a total of n averaged feature vectors with the shape of (1, n_embedding1) are obtained. The description is as follows.
[0057] Average the feature matrices of n encoding layers with the shape of (n_frame, n_embedding1) to obtain n feature vectors with the shape of (1, n_embedding1) after averaging. The feature matrix of the first encoding layer of the first encoder Encoder1 is input into the first mean layer of the first mean module to obtain the first feature vector with the shape of (1, n_embedding1) after averaging, and so on, until the feature matrix of the nth encoding layer of the first encoder Encoder1 is input into the nth mean layer of the first mean module to obtain the nth feature vector with the shape of (1, n_embedding1) after averaging. n_embedding1 is the number of features embedded by the first fully connected layer.
[0058] Input the n feature vectors with the shape of (1, n_embedding1) after averaging into the first splicing module for splicing to obtain the first spliced feature matrix with the shape of (n, n_embedding1). The first spliced feature matrix is the feature of each layer of the low-dimensional feature matrix.
[0059] As Figure 2 As shown on the right side of
[0060] Input the second fully connected matrix embedding2 with the shape of (n_frame, n_embedding2) into each encoding layer of the second encoder Encoder2 in the second conversion and fusion module. One encoding layer obtains an encoding layer feature matrix with the shape of (n_frame, n_embedding2), and a total of n encoding layer feature matrices with the shape of (n_frame, n_embedding2) are obtained. The description is as follows.
[0061] In the second encoder Encoder2 of the second conversion and fusion module, since the Encoder of the Transformer is a sequence-to-sequence model, it does not change the dimension of the input. The second fully-connected matrix embedding2 with the shape of (n_frame, n_embedding2) passes through n encoding layers Encoder and obtains n encoding layer feature matrices with the shape of (n_frame, n_embedding2). The first encoding layer of the second encoder Encoder2 obtains the first encoding layer feature matrix of the second encoder Encoder2, the second encoding layer of the second encoder Encoder2 obtains the second encoding layer feature matrix of the second encoder Encoder2, and so on. The nth encoding layer of the second encoder Encoder2 obtains the nth encoding layer feature matrix of the second encoder Encoder2. The shape of each encoding layer feature matrix of the second encoder Encoder2 is (n_frame, n_embedding2). n_embedding2 is the number of features embedded by the second fully-connected layer.
[0062] Input each encoding layer feature matrix with the shape of (n_frame, n_embedding2) of the second encoder Encoder2 into the corresponding mean layer of the second mean module. Each mean layer of the second mean module obtains an averaged feature vector with the shape of (1, n_embedding2), and a total of n averaged feature vectors with the shape of (1, n_embedding2) are obtained. The description is as follows.
[0063] Average the n encoding layer feature matrices with the shape of (n_frame, n_embedding2) to obtain n averaged feature vectors with the shape of (1, n_embedding2). The first encoding layer feature matrix of the second encoder Encoder2 is input into the first mean layer of the second mean module to obtain the first averaged feature vector with the shape of (1, n_embedding2), and so on, until the nth encoding layer feature matrix of the second encoder Encoder2 is input into the nth mean layer of the second mean module to obtain the nth averaged feature vector with the shape of (1, n_embedding2).
[0064] Input the n averaged feature vectors with the shape of (1, n_embedding2) into the second splicing module for splicing to obtain the second spliced feature matrix with the shape of (n, n_embedding2), and the second spliced feature matrix is the feature of each layer of the high-dimensional feature matrix.
[0065] Step S3 Feature fusion and obtaining classification results:
[0066] AsFigure 2 As shown below in the middle part, the first concatenated feature matrix of shape (n, n_embedding1), i.e., each layer feature of the low-dimensional feature matrix, and the second concatenated feature matrix of shape (n, n_embedding2), i.e., each layer feature of the high-dimensional feature matrix, are input into the third concatenation module for concatenation to obtain a feature fusion matrix of shape (n, n_embedding1 + n_embedding2). The feature fusion matrix of shape (n, n_embedding1 + n_embedding2) is input into the third fully connected layer embedding3 to obtain a third fully connected matrix embedding3 of shape (n, n_embedding3).
[0067] n_embedding3 is the number of features embedded by the third fully connected layer.
[0068] The third fully connected matrix embedding3 of shape (n, n_embedding3) is input into the third encoder of the conversion module. The third fully connected matrix embedding3 passes through the first coding layer and the second coding layer of the third encoder in sequence, and then passes through the third mean module to take the average value to obtain a fused Transformer feature of shape (1, n_embedding3). The fused Transformer feature of shape (1, n_embedding3) is input into the binary classification fully connected layer to obtain a binary classification result, and the binary classification result is normal or abnormal.
[0069] Data example:
[0070] This application uses the publicly available dataset MIMII Dataset to verify the superiority of the Transformer industrial motor equipment anomaly detection method based on multi-layer feature fusion. MIMII Dataset is a reliable dataset for fault industrial machine investigation and inspection. It contains the sounds generated by four industrial machines, namely valves, pumps, fans, and sliding rails. Four models (id0, id2, id4, id6) are publicly available for each valve, pump, fan, and sliding rail, and the sound data with signal-to-noise ratios of -6dB, 0dB, and 6dB are publicly available for each model. The data for each model contains normal sounds and abnormal sounds.
[0071] This application selects the pump motor sound of id0 with a signal-to-noise ratio of 0dB as the experimental data, and cuts the training set and the test set in a ratio of 8:2.
[0072] In feature extraction: n_frame = 5, n_mfcc1 = 32, n_mfcc2 = 64.
[0073] In this data embodiment, the number of second Mel-frequency cepstral coefficients \(n_{mfcc2}\) is twice the number of first Mel-frequency cepstral coefficients \(n_{mfcc1}\).
[0074] Obtain, for each layer of features to be fused in the Transformer: \(n = 4\), \(n_{embedding1}=32\), \(n_{embedding2}=32\).
[0075] In this data embodiment, the 4-layer encoder includes a first encoding layer, a second encoding layer, a third encoding layer, and a fourth encoding layer.
[0076] In feature fusion and obtaining classification results: \(n_{embedding3}=32\).
[0077] Experimental results show:
[0078] Compare a Transformer model that only uses \(n_{mfcc1}=32\) without multi-layer feature fusion, a Transformer model that only uses \(n_{mfcc2}=64\) without multi-layer feature fusion, and a Transformer model based on multi-layer feature fusion. The accuracy of the Transformer model that only uses \(n_{mfcc1}=32\) is 97.39; the accuracy of the Transformer model that only uses \(n_{mfcc1}=64\) is 97.83; the accuracy of the Transformer model based on multi-layer feature fusion is 99.13. Therefore, the Transformer model with multi-layer feature fusion fuses MFCC feature matrices with different numbers of Mel coefficients and performs feature fusion between different layers of the Transformer, increasing the expression ability between features of different layers of the model and improving the accuracy of the model.
[0079] Relative to the above embodiment, the number of second Mel-frequency cepstral coefficients \(n_{mfcc2}\) can also be about twice the number of first Mel-frequency cepstral coefficients \(n_{mfcc1}\). For example, \(n_{mfcc1}=32\), \(n_{mfcc2}=65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79\) or 80, which will not be elaborated here.
[0080] Relative to the above embodiment, it is also possible that \(n = 5\), and the 5-layer encoder includes a first encoding layer, a second encoding layer, a third encoding layer, a fourth encoding layer, and a fifth encoding layer, which will not be elaborated here.
[0081] Relative to the above embodiment, it is also possible that \(n = 6\), and the 6-layer encoder includes a first encoding layer, a second encoding layer, a third encoding layer, a fourth encoding layer, a fifth encoding layer, and a sixth encoding layer, which will not be elaborated here.
[0082] In the feature extraction stage, the time-domain data of the motor is converted into an MFCC feature matrix. Different numbers of Mel coefficients have their own advantages and disadvantages: the MFCC with a smaller number of Mel frequency cepstral coefficients can remove redundant information, reduce the computational cost, and improve the efficiency; the MFCC with a larger number of Mel frequency cepstral coefficients can capture richer speech spectral features and improve the classification effect.
[0083] Therefore, in the feature extraction stage, using only the information of one MFCC dimension may reduce the accuracy of motor sound classification.
[0084] In a neural network, cross-layer feature fusion has the following advantages: 1. Improve training efficiency and avoid overfitting - Feature fusion can reduce redundant calculations, enabling the model to find a better expression between multi-scale features, thus enhancing the generalization ability. 2. Enhance the multi-scale feature expression ability - By sharing information across layers, each layer can utilize the knowledge of the previous layer, thus fully capturing features of different scales.
[0085] To improve the accuracy of the Transformer in the motor sound classification task, the method of this application fuses MFCC feature matrices with different numbers of Mel coefficients and inputs them into the Transformer. At the same time, feature fusion is performed between different layers of the Transformer, combined with the self-attention mechanism, to achieve anomaly detection of industrial motor equipment based on multi-layer feature fusion of the Transformer, thereby improving the classification accuracy.
Claims
1. A multi-layer feature fusion Transformer neural network for detecting motors, characterized in that: It includes a feature extraction module, a conversion fusion module and a fusion classification module, wherein the feature extraction module includes a first MFCC module and a second MFCC module, and the number of Mel-frequency cepstral coefficients of the second MFCC module is greater than the number of Mel-frequency cepstral coefficients of the first MFCC module; The conversion fusion module includes a first conversion fusion module and a second conversion fusion module, the first conversion fusion module includes a first normalization module, a first fully connected layer, a first encoder, a first mean module and a first splicing module connected in sequence, and the first MFCC module is connected to the first normalization module; the second conversion fusion module includes a second normalization module, a second fully connected layer, a second encoder, a second mean module and a second splicing module connected in sequence, and the second MFCC module is connected to the second normalization module; the first encoder and the second encoder are encoders containing multiple layers of coding layers, the first mean module and the second mean module are mean modules containing multiple layers of mean layers, and the number of coding layers in each encoder is the same as the number of mean layers in the mean module; all coding layers in an encoder are connected in sequence, each coding layer in the encoder is connected to a corresponding mean layer in an adjacent mean module, and all mean layers in the mean module are connected to adjacent splicing modules; The first splicing module and the second splicing module are both connected to the fusion classification module.
2. The multi-layer feature fusion Transformer neural network for detecting motors according to claim 1, characterized in that: The number of coding layers in the encoder and the number of mean layers in the mean module are both represented by n, the n-layer encoder includes the first coding layer to the n-th coding layer, and the n-layer mean module includes the first mean layer to the n-th mean layer; The first fully connected layer, the first encoding layer to the nth encoding layer of the first encoder are connected in sequence, one encoding layer of the first encoder is connected to a mean layer of the first mean module, the first encoding layer of the first encoder is connected to the first mean layer of the first mean module, the nth encoding layer of the first encoder is connected to the nth mean layer of the first mean module, and all mean layers of the first mean module are connected to the first concatenation module; The second fully connected layer and the first to nth encoding layers of the second encoder are connected in sequence, one encoding layer of the second encoder is connected to a mean layer of the second mean module, the first encoding layer of the second encoder is connected to the first mean layer of the second mean module, the nth encoding layer of the second encoder is connected to the nth mean layer of the second mean module, and all mean layers of the second mean module are connected to the second splicing module.
3. The multi-layer feature fusion Transformer neural network for detecting motors according to claim 2, characterized in that: The value of n is in the range of 4 to 6.
4. The multi-layer feature fusion Transformer neural network for detecting motors according to claim 1, characterized in that: The fusion classification module includes a third splicing module, a third fully connected layer, a conversion module and a binary fully connected layer connected in sequence, the first splicing module and the second splicing module are both connected to the third splicing module, the conversion module includes a third encoder and a third mean module, and the third fully connected layer, the third encoder, the third mean module and the binary fully connected layer are connected in sequence.
5. The multi-layer feature fusion Transformer neural network for detecting motors according to claim 4, characterized in that: The third encoder includes a first encoding layer and a second encoding layer, and a third fully connected layer, a first encoding layer of the third encoder, a second encoding layer of the third encoder, a third mean module and a binary classification fully connected layer are connected in sequence.
6. The multi-layer feature fusion Transformer neural network for detecting motors according to claim 1, characterized in that: The number of Mel-frequency cepstral coefficients of the second MFCC module is twice the number of Mel-frequency cepstral coefficients of the first MFCC module; the number of Mel-frequency cepstral coefficients of the first MFCC module ranges from 32 to 64, and the number of Mel-frequency cepstral coefficients of the second MFCC module ranges from 64 to 128.
7. A method for detecting a motor, based on a multi-layer feature fusion Transformer neural network for detecting a motor as claimed in any one of claims 1 to 6, characterized in that: The following steps are included: Step S1: obtaining a time domain signal of the motor sound, extracting the time domain signal of the motor sound through a first MFCC module feature to obtain a first spectrum feature matrix, and extracting the time domain signal of the motor sound through a second MFCC module feature to obtain a second spectrum feature matrix; Step S2: inputting the first spectrum feature matrix into the first normalization module, and the first splicing module outputs the first splicing feature matrix; inputting the second spectrum feature matrix into the second normalization module, and the second splicing module outputs the second splicing feature matrix; Step S3: Input the first splicing feature matrix and the second splicing feature matrix into the fusion classification module to obtain a binary classification result, which is normal or abnormal.
8. The method for detecting a motor according to claim 7, characterized in that: In step S2, After the first spectrum feature matrix is input into the first normalization module, the method further includes the following steps: performing mean normalization on the first spectrum feature matrix by the first normalization module to obtain a normalized first normalized matrix; performing mean normalization on the first normalized matrix by the first fully connected layer to obtain a first fully connected matrix; The first fully connected matrix is input to each encoding layer of the first encoder in the first conversion fusion module, and each encoding layer obtains a encoding layer feature matrix; each encoding layer feature matrix of the first encoder is input to a corresponding mean layer of the first mean module, and each mean layer of the first mean module obtains an averaged feature vector; all the averaged feature vectors of the first mean module are input to the first splicing module to obtain a first splicing feature matrix; After the second spectrum feature matrix is input into the second normalization module, the method further includes the following steps: performing mean normalization on the second spectrum feature matrix by the second normalization module to obtain a normalized second normalized matrix; performing mean normalization on the second normalized matrix by the second fully connected layer to obtain a second fully connected matrix; The second fully connected matrix is input into each coding layer of the second encoder in the second conversion fusion module, and each coding layer obtains a coding layer feature matrix; each coding layer feature matrix of the second encoder is input into a corresponding mean layer of the second mean module, and each mean layer of the second mean module obtains an averaged feature vector; all the averaged feature vectors of the second mean module are input into the second splicing module to obtain a second splicing feature matrix.
9. The method for detecting a motor according to claim 7, characterized in that: In the step S1, the shape of the first spectrum feature matrix is (n_frame, n_mfcc1), and the shape of the second spectrum feature matrix is (n_frame, n_mfcc2), n_frame is the number of frames of Fourier transform in MFCC, n_mfcc1 is the number of Mel-frequency cepstral coefficients of the first MFCC module, and n_mfcc2 is the number of Mel-frequency cepstral coefficients of the second MFCC module; In step S2, the shape of the first normalized matrix is (n_frame, n_mfcc1), the shape of the first fully connected matrix is (n_frame, n_embedding1), the shape of the encoding layer feature matrix of the first encoder is (n_frame, n_embedding1), the shape of the feature vector averaged by the first mean module is (1, n_embedding1), and the shape of the first concatenated feature matrix is (n, n_embedding1); the shape of the second normalized matrix is (n_frame, n_mfcc2), the shape of the second fully connected matrix is (n_frame, n_embedding2), the shape of the encoding layer feature matrix of the second encoder is (n_frame, n_embedding2), the shape of the feature vector averaged by the second mean module is (1, n_embedding2), and the shape of the second concatenated feature matrix is (n, n_embedding2); n_embedding1 is the number of features embedded in the first fully connected layer, and n_embedding2 is the number of features embedded in the second fully connected layer.
10. The method for detecting a motor according to claim 9, characterized in that: In step S3, the step of inputting the first splicing feature matrix and the second splicing feature matrix into the fusion classification module to obtain a binary classification result includes the following steps: The first splicing feature matrix and the second splicing feature matrix are input into the third splicing module, and the first splicing feature matrix and the second splicing feature matrix are spliced by the third splicing module to obtain a feature fusion matrix with a shape of (n, n_embedding1+ n_embedding2); The feature fusion matrix with the shape of (n, n_embedding1+n_embedding2) is input into the third fully connected layer to obtain the third fully connected matrix with the shape of (n, n_embedding3); the third fully connected matrix with the shape of (n, n_embedding3) is input into the third encoder of the conversion module, and then the average value is taken through the third mean module to obtain the fused Transformer feature with the shape of (1, n_embedding3); the fused Transformer feature with the shape of (1, n_embedding3) is input into the binary classification fully connected layer to obtain the binary classification result; n_embedding3 is the number of features embedded in the third fully connected layer.
Citation Information
Patent Citations
Voice emotion recognition model and method based on complementary acoustic representation
CN115312080A
Method for detecting motor based on Transform model of UMAP
CN119622609A