Multilayer feature fusion Transform neural network and method for detecting motor
By adopting multi-layer feature fusion technology in the Transformer neural network, the MFCC feature matrix with the number of cepspectral coefficients of different Mel frequencies and the feature fusion between different layers is solved, the problem of low classification accuracy when detecting motor sound time domain signals is achieved, and higher classification accuracy is achieved.
Patent Information
- Application Number
- CN202510422442.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
When the Transformer neural network detects motor sound time domain signals, the classification accuracy is low, resulting in unsatisfactory detection results.
Multi-layer feature fusion Transformer neural network is adopted to fuse the MFCC feature matrix of the number of cepspectral coefficients of different Mel frequencies, and feature fusion is performed between different layers of Transformer.
Through multi-layer feature fusion, the classification accuracy of detecting motor sound time domain signals is improved, and the model's expression ability between different layer features is enhanced.
Smart Images

Figure CN119940417A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of motor testing, and in particular to a multi-layer feature fusion Transformer neural network and method for detecting a motor. Background Art
[0002] Common machine learning models used for motor sound classification include RNN, CNN, and Transformer.
[0003] Compared with traditional neural networks, such as RNN and CNN, Transformer has significant advantages in processing sequence data, such as sound and text, mainly in the following aspects: 1. The self-attention mechanism enables it to focus on all positions in the entire sequence, overcoming the disadvantages of RNN and CNN in long-distance dependency. 2. Strong parallel computing capabilities and faster training speed. 3. Efficient reasoning and low demand for computing resources.
[0004] Due to these advantages, Transformer has been widely used in speech recognition, sound classification, and large conversational language models such as ChatGPT and DeepSeek.
[0005] When Transformer is used to classify motor sounds, relying solely on the output of the last layer may lead to the loss of shallow local information, thereby reducing the classification accuracy.
[0006] Therefore, when using the Transformer neural network to detect the time domain signal of motor sound, the low classification accuracy of the Transformer neural network becomes a technical problem that needs to be solved urgently. Summary of the invention
[0007] The present invention provides a multi-layer feature fusion Transformer neural network and method for detecting a motor, which solves the technical problem of low classification accuracy of the time domain signal of the motor sound detected by the Transformer neural network.
[0008] In order to solve the above technical problems, the technical solutions adopted by the present invention are as follows: A multi-layer feature fusion Transformer neural network for detecting a motor, comprising a feature extraction module, a conversion fusion module and a fusion classification module, wherein the feature extraction module comprises a first MFCC module and a second MFCC module, wherein the number of Mel-frequency cepstral coefficients of the second MFCC module is greater than the number of Mel-frequency cepstral coefficients of the first MFCC module; The conversion fusion module includes a first conversion fusion module and a second conversion fusion module, the first conversion fusion module includes a first normalization module, a first fully connected layer, a first encoder, a first mean module and a first splicing module connected in sequence, and the first MFCC module is connected to the first normalization module; the second conversion fusion module includes a second normalization module, a second fully connected layer, a second encoder, a second mean module and a second splicing module connected in sequence, and the second MFCC module is connected to the second normalization module; the first encoder and the second encoder are encoders containing multiple layers of coding layers, the first mean module and the second mean module are mean modules containing multiple layers of mean layers, and the number of coding layers in each encoder is the same as the number of mean layers in the mean module; all coding layers in an encoder are connected in sequence, each coding layer in the encoder is connected to a corresponding mean layer in an adjacent mean module, and all mean layers in the mean module are connected to adjacent splicing modules; The first splicing module and the second splicing module are both connected to the fusion classification module.
[0009] A further technical solution is that: the number of coding layers in the encoder and the number of mean layers in the mean module are both represented by n, the n-layer encoder includes the first coding layer to the n-th coding layer, and the n-layer mean module includes the first mean layer to the n-th mean layer; The first fully connected layer, the first encoding layer to the nth encoding layer of the first encoder are connected in sequence, one encoding layer of the first encoder is connected to a mean layer of the first mean module, the first encoding layer of the first encoder is connected to the first mean layer of the first mean module, the nth encoding layer of the first encoder is connected to the nth mean layer of the first mean module, and all mean layers of the first mean module are connected to the first concatenation module; The second fully connected layer and the first to nth encoding layers of the second encoder are connected in sequence, one encoding layer of the second encoder is connected to a mean layer of the second mean module, the first encoding layer of the second encoder is connected to the first mean layer of the second mean module, the nth encoding layer of the second encoder is connected to the nth mean layer of the second mean module, and all mean layers of the second mean module are connected to the second splicing module.
[0010] A further technical solution is that: the value range of n is 4-6.
[0011] A further technical solution is that the fusion classification module includes a third splicing module, a third fully connected layer, a conversion module and a binary fully connected layer connected in sequence, the first splicing module and the second splicing module are both connected to the third splicing module, the conversion module includes a third encoder and a third mean module, and the third fully connected layer, the third encoder, the third mean module and the binary fully connected layer are connected in sequence.
[0012] A further technical solution is that the third encoder includes a first encoding layer and a second encoding layer, and the third fully connected layer, the first encoding layer of the third encoder, the second encoding layer of the third encoder, the third mean module and the binary fully connected layer are connected in sequence.
[0013] A further technical solution is that: the number of Mel-frequency cepstral coefficients of the second MFCC module is twice the number of Mel-frequency cepstral coefficients of the first MFCC module; the number of Mel-frequency cepstral coefficients of the first MFCC module ranges from 32 to 64, and the number of Mel-frequency cepstral coefficients of the second MFCC module ranges from 64 to 128.
[0014] A method for detecting a motor, based on the above-mentioned multi-layer feature fusion Transformer neural network for detecting a motor, comprises the following steps: Step S1: obtaining a time domain signal of the motor sound, extracting the time domain signal of the motor sound through a first MFCC module feature to obtain a first spectrum feature matrix, and extracting the time domain signal of the motor sound through a second MFCC module feature to obtain a second spectrum feature matrix; Step S2: inputting the first spectrum feature matrix into the first normalization module, and the first splicing module outputs the first splicing feature matrix; inputting the second spectrum feature matrix into the second normalization module, and the second splicing module outputs the second splicing feature matrix; Step S3: Input the first splicing feature matrix and the second splicing feature matrix into the fusion classification module to obtain a binary classification result, which is normal or abnormal.
[0015] A further technical solution is: in step S2, After the first spectrum feature matrix is input into the first normalization module, the following steps are also included: the first spectrum feature matrix is normalized by the first normalization module to obtain a normalized first normalized matrix; the first normalized matrix is passed through the first fully connected layer to obtain a first fully connected matrix; the first fully connected matrix is input into each coding layer of the first encoder in the first conversion fusion module, and each coding layer obtains a coding layer feature matrix; each coding layer feature matrix of the first encoder is input into a corresponding mean layer of the first mean module, and each mean layer of the first mean module obtains an averaged feature vector; all the averaged feature vectors of the first mean module are input into the first splicing module to obtain a first splicing feature matrix; After the second spectrum feature matrix is input into the second normalization module, the following steps are also included: the second spectrum feature matrix is mean normalized by the second normalization module to obtain a normalized second normalized matrix; the second normalized matrix is passed through the second fully connected layer to obtain a second fully connected matrix; the second fully connected matrix is input into each coding layer of the second encoder in the second conversion fusion module, and each coding layer obtains a coding layer feature matrix; each coding layer feature matrix of the second encoder is input into a corresponding mean layer of the second mean module, and each mean layer of the second mean module obtains an averaged feature vector; all the averaged feature vectors of the second mean module are input into the second splicing module to obtain a second splicing feature matrix.
[0016] A further technical solution is that: in the step S1, the shape of the first spectrum feature matrix is (n_frame, n_mfcc1), the shape of the second spectrum feature matrix is (n_frame, n_mfcc2), n_frame is the number of frames of Fourier transform in MFCC, n_mfcc1 is the number of Mel-frequency cepstral coefficients of the first MFCC module, and n_mfcc2 is the number of Mel-frequency cepstral coefficients of the second MFCC module; In step S2, the shape of the first normalized matrix is (n_frame, n_mfcc1), the shape of the first fully connected matrix is (n_frame, n_embedding1), the shape of the encoding layer feature matrix of the first encoder is (n_frame, n_embedding1), the shape of the feature vector averaged by the first mean module is (1, n_embedding1), and the shape of the first concatenated feature matrix is (n, n_embedding1); the shape of the second normalized matrix is (n_frame, n_mfcc2), the shape of the second fully connected matrix is (n_frame, n_embedding2), the shape of the encoding layer feature matrix of the second encoder is (n_frame, n_embedding2), the shape of the feature vector averaged by the second mean module is (1, n_embedding2), and the shape of the second concatenated feature matrix is (n, n_embedding2); n_embedding1 is the number of features embedded in the first fully connected layer, and n_embedding2 is the number of features embedded in the second fully connected layer.
[0017] A further technical solution is that: in step S3, the step of inputting the first splicing feature matrix and the second splicing feature matrix into the fusion classification module to obtain a binary classification result includes the following steps: The first spliced feature matrix and the second spliced feature matrix are input into the third splicing module. The first spliced feature matrix and the second spliced feature matrix are spliced through the third splicing module to obtain a feature fusion matrix with a shape of (n, n_embedding1+n_embedding2); the feature fusion matrix with a shape of (n, n_embedding1+n_embedding2) is input into the third fully connected layer to obtain a third fully connected matrix with a shape of (n, n_embedding3); the third fully connected matrix with a shape of (n, n_embedding3) is input into the third encoder of the conversion module, and then the average value is taken through the third mean module to obtain a fused Transformer feature with a shape of (1, n_embedding3); the fused Transformer feature with a shape of (1, n_embedding3) is input into the binary classification fully connected layer to obtain a binary classification result; n_embedding3 is the number of features embedded in the third fully connected layer.
[0018] The beneficial effects of adopting the above technical solution are: A multi-layer feature fusion Transformer neural network for detecting a motor comprises a feature extraction module, a conversion fusion module and a fusion classification module, wherein the feature extraction module comprises a first MFCC module and a second MFCC module, the number of Mel-frequency cepstral coefficients of the second MFCC module is greater than the number of Mel-frequency cepstral coefficients of the first MFCC module; the conversion fusion module comprises a first conversion fusion module and a second conversion fusion module, the first conversion fusion module comprises a first normalization module, a first fully connected layer, a first encoder, a first mean module and a first splicing module connected in sequence, and the first MFCC module is connected to the first normalization module; the second conversion fusion module comprises a second normalization module, a second fully connected layer, a second encoder, a second mean module and a second splicing module connected in sequence, and the second MFCC module is connected to the second normalization module; the first splicing module and the second splicing module are both connected to the fusion classification module. Through the first MFCC module, the second MFCC module, the first conversion fusion module and the second conversion fusion module, the Transformer neural network model with multi-layer feature fusion fuses MFCC feature matrices with different numbers of Mel coefficients, and performs feature fusion between different layers of the Transformer, thereby increasing the expression ability of the model between features in different layers and improving the classification accuracy of detecting the time domain signal of the motor sound.
[0019] A method for detecting a motor, based on the above-mentioned multi-layer feature fusion Transformer neural network for detecting a motor, includes the following steps: step S1: obtaining a time domain signal of the motor sound, extracting the time domain signal of the motor sound through a first MFCC module feature to obtain a first spectrum feature matrix, and extracting the time domain signal of the motor sound through a second MFCC module feature to obtain a second spectrum feature matrix; step S2: inputting the first spectrum feature matrix into a first normalization module, and the first splicing module outputs a first splicing feature matrix; inputting the second spectrum feature matrix into a second normalization module, and the second splicing module outputs a second splicing feature matrix; step S3: inputting the first splicing feature matrix and the second splicing feature matrix into a fusion classification module to obtain a binary classification result, and the binary classification result is normal or abnormal. Through steps S1 to S3, the classification accuracy of detecting the time domain signal of the motor sound is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a principle block diagram of embodiment 1 of the present invention; Figure 2 This is the data flow diagram of Example 2 of the present invention. DETAILED DESCRIPTION
[0021] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is by no means intended to limit the present application and its application or use. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0022] In the following description, many specific details are set forth to facilitate a full understanding of the present application, but the present application may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present application. Therefore, the present application is not limited to the specific embodiments disclosed below.
[0023] Embodiment 1:
[0024] like Figure 1As shown, the present invention discloses a multi-layer feature fusion Transformer neural network for detecting a motor, comprising a feature extraction module, a conversion fusion module and a fusion classification module, wherein the feature extraction module comprises a first MFCC module and a second MFCC module, wherein the number n_mfcc1 of Mel-frequency cepstral coefficients of the first MFCC module ranges from 32 to 64, the number n_mfcc2 of Mel-frequency cepstral coefficients of the second MFCC module ranges from 64 to 128, and the number n_mfcc2 of Mel-frequency cepstral coefficients of the second MFCC module is greater than the number n_mfcc1 of Mel-frequency cepstral coefficients of the first MFCC module.
[0025] The conversion fusion module includes a normalization module, a fully connected layer embedding, an encoder, a mean module and a splicing module connected in sequence, the encoder is an n-layer encoder, the n-layer encoder includes the first encoding layer to the n-th encoding layer, the mean module is an n-layer mean module, the n-layer mean module includes the first mean layer to the n-th mean layer, the normalization module, the fully connected layer embedding, the first encoding layer to the n-th encoding layer are connected in sequence, one encoding layer is connected to one mean layer, the first encoding layer is connected to the first mean layer, and all mean layers are connected to the splicing module.
[0026] The value range of n is 4 to 6.
[0027] The number of the conversion fusion modules is two, namely a first conversion fusion module and a second conversion fusion module.
[0028] The first conversion fusion module includes a first normalization module, a first fully connected layer embedding1, a first encoder Encoder1, a first mean module and a first splicing module connected in sequence, and the second conversion fusion module includes a second normalization module, a second fully connected layer embedding2, a second encoder Encoder2, a second mean module and a second splicing module connected in sequence, the first encoder Encoder1 and the second encoder Encoder2 are both n-layer encoders, the n-layer encoders include the first encoding layer to the n-th encoding layer, the first mean module and the second mean module are both n-layer mean modules, and the n-layer mean modules include the first mean layer to the n-th mean layer.
[0029] The first MFCC module, the first normalization module, the first fully connected layer embedding1, the first encoding layer to the nth encoding layer of the first encoder Encoder1 are connected in sequence, an encoding layer of the first encoder Encoder1 is connected to an average layer of the first average module, the first encoding layer of the first encoder Encoder1 is connected to the first average layer of the first average module, and so on, until the nth encoding layer of the first encoder Encoder1 is connected to the nth average layer of the first average module, and all average layers of the first average module are connected to the first splicing module.
[0030] The second MFCC module, the second normalization module, the second fully connected layer embedding2, and the first encoding layer to the nth encoding layer of the second encoder Encoder2 are connected in sequence, an encoding layer of the second encoder Encoder2 is connected to an average layer of the second average module, the first encoding layer of the second encoder Encoder2 is connected to the first average layer of the second average module, and so on, until the nth encoding layer of the second encoder Encoder2 is connected to the nth average layer of the second average module, and all average layers of the second average module are connected to the second splicing module.
[0031] The fusion classification module includes a third splicing module, a third fully connected layer embedding3, a conversion module and a binary classification fully connected layer connected in sequence. The first splicing module and the second splicing module are both connected to the third splicing module. The conversion module includes a third encoder and a third mean module. The third encoder includes a first encoding layer and a second encoding layer. The third fully connected layer embedding3, the first encoding layer of the third encoder, the second encoding layer of the third encoder, the third mean module and the binary classification fully connected layer are connected in sequence.
[0032] The Transformer model of the feature extraction module, conversion fusion module and fusion classification module is a multi-layer feature fusion Transformer neural network.
[0033] Embodiment 2:
[0034] The present invention discloses a method for detecting a motor based on a multi-layer feature fusion Transformer neural network, comprising step S1 of feature extraction, step S2 of obtaining each layer of features to be fused by the Transformer, and step S3 of feature fusion and obtaining a classification result.
[0035] like Figure 2 Shown is the data flow diagram of this application, which is described in detail as follows.
[0036] Step S1 feature extraction: like Figure 2As shown in the upper part, the time domain signal of the motor sound is obtained, and the time domain signal of the motor sound is subjected to MFCC feature extraction of the first Mel frequency cepstral coefficient number n_mfcc1 to obtain a first spectrum feature matrix with a shape of (n_frame, n_mfcc1), and the time domain signal of the motor sound is subjected to MFCC feature extraction of the second Mel frequency cepstral coefficient number n_mfcc2 to obtain a first spectrum feature matrix with a shape of (n_frame, The first spectral feature matrix is a low-dimensional feature matrix, and the second spectral feature matrix is a high-dimensional feature matrix.
[0037] Step S2 obtains each layer feature of Transformer to be fused: like Figure 2 As shown in the left part of , the first spectrum feature matrix with a shape of (n_frame, n_mfcc1) is normalized by the first normalization module to obtain the first normalized matrix with a shape of (n_frame, n_mfcc1). The first normalized matrix is the low-dimensional feature matrix. The first normalized matrix with a shape of (n_frame, n_mfcc1) is passed through the first fully connected layer embedding1 to obtain the first fully connected matrix embedding1 with a shape of (n_frame, n_embedding1).
[0038] The first fully connected matrix embedding1 with a shape of (n_frame, n_embedding1) is input to each encoding layer of the first encoder Encoder1 in the first conversion fusion module, and each encoding layer obtains an encoding layer feature matrix with a shape of (n_frame, n_embedding1), and a total of n encoding layer feature matrices with a shape of (n_frame, n_embedding1) are obtained. The explanation is as follows.
[0039] In the first encoder Encoder1 of the first conversion fusion module, since the Transformer Encoder is a sequence-to-sequence model, the input dimension will not be changed. The first fully connected matrix embedding1 with a shape of (n_frame, n_embedding1) passes through n layers of encoding layers Encoder, and n encoding layer feature matrices with a shape of (n_frame, n_embedding1) are obtained. The first encoding layer of the first encoder Encoder1 obtains the first encoding layer feature matrix of the first encoder Encoder1, the second encoding layer of the first encoder Encoder1 obtains the second encoding layer feature matrix of the first encoder Encoder1, and so on. The nth encoding layer of the first encoder Encoder1 obtains the nth encoding layer feature matrix of the first encoder Encoder1. The shape of each encoding layer feature matrix of the first encoder Encoder1 is (n_frame, n_embedding1).
[0040] Each encoding layer feature matrix of the first encoder Encoder1 with a shape of (n_frame, n_embedding1) is input to a corresponding mean layer of the first mean module. Each mean layer of the first mean module obtains an averaged feature vector with a shape of (1, n_embedding1), and a total of n averaged feature vectors with a shape of (1, n_embedding1) are obtained. The explanation is as follows.
[0041] The n encoding layer feature matrices of shape (n_frame, n_embedding1) are averaged to obtain n averaged feature vectors of shape (1, n_embedding1). The first encoding layer feature matrix of the first encoder Encoder1 is input to the first mean layer of the first mean module to obtain the first averaged feature vector of shape (1, n_embedding1), and so on, until the nth encoding layer feature matrix of the first encoder Encoder1 is input to the nth mean layer of the first mean module to obtain the nth averaged feature vector of shape (1, n_embedding1). n_embedding1 is the number of features embedded in the first fully connected layer.
[0042] The averaged feature vectors of n shapes (1, n_embedding1) are input into the first concatenation module to obtain the first concatenation feature matrix of shape (n, n_embedding1). The first concatenation feature matrix is the feature of each layer of the low-dimensional feature matrix.
[0043] like Figure 2As shown on the right side of , the second spectrum feature matrix with a shape of (n_frame, n_mfcc2) is normalized by the second normalization module to obtain a normalized second normalized matrix with a shape of (n_frame, n_mfcc2). The second normalized matrix is a high-dimensional feature matrix. The second normalized matrix with a shape of (n_frame, n_mfcc2) is passed through the second fully connected layer embedding2 to obtain the second fully connected matrix embedding2 with a shape of (n_frame, n_embedding2).
[0044] The second fully connected matrix embedding2 with a shape of (n_frame, n_embedding2) is input to each encoding layer of the second encoder Encoder2 in the second conversion fusion module, and each encoding layer obtains an encoding layer feature matrix with a shape of (n_frame, n_embedding2), and a total of n encoding layer feature matrices with a shape of (n_frame, n_embedding2) are obtained. The explanation is as follows.
[0045] In the second encoder Encoder2 of the second transformation fusion module, since the Transformer Encoder is a sequence-to-sequence model, the input dimension will not be changed. The second fully connected matrix embedding2 with a shape of (n_frame, n_embedding2) passes through n layers of encoding layers Encoder, and n encoding layer feature matrices with a shape of (n_frame, n_embedding2) are obtained. The first encoding layer of the second encoder Encoder2 obtains the first encoding layer feature matrix of the second encoder Encoder2, the second encoding layer of the second encoder Encoder2 obtains the second encoding layer feature matrix of the second encoder Encoder2, and so on. The nth encoding layer of the second encoder Encoder2 obtains the nth encoding layer feature matrix of the second encoder Encoder2. The shape of each encoding layer feature matrix of the second encoder Encoder2 is (n_frame, n_embedding2). n_embedding2 is the number of features embedded in the second fully connected layer.
[0046] Each encoding layer feature matrix of the second encoder Encoder2 with a shape of (n_frame, n_embedding2) is input to a corresponding mean layer of the second mean module, and each mean layer of the second mean module obtains an averaged feature vector with a shape of (1, n_embedding2), and a total of n averaged feature vectors with a shape of (1, n_embedding2) are obtained. The explanation is as follows.
[0047] The n encoding layer feature matrices of shape (n_frame, n_embedding2) are averaged to obtain n averaged feature vectors of shape (1, n_embedding2). The first encoding layer feature matrix of the second encoder Encoder2 is input to the first mean layer of the second mean module to obtain the first averaged feature vector of shape (1, n_embedding2), and so on, until the nth encoding layer feature matrix of the second encoder Encoder2 is input to the nth mean layer of the second mean module to obtain the nth averaged feature vector of shape (1, n_embedding2).
[0048] The averaged feature vectors of n shapes (1, n_embedding2) are input into the second concatenation module to obtain a second concatenation feature matrix of shape (n, n_embedding2). The second concatenation feature matrix is the feature of each layer of the high-dimensional feature matrix.
[0049] Step S3: feature fusion and classification results: like Figure 2 As shown in the middle part of the figure, the first concatenated feature matrix of shape (n, n_embedding1), i.e., the low-dimensional feature matrix, and the second concatenated feature matrix of shape (n, n_embedding2), i.e., the high-dimensional feature matrix, are input to the third concatenated module to obtain a feature fusion matrix of shape (n, n_embedding1+n_embedding2). The feature fusion matrix of shape (n, n_embedding1+n_embedding2) is input to the third fully connected layer embedding3 to obtain the third fully connected matrix embedding3 of shape (n, n_embedding3).
[0050] n_embedding3 is the number of features embedded in the third fully connected layer.
[0051] The third fully connected matrix embedding3 with a shape of (n, n_embedding3) is input to the third encoder of the conversion module. The third fully connected matrix embedding3 passes through the first encoding layer and the second encoding layer of the third encoder in sequence, and then passes through the third mean module to take the average value to obtain the fused Transformer feature with a shape of (1, n_embedding3). The fused Transformer feature with a shape of (1, n_embedding3) is input to the binary classification fully connected layer to obtain the binary classification result, which is normal or abnormal.
[0052] Data Example: This application uses the public dataset MIMII Dataset to verify the superiority of the Transformer industrial motor equipment anomaly detection method based on multi-layer feature fusion. MIMII Dataset is a reliable dataset for investigation and inspection of faulty industrial machines. It contains the sounds generated by four types of industrial machines, namely valves, pumps, fans, and slides. Each valve, pump, fan, and slide has 4 models (id0, id2, id4, id6) published respectively, and each model has -6db, 0db, 6db signal-to-noise ratio sound data published, and each model's data contains normal and abnormal sounds.
[0053] This application selects the pump motor sound with 0db signal-to-noise ratio and ID0 as experimental data, and divides the training set and test set into a ratio of 8:2.
[0054] During feature extraction: n_frame=5, n_mfcc1=32, n_mfcc2=64.
[0055] In this embodiment, the second number of mel-frequency cepstral coefficients n_mfcc2 is twice the first number of mel-frequency cepstral coefficients n_mfcc1.
[0056] Get the features of each layer of Transformer to be fused: n=4, n_embedding1=32, n_embedding2=32.
[0057] In this data embodiment, the 4-layer encoder includes a first encoding layer, a second encoding layer, a third encoding layer and a fourth encoding layer.
[0058] In feature fusion and classification results: n_embedding3=32.
[0059] Experimental results show: The Transformer model without multi-layer feature fusion using only n_mfcc1=32 and the Transformer model without multi-layer feature fusion using only n_mfcc2=64 are compared with the Transformer model based on multi-layer feature fusion. The accuracy of the Transformer model using only n_mfcc1=32 is 97.39; the accuracy of the Transformer model using only n_mfcc1=64 is 97.83; the accuracy of the Transformer model based on multi-layer feature fusion is 99.13. Therefore, the Transformer model with multi-layer feature fusion integrates MFCC feature matrices with different numbers of Mel coefficients and performs feature fusion between different layers of the Transformer, which increases the model's expressive ability between features at different layers and improves the accuracy of the model.
[0060] Compared with the above embodiment, the second Mel-frequency cepstral coefficient number n_mfcc2 can also be about twice the first Mel-frequency cepstral coefficient number n_mfcc1, for example, n_mfcc1=32, n_mfcc2=65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79 or 80, which will not be repeated here.
[0061] With respect to the above embodiment, n may also be 5, and the 5-layer encoder includes a first encoding layer, a second encoding layer, a third encoding layer, a fourth encoding layer and a fifth encoding layer, which will not be described in detail.
[0062] With respect to the above embodiment, n may also be 6, and the 6-layer encoder includes a first encoding layer, a second encoding layer, a third encoding layer, a fourth encoding layer, a fifth encoding layer and a sixth encoding layer, which will not be described in detail.
[0063] In the feature extraction stage, the time domain data of the motor is converted into the MFCC feature matrix. Different numbers of Mel coefficients have their own advantages and disadvantages: MFCCs with a smaller number of Mel frequency cepstral coefficients can remove redundant information, reduce computing costs, and improve efficiency; MFCCs with a higher number of Mel frequency cepstral coefficients can capture richer speech spectrum features and improve classification results.
[0064] Therefore, in the feature extraction stage, using information from only one MFCC dimension may reduce the accuracy of motor sound classification.
[0065] In neural networks, cross-layer feature fusion has the following advantages: 1. Improve training efficiency and avoid overfitting - feature fusion can reduce redundant calculations, allowing the model to find a better expression between multi-scale features, thereby improving generalization capabilities. 2. Enhance the ability to express multi-scale features - by sharing information across layers, each layer can use the knowledge of the previous layer to fully capture features of different scales.
[0066] In order to improve the accuracy of Transformer in motor sound classification tasks, the method of this application combines MFCC feature matrices with different numbers of Mel coefficients and inputs them into Transformer. At the same time, feature fusion is performed between different layers of Transformer, combined with the self-attention mechanism, to achieve Transformer industrial motor equipment anomaly detection based on multi-layer feature fusion, thereby improving classification accuracy.
Claims
1. A multi-layer feature fusion Transformer neural network for detecting motors, characterized in that: It includes a feature extraction module, a conversion fusion module and a fusion classification module, wherein the feature extraction module includes a first MFCC module and a second MFCC module, and the number of Mel-frequency cepstral coefficients of the second MFCC module is greater than the number of Mel-frequency cepstral coefficients of the first MFCC module; The conversion fusion module includes a first conversion fusion module and a second conversion fusion module, the first conversion fusion module includes a first normalization module, a first fully connected layer, a first encoder, a first mean module and a first splicing module connected in sequence, and the first MFCC module is connected to the first normalization module; the second conversion fusion module includes a second normalization module, a second fully connected layer, a second encoder, a second mean module and a second splicing module connected in sequence, and the second MFCC module is connected to the second normalization module; the first encoder and the second encoder are encoders containing multiple layers of coding layers, the first mean module and the second mean module are mean modules containing multiple layers of mean layers, and the number of coding layers in each encoder is the same as the number of mean layers in the mean module; all coding layers in an encoder are connected in sequence, each coding layer in the encoder is connected to a corresponding mean layer in an adjacent mean module, and all mean layers in the mean module are connected to adjacent splicing modules; The first splicing module and the second splicing module are both connected to the fusion classification module.
2. The multi-layer feature fusion Transformer neural network for detecting motors according to claim 1, characterized in that: The number of coding layers in the encoder and the number of mean layers in the mean module are both represented by n, the n-layer encoder includes the first coding layer to the n-th coding layer, and the n-layer mean module includes the first mean layer to the n-th mean layer; The first fully connected layer, the first encoding layer to the nth encoding layer of the first encoder are connected in sequence, one encoding layer of the first encoder is connected to a mean layer of the first mean module, the first encoding layer of the first encoder is connected to the first mean layer of the first mean module, the nth encoding layer of the first encoder is connected to the nth mean layer of the first mean module, and all mean layers of the first mean module are connected to the first concatenation module; The second fully connected layer and the first to nth encoding layers of the second encoder are connected in sequence, one encoding layer of the second encoder is connected to a mean layer of the second mean module, the first encoding layer of the second encoder is connected to the first mean layer of the second mean module, the nth encoding layer of the second encoder is connected to the nth mean layer of the second mean module, and all mean layers of the second mean module are connected to the second splicing module.
3. The multi-layer feature fusion Transformer neural network for detecting motors according to claim 2, characterized in that: The value of n is in the range of 4 to 6.
4. The multi-layer feature fusion Transformer neural network for detecting motors according to claim 1, characterized in that: The fusion classification module includes a third splicing module, a third fully connected layer, a conversion module and a binary fully connected layer connected in sequence, the first splicing module and the second splicing module are both connected to the third splicing module, the conversion module includes a third encoder and a third mean module, and the third fully connected layer, the third encoder, the third mean module and the binary fully connected layer are connected in sequence.
5. The multi-layer feature fusion Transformer neural network for detecting motors according to claim 4, characterized in that: The third encoder includes a first encoding layer and a second encoding layer, and a third fully connected layer, a first encoding layer of the third encoder, a second encoding layer of the third encoder, a third mean module and a binary classification fully connected layer are connected in sequence.
6. The multi-layer feature fusion Transformer neural network for detecting motors according to claim 1, characterized in that: The number of Mel-frequency cepstral coefficients of the second MFCC module is twice the number of Mel-frequency cepstral coefficients of the first MFCC module; the number of Mel-frequency cepstral coefficients of the first MFCC module ranges from 32 to 64, and the number of Mel-frequency cepstral coefficients of the second MFCC module ranges from 64 to 128.
7. A method for detecting a motor, based on a multi-layer feature fusion Transformer neural network for detecting a motor as claimed in any one of claims 1 to 6, characterized in that: The following steps are included: Step S1: obtaining a time domain signal of the motor sound, extracting the time domain signal of the motor sound through a first MFCC module feature to obtain a first spectrum feature matrix, and extracting the time domain signal of the motor sound through a second MFCC module feature to obtain a second spectrum feature matrix; Step S2: inputting the first spectrum feature matrix into the first normalization module, and the first splicing module outputs the first splicing feature matrix; inputting the second spectrum feature matrix into the second normalization module, and the second splicing module outputs the second splicing feature matrix; Step S3: Input the first splicing feature matrix and the second splicing feature matrix into the fusion classification module to obtain a binary classification result, which is normal or abnormal.
8. The method for detecting a motor according to claim 7, characterized in that: In step S2, After the first spectrum feature matrix is input into the first normalization module, the method further includes the following steps: performing mean normalization on the first spectrum feature matrix by the first normalization module to obtain a normalized first normalized matrix; performing mean normalization on the first normalized matrix by the first fully connected layer to obtain a first fully connected matrix; The first fully connected matrix is input to each encoding layer of the first encoder in the first conversion fusion module, and each encoding layer obtains a encoding layer feature matrix; each encoding layer feature matrix of the first encoder is input to a corresponding mean layer of the first mean module, and each mean layer of the first mean module obtains an averaged feature vector; all the averaged feature vectors of the first mean module are input to the first splicing module to obtain a first splicing feature matrix; After the second spectrum feature matrix is input into the second normalization module, the method further includes the following steps: performing mean normalization on the second spectrum feature matrix by the second normalization module to obtain a normalized second normalized matrix; performing mean normalization on the second normalized matrix by the second fully connected layer to obtain a second fully connected matrix; The second fully connected matrix is input into each coding layer of the second encoder in the second conversion fusion module, and each coding layer obtains a coding layer feature matrix; each coding layer feature matrix of the second encoder is input into a corresponding mean layer of the second mean module, and each mean layer of the second mean module obtains an averaged feature vector; all the averaged feature vectors of the second mean module are input into the second splicing module to obtain a second splicing feature matrix.
9. The method for detecting a motor according to claim 7, characterized in that: In the step S1, the shape of the first spectrum feature matrix is (n_frame, n_mfcc1), and the shape of the second spectrum feature matrix is (n_frame, n_mfcc2), n_frame is the number of frames of Fourier transform in MFCC, n_mfcc1 is the number of Mel-frequency cepstral coefficients of the first MFCC module, and n_mfcc2 is the number of Mel-frequency cepstral coefficients of the second MFCC module; In step S2, the shape of the first normalized matrix is (n_frame, n_mfcc1), the shape of the first fully connected matrix is (n_frame, n_embedding1), the shape of the encoding layer feature matrix of the first encoder is (n_frame, n_embedding1), the shape of the feature vector averaged by the first mean module is (1, n_embedding1), and the shape of the first concatenated feature matrix is (n, n_embedding1); the shape of the second normalized matrix is (n_frame, n_mfcc2), the shape of the second fully connected matrix is (n_frame, n_embedding2), the shape of the encoding layer feature matrix of the second encoder is (n_frame, n_embedding2), the shape of the feature vector averaged by the second mean module is (1, n_embedding2), and the shape of the second concatenated feature matrix is (n, n_embedding2); n_embedding1 is the number of features embedded in the first fully connected layer, and n_embedding2 is the number of features embedded in the second fully connected layer.
10. The method for detecting a motor according to claim 9, characterized in that: In step S3, the step of inputting the first splicing feature matrix and the second splicing feature matrix into the fusion classification module to obtain a binary classification result includes the following steps: The first splicing feature matrix and the second splicing feature matrix are input into the third splicing module, and the first splicing feature matrix and the second splicing feature matrix are spliced by the third splicing module to obtain a feature fusion matrix with a shape of (n, n_embedding1+ n_embedding2); The feature fusion matrix with the shape of (n, n_embedding1+n_embedding2) is input into the third fully connected layer to obtain the third fully connected matrix with the shape of (n, n_embedding3); the third fully connected matrix with the shape of (n, n_embedding3) is input into the third encoder of the conversion module, and then the average value is taken through the third mean module to obtain the fused Transformer feature with the shape of (1, n_embedding3); the fused Transformer feature with the shape of (1, n_embedding3) is input into the binary classification fully connected layer to obtain the binary classification result; n_embedding3 is the number of features embedded in the third fully connected layer.
Citation Information
Patent Citations
Voice emotion recognition model and method based on complementary acoustic representation
CN115312080A
Method for detecting motor based on Transform model of UMAP
CN119622609A
Cited By
Transform neural network and method for multi-feature grouping
CN120804841A
Transform neural network and method for detecting motor based on SE attention mechanism
CN121031671A
Transformer neural network and method for detecting motor based on SE attention mechanism
CN121031671B
A motor abnormal sound detection method based on feature aggregation transformer combined with DenseNet
CN122658357A