A method for detecting motors based on the Transformer model of UMAP

By using UMAP feature compression technology in the Transformer model, the motor sound data with more noise is processed, and the problem of low classification accuracy in the existing technology is solved, and higher motor sound classification accuracy and recognition efficiency are achieved.

CN119622609BActive Publication Date: 2025-06-03SUZHOU ACOUSTIC IND TECH RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510159184.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-03
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

When the prior art processes motor sound data with high noise and low signal-to-noise, the classification accuracy is low, making it difficult to effectively identify abnormal motor sounds.

Method used

UMAP-based Transformer model is adopted to remove redundant information through UMAP feature compression, and multiple UMAPs are used to maintain local details and overall structure of the data. Finally, the feature matrix normalized by input means is classified into the Transformer model.

Benefits of technology

It improves the classification accuracy of motor sound data with more noise and lower signal-to-noise, effectively recognizes abnormal motor sounds, and reduces the occupation of manpower and economic resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119622609B_ABST
    Figure CN119622609B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting motors based on the Transformer model with UMAP, which relates to the technical field of motor testing, especially for motor sound data with relatively low signal-to-noise ratio; it includes obtaining a shape (n_frame, n_mfcc) feature matrix; obtaining K kinds of UMAP; splitting the feature matrix into n_frame shape (n_mfcc, 1) feature vectors; respectively inputting each into each UMAP to obtain K * n_frame trained UMAPs; inputting all feature vectors into each trained UMAP to obtain n_frame shape (n_umap_dim, 1) feature vectors and concatenating them into a shape (n_frame, n_umap_dim * K) feature matrix; obtaining a classification result through the Transformer model after mean normalization; it removes redundant information in the features through K kinds of UMAP to improve the accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of motor testing, and in particular to a method for detecting a motor based on a Transformer model of UMAP. Background Art

[0002] Motors are important electromechanical devices in modern production and life, and are of great significance to the smooth progress of industrial processes and the safe operation of equipment. However, the motors produced are prone to abnormal conditions, resulting in faults and thus potential safety hazards. Quickly discovering abnormal equipment can reduce the number of defective products and prevent the spread of damage. The normality or abnormality of a motor can be judged by the sound of the motor during operation. For manual detection, manual detection of abnormal sounds and other features will greatly consume human and material resources, but machine recognition of abnormal equipment can reduce the losses caused by manpower and economy.

[0003] The steps for a machine to recognize abnormal equipment are: first feature extraction, and then model prediction.

[0004] In the feature extraction step of the motor, the common method is to extract the time-domain data of the motor into a feature matrix through Mel Frequency Cepstral Coefficients (MFCC for short). However, sometimes the motor sound is accompanied by noise, so the extracted feature matrix often has a lot of useless redundant information. Therefore, the redundant information can be removed by means of feature compression. Common feature dimensionality reduction methods include Principal Component Analysis (PCA), t-Distributed Stochastic Neighbor Embedding (t-SNE), and Uniform Manifold Approximation and Projection (UMAP). However, PCA is a linear feature compression, which limits the application scenarios of PCA, and t-SNE is slow and cannot preserve global information.

[0005] Common classification methods include time-domain analysis, frequency-domain analysis, Support Vector Machine (SVM) in machine learning, Convolutional Neural Network (CNN) in deep learning, Long Short Term Memory-Recurrent Neural Networks (LSTM-RNN) in deep learning, and Transformer model in deep learning.

[0006] For the complex situation where abnormal sounds of industrial equipment are accompanied by noise, time domain analysis and frequency domain analysis cannot effectively extract features, thereby reducing the accuracy of classification. Machine learning models such as support vector machines (SVM) effectively fit the data, but their generalization ability is insufficient.

[0007] Therefore, improving the classification accuracy of motor sound data with high noise and low signal-to-noise ratio has become a technical problem that needs to be solved urgently. Summary of the invention

[0008] The present invention provides a method for detecting a motor based on a Transformer model of UMAP, which solves the technical problem of low classification accuracy of motor sound data with more noise and lower signal-to-noise ratio.

[0009] In order to solve the above technical problems, the technical solutions adopted by the present invention are as follows:

[0010] A method for detecting a motor based on a Transformer model of UMAP comprises the following steps:

[0011] Step S1: extracting features from the time domain signal of the motor to obtain a feature matrix of shape (n_frame, n_mfcc), wherein the time domain signal is a time domain signal of a sound signal or a time domain signal of a vibration signal;

[0012] Step S2: UMAP feature compression, step S2 includes steps S201 to S204,

[0013] Step S201: setting K n_neighbor parameters to obtain K different types of UMAPs;

[0014] Step S202: split the feature matrix with a shape of (n_frame, n_mfcc) obtained in step S1 to obtain n_frame feature vectors with a shape of (n_mfcc, 1);

[0015] Step S203: Each feature vector of shape (n_mfcc, 1) is independently input to each UMAP, and one UMAP is trained to obtain n_frame trained UMAPs of the same type and K different types of UMAPs, and a total of K*n_frame trained UMAPs are obtained;

[0016] Step S204: Input n_frame feature vectors of shape (n_mfcc, 1) into each trained UMAP. Each trained UMAP outputs n_frame feature vectors of shape (n_umap_dim, 1). Concatenate the n_frame feature vectors of shape (n_umap_dim, 1) to obtain a feature matrix of shape (n_frame, n_umap_dim). For K different types of UMAPs, a total of K feature matrices of shape (n_frame, n_umap_dim) are obtained. Concatenate the K feature matrices of shape (n_frame, n_umap_dim) to obtain a feature matrix of shape (n_frame, n_umap_dim*K);

[0017] Step S3: Input the feature matrix of shape (n_frame, n_umap_dim*K) obtained in Step S2 into the Transformer model after mean normalization. The Transformer model outputs the classification result.

[0018] A further technical solution lies in: In the above Step S1, Mel Frequency Cepstral Coefficient (MFCC) feature extraction is used.

[0019] A further technical solution lies in: It also includes a training process. Training process: Obtain a training data set. After feature extraction in Step S1, a feature matrix of shape (n_frame, n_mfcc) for training is obtained; Obtain K trained UMAPs through Step S201, Step S202, and Step S203, a total of K*n_frame trained UMAPs are obtained; Obtain a feature matrix of shape (n_frame, n_umap_dim*K) for training through Step S204; Use the feature matrix of shape (n_frame, n_umap_dim*K) obtained in Step S204 to train the normalization module to obtain a trained normalization module. Use the normalized features output by the trained normalization module to train the Transformer model to obtain a trained Transformer model, that is, a trained Transformer model based on K types of UMAPs.

[0020] A further technical solution lies in that: the steps of the training normalization module include calculating the mean and standard deviation of each dimension of the feature matrix with the shape of (n_frame, n_umap_dim * K) for training respectively to obtain the trained mean matrix and standard deviation matrix with the shape of (n_frame, n_umap_dim * K), and subtracting the trained mean matrix from the feature matrix with the shape of (n_frame, n_umap_dim * K) for training and then dividing by the trained standard deviation matrix to obtain the normalized training feature matrix with the shape of (n_frame, n_umap_dim * K); the trained mean matrix and standard deviation matrix form the trained normalization module, and the features in the normalized training feature matrix with the shape of (n_frame, n_umap_dim * K) are the normalized features output by the trained normalization module; the Transformer model includes a fully connected layer, an Encoder layer, a mean function and a fully connected classification layer connected in sequence, and the steps of training the Transformer model include inputting the normalized training feature matrix with the shape of (n_frame, n_umap_dim * K) into the Transformer model, the Transformer model outputs a classification result, and during training, using backpropagation to optimize the parameters of the Transformer model according to the loss value calculated by the loss function, and obtaining the trained Transformer model after multiple iterations.

[0021] A further technical solution lies in that: it further includes a testing process. The testing process: obtain a test data set, through feature extraction in step S1, obtain a feature matrix with the shape of (n_frame, n_mfcc) for testing; through step S202, K * n_frame trained UMAPs and step S204, obtain a feature matrix with the shape of (n_frame, n_umap_dim * K) for testing; use the feature matrix with the shape of (n_frame, n_umap_dim * K) for testing obtained in step S204 to be predicted by the trained normalization model and the trained Transformer model in sequence to obtain the classification result of the test set.

[0022] A further technical solution lies in that: the step of predicting through the trained normalization model and the trained Transformer model in sequence includes subtracting the trained mean matrix from the feature matrix with the shape of (n_frame, n_umap_dim * K) for testing and dividing the result by the trained standard deviation matrix to obtain the normalized test feature matrix with the shape of (n_frame, n_umap_dim * K), and inputting the normalized test feature matrix with the shape of (n_frame, n_umap_dim * K) into the trained Transformer model to obtain the classification result.

[0023] A further technical solution lies in that: in the step S201, K = 2 or K = 3.

[0024] A further technical solution lies in that: in the step S201, when K = 2, the n_neighbor of the first UMAP is 10 and the n_neighbor of the second UMAP is 20; when K = 3, the n_neighbor of the first UMAP is 10, the n_neighbor of the second UMAP is 15, and the n_neighbor of the third UMAP is 20.

[0025] A further technical solution lies in that: in the step S1, Filterbank filter bank features are extracted, or Mel Spectrogram features are extracted.

[0026] The beneficial effects produced by adopting the above technical solutions are as follows:

[0027] A method for detecting motors based on a Transformer model with UMAP, comprising the following steps: Step S1: The time-domain signal of the motor is subjected to feature extraction to obtain a feature matrix with a shape of (n_frame, n_mfcc); Step S2: Set K n_neighbor parameters to obtain K types of UMAP; Split the feature matrix with a shape of (n_frame, n_mfcc) to obtain n_frame feature vectors with a shape of (n_mfcc, 1); Input them into each UMAP respectively, and a total of K*n_frame trained UMAPs are obtained; Input the n_frame feature vectors with a shape of (n_mfcc, 1) into each trained UMAP to obtain n_frame feature vectors with a shape of (n_umap_dim, 1), and splice them to obtain a feature matrix with a shape of (n_frame, n_umap_dim), and a total of K feature matrices with a shape of (n_frame, n_umap_dim) are obtained, and they are spliced to obtain a feature matrix with a shape of (n_frame, n_umap_dim*K); Step S3: After mean normalization, obtain the classification result through the Transformer model; It removes redundant information in the features through K types of UMAP to improve the accuracy.

[0028] See the description in the specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is the flowchart of the present invention;

[0030] Figure 2 is the flowchart of the UMAP feature compression step in the training process and the testing process;

[0031] Figure 3 is the data flow diagram of the UMAP feature compression step in the training process and the testing process. SPECIFIC IMPLEMENTATION MANNER

[0032] The purpose of this application is to provide a method for detecting motors based on a Transformer model with multiple UMAPs for the noisy motor sounds in industry.

[0033] By setting the n_neighbor parameter in UMAP to different values, different types of UMAP can be obtained. The n_neighbor parameter in UMAP is an important parameter affecting the feature compression result, which determines the local neighborhood range of each point in the dataset. A smaller n_neighbor value makes UMAP more inclined to preserve the local details of the data, while a larger n_neighbor value makes the model pay more attention to the overall structure of the data. Therefore, when compressing the feature matrix by UMAP, different types of UMAP with K different n_neighbor parameters can be used simultaneously, so that the compressed feature matrix can preserve both the local details and the overall structure of the data.

[0034] In the field of deep learning, compared with convolutional neural networks (CNNs) and recurrent neural networks (RNNs), Transformers have the advantages of parallel processing, long-distance dependence modeling, stronger context capture ability, and easier expansion and adjustment. Therefore, they have been successfully applied to many fields such as natural language processing, computer vision, and audio processing, and have replaced CNNs and RNNs.

[0035] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and in no way limits the present application and its application or use. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0036] In the following description, many specific details are set forth in order to fully understand the present application. However, the present application may also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.

[0037] As Figures 1 to 3 shown, the present invention discloses a method for detecting motors based on a Transformer model using UMAP, which includes the following steps:

[0038] Step S1: MFCC feature extraction.

[0039] The time-domain signal of the motor is subjected to Mel Frequency Cepstral Coefficient (MFCC) feature extraction to obtain a feature matrix with a shape of (n_frame, n_mfcc).

[0040] Training process:

[0041] Obtain a training dataset. Through step S1: MFCC feature extraction, obtain a feature matrix of shape (n_frame, n_mfcc) for training.

[0042] Testing process:

[0043] Obtain a test dataset. Through step S1: MFCC feature extraction, obtain a feature matrix of shape (n_frame, n_mfcc) for testing.

[0044] Details are as follows:

[0045] As Figure 1 shown, for the time-domain signal of the motor, through Mel Frequency Cepstral Coefficients (MFCC for short), a feature matrix of shape (n_frame, n_mfcc) is obtained. n_frame is the number of frames of the MFCC Fourier transform, and n_mfcc is the dimension of MFCC feature extraction.

[0046] In feature extraction, Mel Frequency Cepstral Coefficients (MFCC for short) is a commonly used method for sound feature extraction. It extracts useful features that can represent the audio signal by imitating the human ear's perception of sounds at different frequencies. The specific steps are as follows: Input the time-domain motor sound data, define the number of frames of the Fourier transform (n_frame) and the dimension of MFCC feature extraction (n_mfcc), and a feature matrix of shape (n_frame, n_mfcc) is obtained.

[0047] Step S2: UMAP feature compression.

[0048] Step S201: Set the K n_neighbor parameters to obtain K different types of UMAPs;

[0049] Step S202: Split the feature matrix of shape (n_frame, n_mfcc) obtained in step S1 to obtain n_frame feature vectors of shape (n_mfcc, 1);

[0050] Step S203: Independently input each feature vector of shape (n_mfcc, 1) into each UMAP. One type of UMAP is trained to obtain n_frame trained UMAPs of the same type. For K different types of UMAPs, a total of K * n_frame trained UMAPs are obtained;

[0051] Step S204: n_frame feature vectors with a shape of (n_mfcc, 1) are input to each trained UMAP, and each trained UMAP outputs n_frame feature vectors with a shape of (n_umap_dim, 1). The n_frame feature vectors with a shape of (n_umap_dim, 1) are concatenated to obtain a feature matrix with a shape of (n_frame, n_umap_dim). For K different types of UMAPs, a total of K feature matrices with a shape of (n_frame, n_umap_dim) are obtained. The K feature matrices with a shape of (n_frame, n_umap_dim) are concatenated to obtain a feature matrix with a shape of (n_frame, n_umap_dim*K).

[0052] Training process:

[0053] The feature matrix of the shape (n_frame, n_mfcc) obtained by MFCC feature extraction in step S1 is subjected to steps S201, S202 and S203 to obtain K trained UMAPs, and a total of K*n_frame trained UMAPs are obtained; and the feature matrix of the shape (n_frame, n_umap_dim*K) used for training is obtained in step S204.

[0054] Testing process:

[0055] The feature matrix with the shape of (n_frame, n_mfcc) obtained by step S1: MFCC feature extraction for testing is subjected to step S202, K*n_frame trained UMAPs and step S204 to obtain a feature matrix with the shape of (n_frame,n_umap_dim*K) for testing.

[0056] Details are as follows:

[0057] In industrial motor anomaly detection, motor sound is often accompanied by large noise, so there is a lot of useless redundant information in the feature extraction. Therefore, feature compression can be used to remove redundant information. UMAP is a feature compression method. The core idea of ​​UMAP is to maintain the structure of the data as much as possible, map high-dimensional data to low-dimensional space by constructing and optimizing graphs, and try to preserve the topological structure in the high-dimensional space.

[0058] UMAP's technical ideas:

[0059] UMAP training steps technical ideas:

[0060] 1) Build a neighborhood graph:

[0061] Input a set of n_train_sample training feature vectors of n_dim_pre dimensions. For each data point in the feature vector set, UMAP will find the nearest n_neighbor neighbor points, calculate the distances, and assign weights to each edge to construct a neighborhood graph.

[0062] 2) Mapped neighborhood graph after fitting:

[0063] Continuously optimize and converge to map the high-dimensional graph to a low-dimensional space, compress a set of n_train_sample training feature vectors of n_dim_pre dimensions into a set of n_train_sample training feature vectors of n_dim_post dimensions, and save the mapped neighborhood graph after fitting.

[0064] Technical idea of UMAP testing steps: Insert a set of n_test_sample test feature vectors of n_dim_pre dimensions into the trained low-dimensional space through the mapped neighborhood graph after fitting to obtain a set of n_test_sample test feature vectors of n_dim_post dimensions.

[0065] As Figure 2 shown, the features extracted by UMAP compression of MFCC are divided into training steps and testing steps, Figure 2 belong to Figure 1 a part of.

[0066] UMAP feature compression is divided into training steps and testing steps, that is, the features extracted by UMAP compression of MFCC are divided into training steps and testing steps. The specific steps are described in detail as follows.

[0067] Since the features after MFCC feature extraction are a feature matrix of shape (n_frame, n_mfcc), and UMAP compresses a set of feature vectors, the (n_frame, n_mfcc) feature matrix is split into n_frame feature vectors of shape (n_mfcc, 1), and then each feature vector is independently input into the UMAP operation.

[0068] A training step for UMAP to compress the MFCC feature matrix is as follows: For the data in the training set, extract features into a feature matrix of shape (n_frame, n_mfcc) through MFCC, and then split it into n_frame feature vectors of shape (n_mfcc, 1). Train n_frame UMAPs, and then input the n_frame feature vectors of shape (n_mfcc, 1) into the trained n_frame UMAPs to obtain n_frame feature vectors of shape (n_umap_dim, 1), and then concatenate them into a feature matrix of (n_frame, n_umap_dim).

[0069] Therefore, the specific steps for training multiple UMAP-compressed MFCC feature matrices are as follows: 1) Input the training feature matrix after MFCC feature extraction with a shape of (n_frame, n_mfcc). 2) Define K different types of UMAP with different n_neighbor parameters. Each different type of UMAP trains n_frame well-trained UMAPs respectively, and at the same time obtains a feature matrix with a shape of (n_frame, n_umap_dim). For K different types of UMAP, a total of K * n_frame well-trained UMAPs and K feature matrices with a shape of (n_frame, n_umap_dim) are obtained. 3) Concatenate the K feature matrices with a shape of (n_frame, n_umap_dim) into a feature matrix with a shape of (n_frame, n_umap_dim * K).

[0070] The testing steps for a UMAP-compressed MFCC feature matrix are as follows: The data of the test set is used to extract features through MFCC into a feature matrix with a shape of (n_frame, n_mfcc), and then split into n_frame feature vectors with a shape of (n_mfcc, 1). Then, the n_frame feature vectors with a shape of (n_mfcc, 1) are input into the n_frame well-trained UMAPs to obtain n_frame feature vectors with a shape of (n_umap_dim, 1), and then concatenated into a feature matrix with a shape of (n_frame, n_umap_dim).

[0071] As Figure 3 shown, a data flow diagram for training and testing a UMAP MFCC feature matrix. Figure 3 is Figure 2 a part of

[0072] At the same time, the n_neighbor parameter in UMAP is an important parameter that affects the dimensionality reduction result and determines the local neighborhood range of each point in the dataset. A smaller n_neighbor value makes UMAP more inclined to preserve the local details of the data, while a larger n_neighbor value makes the model pay more attention to the overall structure of the data. Therefore, when compressing the MFCC feature matrix with UMAP, K different types of UMAP with different n_neighbor parameters can be used simultaneously, so that the compressed MFCC feature matrix can preserve both the local details and the overall structure of the data.

[0073] The specific steps for testing the UMAP-compressed MFCC feature matrix are as follows: 1) Input the test feature matrix after MFCC feature extraction with a shape of (n_frame, n_mfcc). 2) Obtain K feature matrices with a shape of (n_frame, n_umap_dim) from the K * n_frame trained UMAP-compressed test feature matrices. 3) Concatenate the K feature matrices with a shape of (n_frame, n_umap_dim) into a feature matrix of (n_frame, n_umap_dim * K).

[0074] Step S3: Obtain the classification result.

[0075] The feature matrix with a shape of (n_frame, n_umap_dim * K) obtained in step S2 is mean-normalized and then input into the Transformer model, and the Transformer model outputs to obtain the classification result.

[0076] After mean normalization, the shape of the feature matrix remains unchanged.

[0077] Training process:

[0078] Use the feature matrix with a shape of (n_frame, n_umap_dim * K) obtained in step S204 for training the normalization module to obtain the trained normalization module, and use the normalized features output by the trained normalization module to train the Transformer model to obtain the trained Transformer model, that is, the trained Transformer model based on K kinds of UMAP.

[0079] Training process of the normalization module:

[0080] Use the feature matrix with the shape of (n_frame, n_umap_dim * K) for training as the training data, that is, the training feature matrix. In the training stage, calculate the mean and standard deviation for each dimension of the data in the training feature matrix with the shape of (n_frame, n_umap_dim * K) respectively to obtain the mean matrix and standard deviation matrix with the shape of (n_frame, n_umap_dim * K). Subtract the mean matrix with the shape of (n_frame, n_umap_dim * K) from the training feature matrix with the shape of (n_frame, n_umap_dim * K) and divide by the standard deviation matrix with the shape of (n_frame, n_umap_dim * K) to obtain the normalized training feature matrix with the shape of (n_frame, n_umap_dim * K), that is, the normalization module is trained. The features in the normalized training feature matrix with the shape of (n_frame, n_umap_dim * K) are the normalized features output by the trained normalization module. The mean matrix with the shape of (n_frame, n_umap_dim * K) is the trained mean matrix, and the standard deviation matrix with the shape of (n_frame, n_umap_dim * K) is the trained standard deviation matrix. The trained mean matrix and standard deviation matrix form the trained normalization module.

[0081] Further explanation, perform mean normalization on the training feature matrix, that is, for each element in the matrix, subtract the mean of the corresponding dimension and divide by the standard deviation of that dimension to obtain the normalized training matrix. The dimension of the normalized training matrix is still (n_frame, n_umap_dim * K).

[0082] Training process of the Transformer model:

[0083] The Transformer model includes a fully connected layer, an Encoder layer, a mean function, and a fully connected classification layer connected in sequence. Input the normalized training feature matrix with the shape of (n_frame, n_umap_dim * K) into the Transformer model, and the Transformer model outputs the classification result. During training, use backpropagation to optimize the parameters of the Transformer model according to the loss value calculated by the loss function, and obtain the trained Transformer model after multiple iterations. Details are as follows:

[0084] First, use the normalized training feature matrix with the shape of (n_frame, n_umap_dim * K) as the input. This input passes through a fully connected layer with the shape of (n_umap_dim * K, n_umap_dim * K), and outputs a training Transformer fully connected layer matrix with the shape of (n_frame, n_umap_dim * K).

[0085] Next, send the training Transformer fully connected layer matrix into the Encoder layer of the Transformer model. This layer adopts a four-head attention mechanism. Due to its sequence-to-sequence mapping characteristics, the input and output dimensions remain unchanged, and a training Transformer Encoder layer matrix with the shape of (n_frame, n_umap_dim * K) is obtained.

[0086] Then, take the mean of the training Transformer Encoder layer matrix to obtain a training Transformer mean vector with the shape of (1, n_umap_dim * K). This vector is a comprehensive representation of the input sequence.

[0087] Finally, input this training Transformer mean vector into a fully connected classification layer with the shape of (n_umap_dim * K, 2), and map it to a two-dimensional space through a linear transformation to obtain the classification result. During training, use backpropagation to optimize the model parameters according to the loss value calculated by the loss function, and obtain a trained Transformer model after multiple iterations.

[0088] Testing process:

[0089] Use the feature matrix with the shape of (n_frame, n_umap_dim * K) obtained in step S204 for testing and predict it successively through the trained normalization model and the trained Transformer model to obtain the accuracy of the test set.

[0090] Testing process of the normalization module:

[0091] Use the feature matrix with the shape of (n_frame, n_umap_dim * K) for testing as the test data, that is, the test feature matrix. Subtract the trained mean matrix with the shape of (n_frame, n_umap_dim * K) from the test feature matrix for testing and divide it by the trained standard deviation matrix with the shape of (n_frame, n_umap_dim * K) to obtain the normalized test feature matrix with the shape of (n_frame, n_umap_dim * K). Details are as follows:

[0092] First, load the mean matrix and the standard deviation matrix obtained in the training phase, that is, load the trained normalization module.

[0093] Then, perform mean normalization on the test data, that is, for each element in the test feature matrix, subtract the mean of the corresponding dimension and divide it by the standard deviation of that dimension to obtain the normalized test feature matrix. The dimension of the normalized test feature matrix is still (n_frame, n_umap_dim * K).

[0094] Testing process of the Transformer model:

[0095] First, load the trained Transformer model and use the normalized test feature matrix with the shape of (n_frame, n_umap_dim *K) as the input. This input passes through the fully connected layer with the shape of (n_umap_dim * K, n_umap_dim * K) to output the test Transformer fully connected layer matrix with the shape of (n_frame, n_umap_dim * K).

[0096] Next, send the test Transformer fully connected layer matrix into the Encoder layer of the Transformer model. This layer uses a four-head attention mechanism. Due to its sequence-to-sequence mapping characteristics, the input and output dimensions remain unchanged, and the test Transformer Encoder layer matrix with the shape of (n_frame, n_umap_dim * K) is obtained.

[0097] Then, take the mean of the test Transformer Encoder layer matrix to obtain the test Transformer mean vector with the dimension of (1, n_umap_dim *K). This vector is the comprehensive representation of the input sequence.

[0098] Finally, input the mean vector of the tested Transformer into a fully connected classification layer with dimensions (n_umap_dim * K, 2), and through linear transformation, map it to a two-dimensional space to obtain the classification result, which is normal or abnormal.

[0099] Details are as follows:

[0100] As Figure 1 shown, the Transformer model based on multiple UMAPs is trained and tested.

[0101] During the training process, input the training dataset, perform MFCC feature extraction, K kinds of UMAP feature compression training, normalization module training, and Transformer model training, and finally obtain a trained Transformer model based on K kinds of UMAPs.

[0102] During the testing process, input the test dataset, perform MFCC feature extraction, the trained K kinds of UMAP compressed features, normalize the features with the trained normalization model, and predict with the trained Transformer model, and finally obtain the accuracy of the test set.

[0103] The Transformer model is a prior art and is briefly described as follows:

[0104] The Transformer model is divided into training and testing steps.

[0105] The architecture of the Transformer model is: fully connected layer, Transformer encoder, mean calculation, fully connected classification layer. The Transformer encoder is the Encoder layer, and the mean function is for mean calculation.

[0106] The training steps of the Transformer model: Input a training feature matrix of (n_frame, n_umap_dim * K), train the normalization module, and train the Transformer model.

[0107] The testing steps of the Transformer model: Input a test feature matrix of (n_frame, n_umap_dim * K), normalize the test feature matrix with the trained normalization module, and then input the normalized test feature matrix into the trained Transformer model to obtain the classification result. The classification result includes two results: normal or abnormal for the motor.

[0108] Experimental examples:

[0109] This application uses the publicly available dataset MIMII Dataset to verify the superiority of the method for detecting motors based on multiple UMAP-based Transformer models. The MIMII Dataset is a reliable dataset for the investigation and inspection of faulty industrial machines. It contains the sounds generated by four industrial machines, namely valves, pumps, fans, and slides. Four models (id0, id2, id4, id6) are publicly available for each valve, pump, fan, and slide, and sound data with signal-to-noise ratios of -6dB, 0dB, and 6dB are publicly available for each model. At different signal-to-noise ratios of each model, the corresponding data includes normal sounds and abnormal sounds.

[0110] This application selects the sound of the pump motor with an id0 and a signal-to-noise ratio of -6dB, i.e., more noise, as the experimental data, and divides the training set and the test set in a ratio of 8:2.

[0111] In feature extraction: For MFCC, the n_frame of the feature matrix with a shape of (n_frame, n_mfcc) is 5, and n_mfcc is 64.

[0112] In feature compression: K = 2; K n_neighbors are divided into 10 and 20; n_umap_dim = 32.

[0113] Experimental results are shown as follows:

[0114] The Transformer models without using UMAP-compressed features, the Transformer models using only one type of UMAP, and the Transformer models based on multiple UMAPs are compared respectively.

[0115] The accuracy of the Transformer model without using UMAP-compressed features is 95.2.

[0116] When K = 1 and n_neighbor = 10, the accuracy of the Transformer model using one type of UMAP is 95.5.

[0117] When K = 1 and n_neighbor = 20, the accuracy of the Transformer model using one type of UMAP is 95.4.

[0118] When K = 2; n_neighbor of the first type of UMAP is 10, and n_neighbor of the second type of UMAP is 20, the accuracy of the Transformer model based on multiple UMAPs is 95.7.

[0119] K = 3; the n_neighbor of the first UMAP is 10, the n_neighbor of the second UMAP is 15, and the n_neighbor of the third UMAP is 20. The accuracy of using the Transformer model based on multiple UMAPs is also 95.7. It can be seen that when K is greater than 2, the accuracy does not increase, but the running efficiency of the model decreases. Therefore, the optimal parameter of K is 2.

[0120] When K ≥ 4, the running speed will decrease significantly, which will not be elaborated here.

[0121] It can be seen that in the dataset with relatively high motor sound noise, by compressing features with UMAP to remove redundant features, and using different types of UMAP with K different n_neighbor parameters, the classification of motor sounds with relatively low signal-to-noise ratio can be improved.

[0122] Technical effects:

[0123] It has relatively high accuracy for motor sound data with relatively low signal-to-noise ratio.

[0124] UMAP is used to remove redundant information with more noise.

[0125] At the same time, different types of UMAP with K different n_neighbor parameters are used, so that the compressed MFCC feature matrix can simultaneously maintain the local details and the overall structure of the data, which can improve the accuracy.

[0126] UMAP and Transformer are combined and applied to the classification of motor sound data with relatively high noise.

[0127] Compared with the above embodiments, in step S1, Filterbank filter bank feature extraction can also be used to replace MFCC feature extraction, or Mel Spectrogram feature extraction can also be used to replace MFCC feature extraction.

[0128] Mel Spectrogram is the Mel spectrogram; Mel-Frequency Cepstral Coefficients, abbreviated as MFCC, Mel frequency cepstral coefficients.

[0129] The time-domain signal of the motor is the data manifested by the motor on the time axis, usually generated by the vibration of the motor during operation, and is a time series that changes with time. The time-domain signal of the motor is a sound signal, and the corresponding time-domain signal of the motor is obtained by recording through a microphone, or the time-domain signal of the motor is a vibration signal, and the corresponding time-domain signal of the motor is obtained by a vibration sensor.

[0130] Common time-domain signal acquisition methods include:

[0131] Airborne sound signals obtained by microphone sensors: The pressure changes of sound waves in the air are captured by the microphone and converted into corresponding voltage signals.

[0132] Motor vibration signals obtained by vibration sensors: The mechanical vibrations generated during the operation of the motor are captured by the sensor and converted into voltage signals.

Claims

1. A method for detecting a motor based on a Transformer model of UMAP, characterized in that: The following steps are included: Step S1: extracting features from the time domain signal of the motor to obtain a feature matrix of shape (n_frame, n_mfcc), wherein the time domain signal is a time domain signal of a sound signal or a time domain signal of a vibration signal; Step S2: UMAP feature compression, step S2 includes steps S201 to S204, Step S201: setting K n_neighbor parameters to obtain K different types of UMAPs; Step S202: split the feature matrix with a shape of (n_frame, n_mfcc) obtained in step S1 to obtain n_frame feature vectors with a shape of (n_mfcc, 1); Step S203: Each feature vector of shape (n_mfcc, 1) is independently input to each UMAP, and one UMAP is trained to obtain n_frame trained UMAPs of the same type and K different types of UMAPs, and a total of K*n_frame trained UMAPs are obtained; Step S204: input n_frame feature vectors of shape (n_mfcc, 1) to each trained UMAP, each trained UMAP outputs n_frame feature vectors of shape (n_umap_dim, 1), concatenate n_frame feature vectors of shape (n_umap_dim, 1) to obtain a feature matrix of shape (n_frame, n_umap_dim), K different types of UMAP, obtain K feature matrices of shape (n_frame, n_umap_dim), concatenate K feature matrices of shape (n_frame, n_umap_dim) to obtain a feature matrix of shape (n_frame, n_umap_dim*K); Step S3: normalize the mean of the feature matrix of shape (n_frame, n_umap_dim*K) obtained in step S2 and input it into the Transformer model, and the Transformer model outputs the classification result; The method also includes a training process, which includes: obtaining a training data set, extracting features through step S1, and obtaining a feature matrix with a shape of (n_frame, n_mfcc) for training; obtaining K types of trained UMAPs through steps S201, S202, and S203, and obtaining K*n_frame trained UMAPs in total; obtaining a feature matrix with a shape of (n_frame, n_umap_dim*K) for training through step S204; training a normalization module with the feature matrix with a shape of (n_frame, n_umap_dim*K) obtained in step S204 to obtain a trained normalization module, and training a Transformer model with the normalized features output by the trained normalization module to obtain a trained Transformer model, that is, a trained Transformer model based on K types of UMAPs.

2. The method for detecting a motor using a Transformer model based on UMAP according to claim 1, characterized in that: In the step S1, Mel-frequency cepstral coefficients (MFCC) are used for feature extraction.

3. The method for detecting a motor based on the Transformer model of UMAP according to claim 1, characterized in that: The step of training the normalization module includes calculating the mean and standard deviation of each dimension of the feature matrix of the shape (n_frame, n_umap_dim * K) used for training to obtain a trained mean matrix and a standard deviation matrix of the shape (n_frame, n_umap_dim * K), and subtracting the trained mean matrix from the feature matrix of the shape (n_frame, n_umap_dim * K) used for training and dividing it by the trained standard deviation matrix to obtain a normalized training feature matrix of the shape (n_frame, n_umap_dim * K); the trained mean matrix and standard deviation matrix form a trained normalization module, and the features in the normalized training feature matrix of the shape (n_frame, n_umap_dim * K) are the normalized features output by the trained normalization module; the Transformer model includes a fully connected layer, an Encoder layer, a mean function and a fully connected classification layer connected in sequence, and the step of training the Transformer model includes The normalized training feature matrix of n_umap_dim * K) is input into the Transformer model, and the Transformer model outputs the classification result. Back propagation is used in training to optimize the Transformer model parameters according to the loss value calculated by the loss function. After multiple iterations, the trained Transformer model is obtained.

4. The method for detecting a motor based on the Transformer model of UMAP according to claim 3 is characterized in that: The test process is also included. The test process is as follows: obtaining a test data set, extracting features through step S1, and obtaining a feature matrix with a shape of (n_frame, n_mfcc) for testing; After step S202, K*n_frame trained UMAPs and step S204, a feature matrix with a shape of (n_frame, n_umap_dim*K) for testing is obtained; the feature matrix with a shape of (n_frame, n_umap_dim*K) for testing obtained in step S204 is predicted by the trained normalized model and the trained Transformer model in turn to obtain the classification result of the test set.

5. The method for detecting a motor using a Transformer model based on UMAP according to claim 4, characterized in that: The step of predicting by the trained normalized model and the trained Transformer model in sequence includes obtaining a normalized test feature matrix with a shape of (n_frame, n_umap_dim * K) by subtracting the trained mean matrix from the feature matrix with a shape of (n_frame, n_umap_dim * K) for testing and dividing it by the trained standard deviation matrix, and inputting the normalized test feature matrix with a shape of (n_frame, n_umap_dim * K) into the trained Transformer model to obtain a classification result.

6. The method for detecting a motor based on the Transformer model of UMAP according to claim 1, characterized in that: In step S201, K=2, or K=3.

7. The method for detecting a motor based on the Transformer model of UMAP according to claim 6, characterized in that: In the step S201, when K=2, n_neighbor of the first UMAP=10, and n_neighbor of the second UMAP=20; when K=3, n_neighbor of the first UMAP=10, n_neighbor of the second UMAP=15, and n_neighbor of the third UMAP=20.

8. The method for detecting a motor based on the Transformer model of UMAP according to claim 1, characterized in that: In the step S1, filterbank filter bank feature extraction is performed, or Mel Spectrogram feature extraction is performed.

Citation Information

Patent Citations

  • Motor imagery electroencephalogram signal classification system based on TimesNet and convolutional neural network

    CN118094317A

  • Data processing device, data processing method, program, and model

    JP2021089483A