Method for Detecting Electric Motors Based on a Hybrid Autoencoder Ensemble Learning Model

By building a hybrid autoencoder integrated learning model, combining multiple feature extraction methods and autoencoder AU architecture, the problem of low motor abnormal detection accuracy is solved and higher detection accuracy is achieved.

CN119476368BActive Publication Date: 2025-06-24SUZHOU ACOUSTIC IND TECH RES INST CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411591732.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-08
Publication Date
2025-06-24
Estimated Expiration
2044-11-08

AI Technical Summary

Technical Problem

In the prior art, the autoencoder AU architecture is difficult to determine and the feature extraction method is difficult to determine, resulting in low motor abnormal detection accuracy.

Method used

Using a method based on a hybrid autoencoder integrated learning model, an integrated model including multiple audio feature extraction methods and multiple autoencoder AU architectures is constructed. By training and combining different basic models, the detection accuracy is improved.

Benefits of technology

Through the integrated learning model, the randomness of the model is reduced, the advantages of each basic model are exerted, and the accuracy of motor abnormality detection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119476368B_ABST
    Figure CN119476368B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting motors based on a hybrid autoencoder ensemble learning model, which relates to the technical field of motor testing; it includes: Step S1: Construct and obtain a hybrid autoencoder ensemble learning model, which includes m*n types of base models, and one type of base model includes one audio feature extraction method and one autoencoder AU; Step S2: Train the autoencoder AU of each base model to obtain a trained base model, and input the normal audio data of the motors in the training set into each trained base model to train and obtain the best base model; Step S3: Input the normal and abnormal audio data of the motors in the test set into each best base model to obtain the abnormal scores of the test set of each best base model, and average all the abnormal scores to obtain the abnormal score of the test set; the greater the difference between the base models, the higher the accuracy of the base models, and the higher the accuracy of the hybrid autoencoder ensemble learning model, thereby making the motor abnormal detection accuracy higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of motor testing, and particularly relates to a method for detecting motors based on a hybrid autoencoder ensemble learning model. Background Art

[0002] The process of machine recognition of motor sounds is as follows: input the sound of the motor into a feature extraction method, then use the feature extraction method to extract the motor sound features, and finally input the motor sound features into a machine learning model for classification.

[0003] The existing technical solutions are as follows:

[0004] The first existing technical solution: Yang Tianjia. Research on Abnormal Detection and Time-Series Prediction Algorithm of Forklift AGV DC Motor Based on Hybrid Samples [D]. Qilu University of Technology, 2024. DOI: 10.27278 / d.cnki.gsdqc.2024.000872.

[0005] The second existing technical solution: Tao Yonghui, Wang Yong. Abnormal Data Detection of Wind Turbines Based on Improved K-means [J]. Foreign Electronic Measurement Technology, 2023, 42(04): 141-148. DOI: 10.19652 / j.cnki.femt.2304738.

[0006] The third existing technical solution: Liu Fawei. Abnormal Detection of Rotating Motors Based on CAE-OC-SVM [J]. Industrial Control Computer, 2024, 37(03): 133-135.

[0007] The fourth existing technical solution: Agrawal V K, Maurya S S. Unsupervised detection of anomalous sounds for machine condition monitoring [J]. DCASE 2020 Challenge, Tech. Rep, 2020.

[0008] The fifth existing technical solution: Yuan Xuefeng, Ma Chenglong, Chen Shihe. Research on Equipment Parameter Early Warning Based on GMM and NSET Optimization Algorithm [J]. Control Engineering, 2022, 29(06): 1058-1064. DOI: 10.14107 / j.cnki.kzgc.20210731.

[0009] Sixth prior art solution: Breunig M M, Kriegel H P, Ng R T, et al. Proceedings of the 2000 ACM SIGMOD international conference on Management of data[C] / / Proceedings of the 2000 Acm Sigmod International Conference on Management of Data. 2000: 93-104。

[0010] Seventh prior art solution: Jiang W, Hong Y, Zhou B, et al. A GAN-based anomaly detection approach for imbalanced industrial time series[J]. IEEE Access, 2019, 7: 143608-143619。

[0011] Eighth prior art solution: Marchi E, Vesperini F, Weninger F, et al. Non-linear prediction with LSTM recurrent neural networks for acoustic novelty detection[C] / / 2015 International Joint Conference on Neural Networks (IJCNN). IEEE, 2015: 1-7。

[0012] Ninth prior art solution: Qi Y, Shen C, Wang D, et al. Stacked sparse autoencoder-based deep network for fault diagnosis of rotating machinery[J]. Ieee Access, 2017, 5: 15066-15079。

[0013] Tenth prior art solution: Principia E, Rossetti D, Squartini S, et al. Unsupervised electric motor fault detection by using deep autoencoders[J]. IEEE / CAA Journal of Automatica Sinica, 2019, 6(2): 441-451。

[0014] Eleventh prior art solution: Atal B S, Hanauer S L. Speech analysis and synthesis by linear prediction of the speech wave[J]. The journal of the acoustical society of America, 1971, 50(2B): 637-655。

[0015] Twelfth prior art solution: Davis S, Mermelstein P. Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences[J]. IEEE transactions on acoustics, speech, and signal processing, 1980, 28(4): 357-366。

[0016] Thirteenth prior art solution: Elmasry W, Wadi M. Edla-efds: A novel ensemble deep learning approach for electrical fault detection systems[J]. Electric Power Systems Research, 2022, 207: 107834。

[0017] Fourteenth prior art solution: Dong X, Yu Z, Cao W, et al. A survey on ensemble learning[J]. Frontiers of Computer Science, 2020, 14: 241-258。

[0018] In the prior art solutions, during the training process, autoencoder AU models with different numbers of neurons in different hidden layers have certain differences in their architectures. An AU architecture with a larger number of neurons will not be underfitted, but is prone to overfitting some data. An autoencoder AU architecture with a smaller number of neurons will not overfit, but is prone to underfitting some data. Therefore, autoencoder AU models with different architectures have their own advantages and disadvantages. At the same time, since the parameters of the autoencoder AU are randomly initialized during training, the training of the autoencoder AU is random. Even if the same autoencoder AU architecture is trained multiple times, there are still differences in quality.

[0019] The authorized announcement number is CN118226254B, and the name is a method for detecting motors based on an autoencoder ensemble learning model. To address the randomness problem of the autoencoder AU, a single autoencoder AU architecture and a single feature extraction method are adopted. Ensemble learning can reduce the randomness of the model and leverage the advantages of each base model, thereby improving the overall accuracy. However, in the detection of abnormal motor sounds, ensemble learning is composed of multiple base models. One type of base model represents a combination of an audio feature extraction method and a machine learning model architecture. One type of base model can be trained multiple times to generate multiple base models. Therefore, it is difficult to determine the autoencoder AU architecture, the training of the autoencoder AU is random, and it is difficult to determine the feature extraction method, resulting in a low accuracy in motor anomaly detection. Summary of the Invention

[0020] The present invention provides a method for detecting motors based on a hybrid autoencoder ensemble learning model to solve the technical problem of low accuracy in motor anomaly detection.

[0021] To solve the above technical problem, the technical solutions adopted by the present invention are as follows:

[0022] A method for detecting motors based on a hybrid autoencoder ensemble learning model includes the following steps:

[0023] Step S1: Construct a hybrid autoencoder ensemble learning model. The hybrid autoencoder ensemble learning model includes m*n types of base models, where m is the number of types of audio feature extraction methods and n is the number of types of autoencoder AU architectures. One type of base model includes an audio feature extraction method and an autoencoder AU;

[0024] Step S2: Train the autoencoder AU of each base model in the hybrid autoencoder ensemble learning model to obtain trained base models, and input the normal audio data of the motor in the training set into each trained base model to train and obtain each optimal base model;

[0025] Step S3: Input the normal and abnormal audio data of the motors in the test set into each optimal base model in the hybrid autoencoder ensemble learning model to obtain the abnormal scores of the test set for each optimal base model, and average all the abnormal scores of the test set to obtain the abnormal score of the test set.

[0026] A further technical solution lies in that: in the step S1, the steps of constructing the hybrid autoencoder ensemble learning model include the following steps.

[0027] Step S101: Construct and obtain the autoencoder AU.

[0028] n = 2. The autoencoder AU includes the first autoencoder modelA and the second autoencoder modelB, a total of two types of autoencoders. The first autoencoder modelA and the second autoencoder modelB are a type of autoencoder, namely the basic autoencoder.

[0029] Step S102: Construct and obtain the feature extraction method.

[0030] m = 2. The feature extraction methods include Mel Frequency Cepstral Coefficients MFCC and Frequency Filter Bank FB, a total of two feature extraction methods.

[0031] The base models include the MFCC_modelA base model, the FB_modelA base model, the MFCC_modelB base model, and the FB_modelB base model, a total of four base models. The four base models form the hybrid autoencoder ensemble learning model and are marked as the mul_fea_model base model.

[0032] A further technical solution lies in that: in the step S102, the following steps are further included.

[0033] The steps of the frequency filter bank FB extracting features include that the sound signal of the motor is sequentially subjected to short-time Fourier transform, processed by the frequency filter bank FB, logarithmic transform, and constructing a feature vector to obtain a frequency band feature vector.

[0034] The steps of the Mel Frequency Cepstral Coefficients MFCC extracting features include that the sound signal of the motor is sequentially subjected to short-time Fourier transform, processed by the Mel filter bank MF, logarithmic transform, discrete cosine transform, and constructing a feature vector to obtain a Mel Frequency Cepstral Coefficients feature vector.

[0035] A further technical solution lies in that: in the step S2, the steps of training the autoencoder AU in each base model of the hybrid autoencoder ensemble learning model to obtain the trained base models include obtaining the gradient of the autoencoder AU based on the loss function, optimizing the gradient with the optimization function Adam, and backpropagating to obtain the weights and thresholds of each autoencoder AU in the hybrid autoencoder ensemble learning model until multiple trained autoencoders AU are obtained.

[0036] The steps of inputting the audio data of the normal motor in the training set into each trained base model to obtain each optimal base model include inputting the audio data of the normal motor in the training set into each trained base model. The audio data of the normal motor is extracted by the feature extraction method in the trained base model to obtain the audio features of the normal motor. The audio features of the normal motor are used as the training input features and input into the autoencoder AU in the trained base model to obtain the training output features of the autoencoder AU in the trained base model. Calculate the mean square error between the training input features and the training output features of the autoencoder AU in each trained base model as the training error, obtain the reciprocal of each training error, normalize it and obtain the integrated weight, obtain each optimal autoencoder AU, and then obtain each optimal base model.

[0037] A further technical solution lies in that: in the step S2,

[0038] The loss function is:

[0039] (1)

[0040] In formula (1): is the loss function; p is the batch size; q is the dimension of the input features; is the i-th feature of the j-th input feature vector in the batch; is the i-th feature of the j-th output feature vector in the batch; j ∈ [1, p], i ∈ [1, q];

[0041] The training error function is:

[0042] (2)

[0043] In formula (2): is the training error function; z is the number of training samples; q is the dimension of the input features; i is the i-th dimension of the feature vector, and j is the feature vector of the j-th training data; is the i-th feature of the j-th input feature vector in the training set; is the i-th feature of the j-th reconstructed feature vector in the training set; j ∈ [1, z], i ∈ [1, q].

[0044] A further technical solution lies in that: in the step S3, the steps of obtaining the test set anomaly scores of each optimal base model include inputting the normal and abnormal audio data of the motor in the test set into each optimal base model in the hybrid autoencoder ensemble learning model. The normal and abnormal audio data of the motor are extracted by the feature extraction method in the optimal base model to obtain the normal and abnormal audio features of the motor. The normal and abnormal audio features of the motor are used as the test input features and input into the autoencoder AU in the optimal base model to obtain the test output features of the autoencoder AU in the optimal base model, that is, the reconstructed features. Calculate the mean square error between the test input features and the test output features of the autoencoder AU in each optimal base model as the reconstruction error, and the reconstruction error is the test set anomaly score.

[0045] The beneficial effects produced by adopting the above technical solutions are as follows:

[0046] A method for detecting a motor based on a hybrid autoencoder ensemble learning model includes: Step S1: Construct and obtain a hybrid autoencoder ensemble learning model, which includes m*n types of base models, where m is the number of types of audio feature extraction methods and n is the number of types of autoencoder AU architectures. One type of base model includes one type of audio feature extraction method and one type of autoencoder AU; Step S2: Train the autoencoder AU of each base model to obtain the trained base models, and input the normal audio data of the motor in the training set into each trained base model to train and obtain each optimal base model; Step S3: Input the normal and abnormal audio data of the motor in the test set into each optimal base model in the hybrid autoencoder ensemble learning model to obtain the test set anomaly scores of each optimal base model, and average all the test set anomaly scores to obtain the anomaly score of the test set. The greater the difference between the base models, the higher the accuracy of the base models, and the higher the accuracy of the hybrid autoencoder ensemble learning model, thereby making the motor anomaly detection accuracy higher.

[0047] See the description in the specific implementation part for details. Description of the Drawings

[0048] Figure 1 is the flowchart of the present invention;

[0049] Figure 2 is the data flow diagram of the present invention;

[0050] Figure 3 is the data flow diagram for obtaining the trained autoencoder AU;

[0051] Figure 4 is the data flow diagram for obtaining the optimal base model. Specific Implementation Manner

[0052] The authorized announcement number is CN118226254B, and the name is a method for detecting motors based on an integrated learning model of autoencoders, hereinafter referred to as the original technical solution. The technical solution of this application is a technical solution obtained by further improving the original technical solution, hereinafter referred to as the improved technical solution. The original technical solution uses an audio feature extraction method and an autoencoder AU architecture. The improved technical solution uses the integration of multiple feature extraction methods and multiple autoencoder AU architectures, solving the technical problem that it is difficult to determine the autoencoder AU architecture, solving the technical problem that it is difficult to determine the feature extraction method, and thus solving the technical problem of low accuracy in motor anomaly detection.

[0053] The autoencoder AU is a commonly used unsupervised learning model that can be used in many fields such as anomaly detection and image generation. The AU maps features based on the weights, thresholds, and activation functions of neurons and has relatively strong fitting and generalization capabilities. Since the distribution of abnormal data is different from that of normal data, the AU is trained with normal data. Therefore, the AU has a small reconstruction error, also known as the anomaly score, for normal data and a large reconstruction error for abnormal data.

[0054] Due to the difficulty in determining the autoencoder AU architecture and the feature extraction method, the accuracy of motor anomaly detection is reduced. Ensemble learning is a method of combining multiple base models to improve the performance of the overall model. By combining the outputs of multiple models, the advantages and disadvantages of each base model are fully utilized. Therefore, ensemble learning can obtain better results than a single model. The accuracy of ensemble learning depends on the accuracy of the base models and the differences between the base models. The greater the differences between the base models and the higher the accuracy of the base models, the higher the accuracy of ensemble learning. Different motor sound feature extraction methods and AUs with different architectures can be combined into different types of base models with relatively large differences. By training each type of base model multiple times, the best base model in each type of base model can be selected, and then the best base models in each type of base model are integrated to form an improved method for detecting motors based on a hybrid autoencoder ensemble learning model.

[0055] By making full use of the advantages and disadvantages of various autoencoder AU architectures and different feature extraction methods, a method for detecting motors based on a hybrid autoencoder ensemble learning model is proposed. This method trains each type of base model multiple times, selects the best base model in each type of base model, and then combines the best base models in each type of base model into an ensemble learning model.

[0056] Common feature extraction methods include the Filter bank (FB for short), the Mel Frequency Cepstral Coefficients (MFCC for short). Both the Filter bank (FB) and the Mel Frequency Cepstral Coefficients (MFCC) fit well with the human ear's perception. However, MFCC has better robustness, and FB retains more detailed information for each frequency band. Therefore, each has its own advantages and disadvantages.

[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and in no way constitutes a limitation on the present application and its application or use. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0058] Many specific details are set forth in the following description to facilitate a thorough understanding of the present application. However, the present application may be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.

[0059] Embodiment 1:

[0060] As Figure 1 shown, the present invention discloses a method for detecting a motor based on a hybrid autoencoder ensemble learning model, which includes the following steps:

[0061] Step S1: Construct and obtain a hybrid autoencoder ensemble learning model.

[0062] Construct and obtain a hybrid autoencoder ensemble learning model. The hybrid autoencoder ensemble learning model includes m * n basic models, where m is the number of types of audio feature extraction methods, and n is the number of types of the autoencoder AU architecture. The basic model includes a feature extraction method and an autoencoder AU.

[0063] As Figure 2As shown in the figure, the hybrid autoencoder ensemble learning model proposed in this application is an improved multi-feature and multi-model autoencoder ensemble learning model. It adopts m feature extraction methods and n autoencoder AU architectures, and constructs a total of m * n base models. One base model includes one audio feature extraction method and one autoencoder AU. Different autoencoder AU architectures and different feature extraction methods are constructed. One audio feature extraction method and one architecture of the autoencoder AU form one base model, thus combining into a variety of base models with different and significant differences.

[0064] The accuracy of the hybrid autoencoder ensemble learning model depends on the accuracy of the base models and the differences between the base models. The greater the differences between the base models, the higher the accuracy of the base models, and the higher the accuracy of the hybrid autoencoder ensemble learning model.

[0065] Step S101: Construct and obtain the autoencoder AU.

[0066] Construct and obtain n autoencoders AU, where n = 2.

[0067] The autoencoder AU is a type of autoencoder model, with a quantity of two, including the first autoencoder model A and the second autoencoder model B.

[0068] Step S102: Construct and obtain the feature extraction method.

[0069] Construct and obtain m feature extraction methods, where m = 2.

[0070] The feature extraction methods include Mel Frequency Cepstral Coefficients MFCC and Frequency Filter Bank FB, a total of two.

[0071] The steps for the Frequency Filter Bank FB to extract features include that the sound signal of the motor is successively subjected to short-time Fourier transform, processed by the Frequency Filter Bank FB, logarithmic transform, and construction of a feature vector to obtain a frequency band feature vector, and the frequency band feature vector is an audio feature.

[0072] The steps for Mel Frequency Cepstral Coefficients MFCC to extract features include that the sound signal of the motor is successively subjected to short-time Fourier transform, processed by the Mel filter bank MF, logarithmic transform, discrete cosine transform, and construction of a feature vector to obtain a Mel Frequency Cepstral Coefficient feature vector, and the Mel Frequency Cepstral Coefficient feature vector is an audio feature.

[0073] Details are as follows:

[0074] Common machine learning motor feature extraction methods include: Filter bank, abbreviated as FB, which is an audio feature extraction method, and Mel Frequency Cepstral Coefficients, abbreviated as MFCC, which is also an audio feature extraction method.

[0075] FB is a set of filters used to decompose a signal into different frequency bands and is a key step in signal processing, especially in audio and speech processing. In speech signal processing, filter banks are usually used to extract meaningful frequency band information from frequency domain signals.

[0076] Step S1021: The frequency filter bank FB extracts features.

[0077] The acoustic signal of the motor is successively processed by short-time Fourier transform, frequency filter bank FB, logarithmic transform, and constructing a feature vector to obtain a frequency band feature vector, which is detailed as follows.

[0078] Step S10211: Short-time Fourier transform.

[0079] The acoustic signal of the motor is framed, windowed, and Fourier-transformed to obtain a spectrogram.

[0080] Step S10212: Frequency filter bank FB processing.

[0081] The spectrogram is filtered by the frequency filter bank FB to obtain the spectrogram of the corresponding frequency band in the spectrum.

[0082] To better conform to the human ear's hearing, the frequency filter bank FB is used to filter the spectrogram to obtain the sum of the energies of the corresponding frequency bands in the spectrum.

[0083] Step S10213: Logarithmic transform.

[0084] The spectrogram obtained by filtering with the frequency filter bank in step S10212 is logarithmically transformed to obtain the logarithmically transformed spectrogram.

[0085] After calculating the energy output by each filter bank, these energies are logarithmically transformed. This can compress the dynamic range of the audio signal, reducing the gap between large and small energies and further simulating the human ear's perception of loudness.

[0086] Step S10214: Constructing a feature vector.

[0087] The vectors of each frame in the logarithmically transformed spectrogram obtained in step S10213 are concatenated into a total frequency band feature vector.

[0088] Since the spectrogram after taking the logarithm is composed of multiple frame vectors, in order to conform to the input of the autoencoder AU, the vectors of each frame are concatenated into a total feature vector.

[0089] Mel Frequency Cepstral Coefficients (MFCC) is a feature extraction method widely used in the fields of speech recognition, audio processing, etc. It extracts useful features that can represent audio signals by mimicking the human ear's perception of sounds at different frequencies. Mel Frequency Cepstral Coefficients (MFCC) perform an additional discrete cosine transform operation based on the frequency filter bank (FB).

[0090] Step S1022: Extract features using Mel Frequency Cepstral Coefficients (MFCC).

[0091] The acoustic signal of the motor is successively subjected to short-time Fourier transform, Mel filter bank (MF) processing, logarithmic transform, discrete cosine transform, and construction of a feature vector to obtain the Mel Frequency Cepstral Coefficient feature vector, which is detailed as follows.

[0092] Step S10221: Short-time Fourier transform.

[0093] The acoustic signal of the motor is framed, windowed, and Fourier-transformed to obtain a spectrogram.

[0094] Step S10222: Mel filter bank (MF) processing.

[0095] The spectrogram is filtered through a Mel filter bank to obtain the spectrogram of the corresponding frequency band in the spectrum.

[0096] To better conform to the human ear's hearing, the Mel filter bank is used to filter the spectrogram to obtain the sum of the energies of the corresponding frequency bands in the spectrum.

[0097] Step S10223: Logarithmic transform.

[0098] The spectrogram after Mel filter bank filtering in step S10222 is logarithmically transformed to obtain the logarithmically transformed spectrogram.

[0099] After calculating the energy output by each filter bank, a logarithmic transform is performed on these energies. This can compress the dynamic range of the audio signal, reducing the gap between large and small energies and further simulating the human ear's perception of loudness.

[0100] Step S10224: Discrete cosine transform.

[0101] The spectrogram obtained in step S10223 is discretely cosine-transformed to obtain cepstral coefficients.

[0102] Perform a discrete cosine transform (DCT) on the logarithmic energy spectrum to convert it into cepstral coefficients.

[0103] Step S10225: Construct a feature vector.

[0104] Concatenate the vectors of each frame in the cepstral coefficients of step S10224 into a total Mel-frequency cepstral coefficient feature vector.

[0105] Since the cepstral coefficients of the discrete cosine transform are composed of multiple frame vectors, in order to conform to the input of the autoencoder AU, the vectors of each frame are concatenated into a total feature vector.

[0106] Both the frequency filter bank FB and the Mel-frequency cepstral coefficients MFCC fit well with the human ear's perception, but MFCC has better robustness, and FB retains more detailed information in each frequency band. Therefore, each has its own advantages and disadvantages.

[0107] Step S2: Train to obtain the best base model.

[0108] For each base model's autoencoder AU in the hybrid autoencoder ensemble learning model obtained by training in step S1, obtain the trained base model. Input the audio data of normal motors in the training set into each trained base model to train and obtain each best base model.

[0109] The steps to train the autoencoder AU of each base model in the hybrid autoencoder ensemble learning model to obtain the trained base model include: based on the loss function, use the gradient descent method to obtain the gradient of the autoencoder AU, use the optimization function Adam to optimize the gradient, and backpropagate to obtain the weights and thresholds of each autoencoder AU in the hybrid autoencoder ensemble learning model until multiple trained autoencoders AU are obtained.

[0110] The steps to input the audio data of normal motors in the training set into each trained base model to train and obtain each best base model include: input the audio data of normal motors in the training set into each trained base model. The audio features of normal motors are extracted through the feature extraction method in the trained base model. The audio features of normal motors are used as the training input features and input into the autoencoder AU in the trained base model to obtain the training output features of the autoencoder AU in the trained base model. Calculate the mean square error between the training input features and the training output features of the autoencoder AU in each trained base model as the training error. Obtain the reciprocal of each training error, normalize it, and obtain the integrated weights. Obtain each best autoencoder AU, and then obtain each best base model.

[0111] Details are as follows:

[0112] Training process of the autoencoder AU: Construct the architecture of the autoencoder AU and determine the loss function.

[0113] As Figure 3As shown in the figure, the architecture of the autoencoder AU. The feature vectors extracted by the feature extraction method are input into the encoder layer to obtain feature embeddings, and then the decoder layer is used to parse the feature embeddings to obtain the reconstructed feature vectors, where the input feature vectors and the reconstructed feature vectors have the same dimension. Different numbers of neurons in the encoder layer and the decoder layer determine different autoencoder AU architectures.

[0114] The loss function can be expressed as:

[0115] (1)

[0116] In formula (1): is the loss function; p is the batch size; q is the dimension of the input features; is the i-th feature of the j-th input feature vector in the batch; is the i-th feature of the j-th output feature vector in the batch; j ∈ [1, p], i ∈ [1, q].

[0117] As Figure 2 shown, there are a total of m * n base models. One base model can be trained multiple times to become multiple base models. K is the number of times a base model is trained, and this base model generates K base models. The same base model is trained K times in total, and then the best base model is selected from the K base models of each base model. Then, the best base models of each base model are integrated, and there are a total of m * n best base models.

[0118] As Figure 4 shown, for one base model, the process of training it K times and then selecting the best base model according to the reconstruction error of the training set. The normal data in the training set undergoes feature extraction to obtain the input feature vectors. The input feature vectors are input into the autoencoder AU for training, and then the reconstructed feature vectors are obtained. The mean square error (MSE) is calculated between the input feature vectors and the reconstructed feature vectors, and the training error of the autoencoder AU is obtained. Finally, the trained autoencoder AU with the smallest training error is selected.

[0119] Figure 2 In, the figure of m * n base models, each base model is trained K times. Figure 4 In, the figure of one model trained K times and selecting the best base model; Figure 4 is Figure 2 a part of.

[0120] The reconstruction error is calculated based on the mean square error between the input features and the output features of the autoencoder AU in the training set. This part is the same as the method steps in the original technical solution. The reconstruction error is shown in formula (2):

[0121] (2)

[0122] In formula (2): is the training error function; z is the number of training samples; q is the dimension of the input features; i is the i-th dimension of the feature vector, and j is the feature vector of the j-th training data; is the i-th feature of the j-th input feature vector in the training set; is the i-th feature of the j-th reconstructed feature vector in the training set; j ∈ [1, z], i ∈ [1, q].

[0123] Step S3: Obtain the anomaly score through testing.

[0124] Input the audio data of normal and abnormal motors in the test set into each optimal base model in the hybrid autoencoder ensemble learning model to obtain the test set anomaly scores of each optimal base model, and average all the test set anomaly scores to obtain the anomaly score.

[0125] The steps to obtain the test set anomaly scores of each optimal base model include inputting the audio data of normal and abnormal motors in the test set into each optimal base model in the hybrid autoencoder ensemble learning model. The audio features of normal and abnormal motors are extracted through the feature extraction method in the optimal base model to obtain the audio features of normal and abnormal motors. The audio features of normal and abnormal motors are used as the test input features and input into the autoencoder AU in the optimal base model to obtain the test output features of the autoencoder AU in the optimal base model, that is, the reconstructed features. Calculate the mean square error between the test input features and the test output features of the autoencoder AU in each optimal base model and use it as the reconstruction error. The reconstruction error is the test set anomaly score.

[0126] Details are as follows:

[0127] There are m * n types of base models, and each type of base model is trained K times. The optimal base model in each type of base model is selected through the reconstruction error of the training set. Input the data of the test set into the optimal m * n types of base models to obtain the reconstruction errors of m * n test sets. Take the average of the reconstruction errors of m * n test sets, which is the anomaly score of the test set in the technical solution of this application.

[0128] The area under the curve AUC is a machine learning metric that can be obtained only by giving the labels of the test set and the anomaly scores predicted by the base model.

[0129] When training each base model, it needs to be trained multiple times. The normal data in the training set is extracted with features by the feature extraction method under this base model, and then input into the autoencoder AU architecture of this base model for training. According to the training reconstruction errors of multiple base models, the best base model in this base model is determined.

[0130] When testing, the average value of the test reconstruction errors of the best base models in each base model is taken to obtain the anomaly scores of the test set.

[0131] The technical solution of this application uses the publicly available dataset MIMII Dataset to verify the superiority of the improved method for detecting motors based on the hybrid autoencoder ensemble learning model. MIMII Dataset is a reliable dataset for fault investigation and inspection of industrial machines. It contains the sounds generated by four industrial machines, namely valves, pumps, fans, and sliding rails.

[0132] The technical solution of this application selects the motor sound of a fan with a signal-to-noise ratio of 6 db as the experimental data. The normal data is cut at a ratio of 8:2 to obtain the training set and the test set, and the abnormal data is placed in the test set.

[0133] The specific parameters of the improved autoencoder ensemble learning model based on multiple features and multiple models are as follows:

[0134] m = 2, there are two feature extraction methods, and the feature extraction methods include Mel Frequency Cepstral Coefficients (MFCC) and Frequency Filter Banks (FB); n = 2, there are two autoencoders AU. The autoencoder AU is a type of autoencoder model, and the number is two, including the first autoencoder model A and the second autoencoder model B, forming a hybrid autoencoder ensemble learning model, that is, the mul_fea_model base model.

[0135] Model A and model B are an autoencoder AU architecture.

[0136] The mul_fea_model base model includes the MFCC_modelA base model, the FB_modelA base model, the MFCC_modelB base model, and the FB_modelB base model. The i_mul_fea_model includes the best MFCC_modelA base model, the FB_modelA base model, the MFCC_modelB base model, and the FB_modelB base model.

[0137] 1. Feature extraction methods: Mel Frequency Cepstral Coefficients (MFCC) and Frequency Filter Banks (FB).

[0138] 2. Autoencoder AU architecture: In the autoencoder AU, different numbers of neurons result in different types and architectures of the autoencoder AU.

[0139] 3. Base model: An audio feature extraction method + an autoencoder AU constitute a base model. However, a base model can be trained multiple times, and each training results in a base model.

[0140] 4. MFCC_modelA base model: A base model composed of the MFCC feature extraction method + the modelA AU architecture. However, the base model can be trained multiple times, so MFCC_modelA can contain multiple base models, but they are all of the same type of base model.

[0141] 5. FB_modelA base model: A base model composed of the FB feature extraction method + the modelA AU architecture.

[0142] 6. MFCC_modelB base model: A base model composed of the MFCC feature extraction method + the modelB AU architecture.

[0143] 7. FB_modelB base model: A base model composed of the FB feature extraction method + the modelB AU architecture.

[0144] 8. mul_fea_model base model: It includes four types of base models: the MFCC feature extraction method + the modelA AU architecture; the FB feature extraction method + the modelA AU architecture; the MFCC feature extraction method + the modelB AU architecture; the FB feature extraction method + the modelB AU architecture. Each type of base model can be trained multiple times.

[0145] 9. i_mul_fea_model base model: It includes four types of base models. After each type of base model is trained, the best base model in this type of base model is selected.

[0146] 10. m is the number of types of audio feature extraction methods; n is the number of types of autoencoder AU architectures; so there are a total of m * n types of base models; however, a base model can be trained multiple times to become multiple base models. K is the number of times a base model is trained, and this base model generates K base models.

[0147] In step S101, different autoencoder AU architectures are constructed: The first autoencoder modelA and the second autoencoder modelB of the same category are constructed, a total of two types of autoencoder AUs.

[0148] AUs with different numbers of layers and different numbers of neurons can form different architectures. The neuron topology of the first autoencoder modelA is as follows: the input dimension is 320; there are 5 hidden layers in the middle, and the dimensions of the hidden layers are 64, 64, 8, 64, and 64 respectively; the output dimension is 320. The neuron topology of the second autoencoder modelB is as follows: the input dimension is 320; there are 5 hidden layers in the middle, and the dimensions of the hidden layers are 32, 32, 8, 32, and 32 respectively; the output dimension is 320.

[0149] The type of the autoencoder AU used in Example 1 is a basic autoencoder. Both the first autoencoder modelA and the second autoencoder modelB are basic autoencoders. The difference between the first autoencoder modelA and the second autoencoder modelB lies in the specific data of the dimensions of the hidden layers. Therefore, the first autoencoder modelA and the second autoencoder modelB are two types of autoencoders.

[0150] In step S102, different feature extraction methods are constructed: the technical solution of this application adopts two feature extraction methods, FB and MFCC. FB uses 1024 points as the frame length, 512 points as the step length, and the number of Mel bins is 64. MFCC uses 1024 points as the frame length, 512 points as the step length, and the number of MFCC coefficients is 64. Since the collected signal is a steady-state signal, the continuous 5 frames of the features extracted by FB and MFCC are stretched into a feature vector to facilitate the input of the autoencoder AU model. Finally, the dimensions of the vectors of FB and MFCC are both 64 multiplied by 5, which is 320 dimensions.

[0151] In step S2, the best trained base models are selected: the two feature extraction methods, FB and MFCC, and the two AU model architectures, modelA and modelB, form four base models, namely MFCC_modelA, FB_modelA, MFCC_modelB, and FB_modelB. Each base model is trained three times repeatedly. Using Adam as the optimization function, the mean squared error as the loss function, the batch size is 64, and the base models are trained with early stopping as the iteration stopping condition. The best trained base models are selected according to the training reconstruction error.

[0152] In step S3, the best base models are integrated: the test set is input into the best base models to obtain the reconstruction errors of the test sets of each base model, and the average reconstruction error is the anomaly score.

[0153] Experimental result demonstration: For the AU randomness problem in CN118226254B, a single autoencoder AU architecture and a single feature extraction method are adopted. There are technical problems that it is difficult to determine the autoencoder AU architecture, the training of the autoencoder AU has randomness, and it is difficult to determine the feature extraction method, resulting in low accuracy of motor anomaly detection. The technical solution of this application uses the integration of multiple feature extraction methods and multiple autoencoder AU architectures, which not only solves the problem of difficult determination of the autoencoder AU architecture, but also solves the problem of difficult determination of the feature extraction method, and further solves the problem of low accuracy of motor anomaly detection.

[0154] To demonstrate the superiority of the proposed multi-feature and multi-model autoencoder ensemble learning model for motor detection, the following experiments were conducted.

[0155] A total of four base models were defined, namely:

[0156] The base model MFCC_modelA composed of MFCC as the feature extraction method and the AU architecture of modelA;

[0157] The base model FB_modelA composed of FB as the feature extraction method and the AU architecture of modelA;

[0158] The base model MFCC_modelB composed of MFCC as the feature extraction method and the AU architecture of modelB;

[0159] The base model FB_modelB composed of FB as the feature extraction method and the AU architecture of modelB.

[0160] A total of six ensemble learning methods were defined, namely:

[0161] MFCC_modelA ensemble learning: Only train the MFCC_modelA base model three times to obtain three MFCC_modelA base models. Then, for the anomaly score of the test set, input the test set into the three trained MFCC_modelA base models, and take the average of the anomaly scores, which is the anomaly score of the test set for MFCC_modelA ensemble learning.

[0162] FB_modelA ensemble learning: Only train the FB_modelA base model three times to obtain three FB_modelA base models. Then, for the anomaly score of the test set, input the test set into the three trained FB_modelA base models, and take the average of the anomaly scores, which is the anomaly score of the test set for FB_modelA ensemble learning.

[0163] MFCC_modelB Ensemble Learning: Only use one base model MFCC_modelB to train three times to obtain three base models of MFCC_modelB. Then, the anomaly score of the test set is obtained by inputting the test set into the three trained base models of MFCC_modelB and taking the average of the anomaly scores, which is the anomaly score of the test set for MFCC_modelB ensemble learning.

[0164] FB_modelB Ensemble Learning: Only use one base model FB_modelB to train three times to obtain three base models of FB_modelB. Then, the anomaly score of the test set is obtained by inputting the test set into the three trained base models of FB_modelB and taking the average of the anomaly scores, which is the anomaly score of the test set for FB_modelB ensemble learning.

[0165] mul_fea_model (Multiple features and multiple models): Simultaneously use four base models, MFCC_modelA, FB_modelA, MFCC_modelB, and FB_modelB. Each base model is trained three times to obtain twelve base models. Then, the anomaly score of the test set is obtained by inputting the test set into the twelve trained base models and taking the average of the anomaly scores, which is the anomaly score of the test set for mul_fea_model.

[0166] i_mul_fea_model (Improved multiple features and multiple models): Simultaneously use four base models, MFCC_modelA, FB_modelA, MFCC_modelB, and FB_modelB. Each base model is trained three times to obtain twelve base models. Then, based on the reconstruction error of the training set, select the best base model from each base model, a total of four. Then, the anomaly score of the test set is obtained by inputting the test set into the four trained best base models and taking the average of the anomaly scores, which is the anomaly score of the test set for i_mul_fea_model.

[0167] The metric used in this experiment is the area under the curve AUC (area under curve) to evaluate the quality of the model. By giving the anomaly scores of the model predicting the test set and the true labels of the test set to the AUC calculation formula, the area value under the AUC curve can be obtained. The value range of AUC is from 0 to 1. The closer AUC is to 1, the better the model performance.

[0168] Due to the randomness of the autoencoder AU model, three experiments were conducted on six ensemble learning methods, and the experimental results are shown in Tables 1 and 2.

[0169] Table 1: AUC result table of the model

[0170]

[0171] Table 2: Average AUC result table of the model

[0172]

[0173] Table 1 compares the AUC results of three experiments of MFCC_modelA ensemble learning, FB_modelA ensemble learning, MFCC_modelB ensemble learning, FB_modelB ensemble learning, mul_fea_model, and i_mul_fea_model.

[0174] Table 2 shows the average values of AUC in three experiments for six ensemble learning methods, which are the average values of each ensemble learning in Table 1. It can be seen that since the four ensemble learning methods of MFCC_modelA ensemble learning, FB_modelA ensemble learning, MFCC_modelB ensemble learning, and FB_modelB ensemble learning only consider one audio feature extraction method and one AU architecture, the average value of AUC in the three experiments is 0.92. mul_fea_model considers multiple feature extraction methods and multiple AU architectures, and integrates different types of base models, so the average value of AUC in the three experiments is 0.93. i_mul_fea_model not only considers multiple feature extraction methods and multiple AU architectures, but also considers selecting the best base model, so it has the best effect, and the average value of AUC in the three experiments is 0.94.

[0175] Technical effect:

[0176] It is shown by the mimii dataset that i_mul_fea_model considers different feature extraction methods and different autoencoder AU architectures, and also considers selecting the best base model. Therefore, compared with other ensemble learning methods, the AUC value of the method in this application is the highest. So it can be concluded that the i_mul_fea_model method can improve the abnormal sound detection accuracy of industrial equipment.

[0177] In the prior art, the categories of autoencoders AU for audio processing include basic autoencoders, regular autoencoders, and variational autoencoders. Regular autoencoders and variational autoencoders are both autoencoders obtained by improving the basic autoencoder. The basic autoencoder and the regular autoencoder can both be used for anomaly detection of motor sounds. The variational autoencoder is generally used for data generation and not for anomaly detection of motors.

[0178] In the prior art, the feature extraction methods for audio processing include Mel Frequency Cepstral Coefficients MFCC, Frequency Filter Banks FB, and mel spectrograms. Mel Frequency Cepstral Coefficients MFCC and Frequency Filter Banks FB are the most commonly used feature extraction methods in machine learning audio processing.

[0179] Embodiment 2:

[0180] Compared with Embodiment 1, the feature extraction method can also adopt the mel spectrogram feature extraction method. The technical solution is that m = 3, and the feature extraction methods include Mel Frequency Cepstral Coefficients MFCC, Frequency Filter Banks FB, and mel spectrograms, a total of three feature extraction methods; n = 2, and the autoencoder AU includes a first autoencoder modelA and a second autoencoder modelB, a total of two autoencoders AU. The same points will not be elaborated.

[0181] Compared with the above embodiments, in addition to the two AU architectures mentioned in Embodiment 1, there can be countless architectures for the autoencoder AU. Two of them are listed in Embodiment 1. Therefore, there can be many combinations of base models, and four base models are listed in Embodiment 1. The same points will not be elaborated.

[0182] Compared with the above embodiments, the architecture of the autoencoder AU can also adopt the architecture category of regular autoencoder. The autoencoder AU includes a first autoencoder and a second autoencoder, a total of one category of regular autoencoders. The difference between the first regular autoencoder and the second regular autoencoder lies in the specific data, so there are two regular autoencoders. The same points will not be elaborated.

[0183] Compared with the above embodiments, the architecture of the autoencoder AU can also adopt one architecture category of basic autoencoder and another architecture category of regular autoencoder. The autoencoder AU includes a first autoencoder and a second autoencoder, a total of two categories of autoencoders. The first autoencoder is a basic autoencoder, and the second autoencoder is a regular autoencoder. The same points will not be elaborated.

[0184] Compared with the above embodiments, the feature extraction method can also arbitrarily select two methods from the three methods of Mel Frequency Cepstral Coefficients MFCC, Frequency Filter Banks FB, and mel spectrograms for combined use. The same points will not be elaborated.

Claims

1. A method for detecting a motor based on a hybrid autoencoder ensemble learning model, characterized in that: The following steps are included: Step S1: constructing a hybrid autoencoder ensemble learning model, wherein the hybrid autoencoder ensemble learning model includes m*n base models, where m is the number of types of audio feature extraction methods, n is the number of types of autoencoder AU architectures, and one base model includes one audio feature extraction method and one autoencoder AU; In step S1, the step of constructing a hybrid autoencoder ensemble learning model includes the following steps: Step S101: construct an autoencoder AU; n=2, the autoencoder AU includes a first autoencoder modelA and a second autoencoder modelB, two types of autoencoders in total, the first autoencoder modelA and the second autoencoder modelB are a type of autoencoder, namely, a basic autoencoder; Step S102: constructing a feature extraction method; m=2, the feature extraction methods include Mel frequency cepstral coefficient MFCC and frequency filter bank FB, a total of two feature extraction methods; The base models include MFCC_modelA base model, FB_modelA base model, MFCC_modelB base model and FB_modelB base model, a total of four base models, which form a hybrid autoencoder ensemble learning model and are marked as mul_fea_model base models; Step S2: training the autoencoder AU of each base model in the hybrid autoencoder ensemble learning model to obtain a trained base model, inputting the normal audio data of the motor in the training set into each trained base model to train and obtain each optimal base model; Step S3: Input the normal and abnormal audio data of the test set motor into each best base model in the hybrid autoencoder ensemble learning model, obtain the test set anomaly score of each best base model, and average all the test set anomaly scores to obtain the test set anomaly score.

2. The method for detecting a motor based on a hybrid autoencoder integrated learning model according to claim 1, characterized in that: The step S102 further includes the following steps: The step of extracting features by the frequency filter bank FB includes sequentially subjecting the acoustic signal of the motor to short-time Fourier transform, frequency filter bank FB processing, logarithmic transformation and feature vector construction to obtain a frequency band feature vector; The steps of extracting features of Mel frequency cepstrum coefficients (MFCC) include subjecting the acoustic signal of the motor to short-time Fourier transform, Mel filter bank MF processing, logarithmic transform, discrete cosine transform and feature vector construction to obtain a Mel frequency cepstrum coefficient feature vector.

3. The method for detecting a motor based on a hybrid autoencoder integrated learning model according to claim 1, characterized in that: In step S2, the step of training the autoencoder AU of each base model in the hybrid autoencoder ensemble learning model to obtain a trained base model includes obtaining the gradient of the autoencoder AU by using a gradient descent method based on a loss function, optimizing the gradient by using an optimization function Adam, and back-propagating to obtain the weight and threshold of each autoencoder AU in the hybrid autoencoder ensemble learning model until a plurality of trained autoencoder AUs are obtained; The step of inputting the normal audio data of the motor in the training set into each trained base model to train and obtain each optimal base model includes inputting the normal audio data of the motor in the training set into each trained base model, extracting the normal audio data of the motor through the feature extraction method in the trained base model to obtain the normal audio features of the motor, inputting the normal audio features of the motor as training input features into the autoencoder AU in the trained base model, obtaining the training output features of the autoencoder AU in the trained base model, calculating the mean square error between the training input features and the training output features of the autoencoder AU in each trained base model and using it as the training error, obtaining the inverse of each training error, normalizing and obtaining the integrated weight, obtaining each optimal autoencoder AU, and then obtaining each optimal base model.

4. The method for detecting a motor based on a hybrid autoencoder integrated learning model according to claim 3 is characterized in that: In the step S2, The loss function is: (1) In formula (1): is the loss function; p is the batch size; q is the dimension of the input feature; is the i-th feature of the j-th input feature vector in the batch; is the i-th feature of the j-th output feature vector in the batch; j∈[1,p],i∈[1,q]; The training error function is: (2) In formula (2): is the training error function; z is the number of training samples; q is the dimension of the input feature; i is the i-dimensional feature vector, j is the feature vector of the j-th training data; is the i-th feature of the j-th input feature vector in the training set; is the i-th feature of the j-th reconstructed feature vector in the training set; j∈[1,z], i∈[1,q].

5. The method for detecting a motor based on a hybrid autoencoder integrated learning model according to claim 1, characterized in that: In the step S3, the step of obtaining the test set anomaly score of each best base model includes inputting the normal and abnormal audio data of the test set motor into each best base model in the hybrid autoencoder ensemble learning model, extracting the normal and abnormal audio data of the motor by the feature extraction method in the best base model to obtain the normal and abnormal audio features of the motor, inputting the normal and abnormal audio features of the motor as test input features into the autoencoder AU in the best base model, obtaining the test output features of the autoencoder AU in the best base model, that is, the reconstructed features, calculating the mean square error between the test input features and the test output features of the autoencoder AU in each best base model and using it as the reconstruction error, and the reconstruction error is the test set anomaly score.

Citation Information

Patent Citations

  • Method for detecting motor based on autoencoder ensemble learning model

    CN118226254B

  • Rolling bearing fault diagnosis method based on parallel feature learning and multiple classifiers

    CN110110768A

  • Integrated learning soft measurement modeling method based on auto-encoder diversity generation mechanism

    CN112989635A

  • Method for detecting motor based on self-encoder integrated learning model

    CN118226254A