A machine abnormal noise recognition and detection method based on deep learning
Through the machine noise recognition and detection method based on deep learning, the problems of difficulty in data collection, limited feature extraction and inability to adapt to massive data in machine equipment fault diagnosis are solved, and more efficient and accurate fault abnormal sound recognition is achieved.
Patent Information
- Application Number
- CN202210426476.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-22
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-04-22
AI Technical Summary
The existing machine and equipment fault diagnosis methods have problems such as difficulty in data collection, limited feature extraction and inability to adapt to massive data, resulting in low accuracy of fault diagnosis.
Using deep learning-based machine noise recognition detection method, neural network models are constructed and trained through data preprocessing and feature extraction, including encoder, decoder and loss functions to identify fault abnormal sounds of machine equipment.
The process of fault abnormal sound recognition is simplified, and the efficiency and recognition accuracy of fault abnormal sound prediction are improved.
Smart Images

Figure CN114997210B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a machine abnormal sound recognition and detection method based on deep learning, and belongs to the technical field of deep learning. Background Art
[0002] The current development of intelligent manufacturing has led to an increasing number of modern machine equipment failures. When a machine equipment fails, it will cause factory production to stagnate and have a significant impact on the factory's economic benefits. The occurrence of major failures may also seriously affect the safety of life and property. Therefore, the study of intelligent fault diagnosis technology for machine equipment can not only reduce or avoid major accidents in the factory's production process as much as possible, but is also an urgent need to ensure life safety and reduce property losses.
[0003] Faced with the increasing demand for higher-quality machine equipment fault diagnosis in the industrial sector, the emergence of smarter and more efficient methods has become more urgent. The rapid development of science and technology in many fields has not only improved the performance of computer hardware, but also reduced its price. At the same time, the introduction and improvement of various algorithms have made it possible to store and process massive data, providing a good foundation for the implementation and promotion of deep learning. Driven by big data, the application of deep learning has well solved the shortcomings of traditional methods. The neural network model can better describe the feature space of data, making it better applied to the processing, classification and prediction of massive data. Deep learning simplifies the process of early feature extraction of traditional methods when processing massive data, and can better mine deeper data abstract features in the sound signals of machine equipment. Deep learning makes it possible to promote more sound recognition products, and continuously promotes the development of sound recognition related technologies. Improving the timeliness and accuracy of fault diagnosis can bring more practical and economic value to mankind.
[0004] However, there are currently several problems with machine equipment fault diagnosis:
[0005] (1) The structure of machinery and equipment is highly complex, there are many types of machine and equipment failures, and the probability of failure is small, so it is difficult to collect a large amount of failure sample data.
[0006] (2) The machine equipment fault features extracted through manual experience are limited, which restricts the performance of the classification model and makes the machine equipment fault diagnosis accuracy low.
[0007] (3) Traditional machine learning classification methods cannot adapt to massive amounts of data. Summary of the invention
[0008] The purpose of the present invention is to provide a machine abnormal sound recognition and detection method based on deep learning, which simplifies the process of abnormal sound recognition of machine equipment failures, and also improves the efficiency of abnormal sound prediction and recognition accuracy.
[0009] In order to achieve the above object, the present invention is implemented through the following technical solutions:
[0010] 1) Data preprocessing, feature extraction, data cleaning, reconstruction and denoising, and division of training data and test data;
[0011] 2) Build and train a neural network and use a dataset to train the model;
[0012] 3) Test the neural network model using the test data in step 1;
[0013] 4) Use the model to identify the sound and determine whether it is an abnormal sound.
[0014] Preferably, the neural network includes an encoder, a decoder and a loss function, which are as follows:
[0015] (1) Encoder:
[0016] The encoder consists of a convolution layer and a pooling layer. The original data x is input into the convolution layer, and the data feature map h is generated by convolution calculation. The specific formula is as follows:
[0017] h k =σ(x*W k +b k )
[0018] In the formula, * represents two-dimensional convolution, b k is the bias of the kth channel in the encoding stage, σ is a nonlinear function, using the sigmoid function, and w k is the kth convolution kernel of the convolution layer;
[0019] The output of the convolutional layer is passed into the pooling layer as input, using maximum pooling, that is:
[0020] h i+1x,y =max (i,j)∈N(x,y) (h i,j,k )
[0021] Where i represents the number of layers, (x, y) represents the coordinates after pooling, (j, k) represents the pooled coordinates of the feature map of the previous layer, and N(x, y) is the pooling block;
[0022] (2) Decoder
[0023] The decoder consists of a deconvolution layer and a depooling layer, and its calculation process is as follows:
[0024]
[0025] (3) Loss Function
[0026] The convolutional autoencoder encodes and decodes the original data, and continuously learns to adjust the convolution kernel W and bias b to minimize the reconstruction error between the decoded output and the encoded input, thereby achieving the purpose of optimizing the loss function. The loss function is shown in the following figure:
[0027]
[0028] Preferably, the specific steps of data preprocessing and feature extraction are as follows:
[0029] ① Perform discrete Fourier transform on the preprocessed data sound signal to obtain the signal's spectrum distribution information:
[0030]
[0031] Where x(n) is the preprocessed audio signal input, and N is the number of Fourier transform points;
[0032] ② Take the square of the modulus to obtain the energy spectrum:
[0033]
[0034] Where E(k) is the energy spectrum.
[0035] ③ Pass the obtained energy spectrum through the Mel filter bank;
[0036] ④ Take the logarithm of the output signal of the Mel filter bank to get the logarithmic energy of the output of the mth Mel filter.
[0037] Quantity, as shown in the formula:
[0038]
[0039] Where Hm(k) is the transfer function of the filter bank and m is the number of Mel filters.
[0040] The advantages of the present invention are that the present invention not only simplifies the process of identifying abnormal sound of machine equipment failure, but also improves the efficiency of abnormal sound prediction and the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0042] Figure 1 It is a schematic diagram of the process structure of the present invention.
[0043] Figure 2 Schematic diagram of a simple model of the automatic encoder of the present invention.
[0044] Figure 3 This is a schematic diagram of the structure of the convolutional autoencoder of the present invention.
[0045] Figure 4 Schematic diagram of the sound recognition model of the present invention.
[0046] Figure 5 It is a schematic diagram of the feature extraction process of the present invention.
[0047] Figure 6 This is a schematic diagram of the convolutional autoencoder model structure of the present invention. DETAILED DESCRIPTION
[0048] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0049] Convolutional Autoencoder
[0050] Since the number of neurons in the bottleneck layer of the autoencoder is much lower than the number of neurons in the layers before and after this layer, the autoencoder cannot extract the optimal representation of the original data input from the noisy data. Moreover, when observing the spectrum diagram reconstructed by the ordinary autoencoder, it can be found that the reconstruction error of this method is large and it is impossible to learn the original data. Oleg et al. used convolutional neural networks as feature extractors in image anomaly detection to complete the problem of image anomaly detection. Based on this idea, the autoencoder was combined with the convolutional neural network to design a convolutional autoencoder.
[0051] 1) Data preprocessing, feature extraction, data cleaning, reconstruction and denoising, and data division for training and testing; the specific steps of data preprocessing and feature extraction are as follows:
[0052] ① Perform discrete Fourier transform on the preprocessed data sound signal to obtain the signal's spectrum distribution information:
[0053]
[0054] Where x(n) is the preprocessed audio signal input, and N is the number of Fourier transform points;
[0055] ② Take the square of the modulus to obtain the energy spectrum:
[0056]
[0057] Where E(k) is the energy spectrum.
[0058] ③ Pass the obtained energy spectrum through the Mel filter bank;
[0059] ④ Take the logarithm of the output signal of the Mel filter bank to get the logarithmic energy of the output of the mth Mel filter.
[0060] Quantity, as shown in the formula:
[0061]
[0062] Where Hm(k) is the transfer function of the filter bank and m is the number of Mel filters.
[0063] 2) Build and train a neural network and use a dataset to train the model;
[0064] The neural network includes an encoder, a decoder and a loss function, which are as follows:
[0065] (1) Encoder:
[0066] The encoder consists of a convolution layer and a pooling layer. The original data x is input into the convolution layer, and the data feature map h is generated by convolution calculation. The specific formula is as follows:
[0067] h k =σ(x*W k +b k )
[0068] In the formula, * represents two-dimensional convolution, b k is the bias of the kth channel in the encoding stage, σ is a nonlinear function, using the sigmoid function, and w k is the kth convolution kernel of the convolution layer;
[0069] The output of the convolutional layer is passed into the pooling layer as input, using maximum pooling, that is:
[0070] h i+1,x,y =max (i,j)∈N(x,y) (h i,j,k )
[0071] Where i represents the number of layers, (x, y) represents the coordinates after pooling, (j, k) represents the pooled coordinates of the feature map of the previous layer, and N(x, y) is the pooling block;
[0072] (2) Decoder
[0073] The decoder consists of a deconvolution layer and a depooling layer, and its calculation process is as follows:
[0074]
[0075] (3) Loss Function
[0076] The convolutional autoencoder encodes and decodes the original data, and continuously learns to adjust the convolution kernel W and bias b to minimize the reconstruction error between the decoded output and the encoded input, thereby achieving the purpose of optimizing the loss function. The loss function is as follows:
[0077]
[0078] 3) Test the neural network model using the test data in step 1;
[0079] 4) Use the model to identify the sound and determine whether it is an abnormal sound.
[0080] Convolutional Autoencoder Model Structure:
[0081] The detailed structure of the designed convolutional autoencoder is as follows Figure 6 As shown in the figure. There are nine layers in the convolutional autoencoder structure, with four layers in the encoder and four layers in the decoder. Since ReLU can solve the gradient vanishing problem, ReLU is added as the activation function after each layer. The number of filters in the convolutional layer is 128, 64, 32, and 16. In the encoder convolutional layer, a 3*3 convolution kernel is used, and the downsampling size is 2*2. In the decoder convolutional layer, a 3*3 convolution kernel and a 2*2 upsampling are used.
Claims
1. A machine abnormal noise recognition and detection method based on deep learning, It is characterized in that The following steps are involved: 1) Data preprocessing, feature extraction, data cleaning, reconstruction and denoising, and division of training data and test data; 2) Build and train a neural network and use a dataset to train the model; 3) Test the neural network model using the test data in step 1; 4) Use the model to identify the sound and determine whether it is an abnormal sound; The neural network includes an encoder, a decoder and a loss function, which are as follows: (1) Encoder: The encoder consists of a convolution layer and a pooling layer. The original data x is input into the convolution layer, and the data feature map h is generated by convolution calculation. The specific formula is as follows: In the formula, * represents two-dimensional convolution, is the bias of the kth channel in the encoding stage, σ is a nonlinear function, using the sigmoid function, is the kth convolution kernel of the convolution layer; The output of the convolutional layer is passed into the pooling layer as input, using maximum pooling, that is: In the formula, i represents the number of layers, (x, y) represents the coordinates after pooling, (j, k) represents the pooled coordinates of the feature map of the previous layer, and N(x, y) is the pooling block; (2) Decoder The decoder consists of a deconvolution layer and a depooling layer, and its calculation process is as follows: (3) Loss Function The convolutional autoencoder encodes and decodes the original data, and continuously learns to adjust the convolution kernel W and bias b to minimize the reconstruction error between the decoded output and the encoded input, thereby achieving the purpose of optimizing the loss function. The loss function is shown in the following figure: ; The specific steps of data preprocessing and feature extraction are as follows: ① Perform discrete Fourier transform on the preprocessed data sound signal to obtain the signal's spectrum distribution information: Where x(n) is the preprocessed audio signal input, and N is the number of Fourier transform points; ② Take the square of the modulus to obtain the energy spectrum: Where E(k) is the energy spectrum. ③ Pass the obtained energy spectrum through the Mel filter bank; ④ Take the logarithm of the output signal of the Mel filter bank to get the logarithmic energy of the output of the mth Mel filter. Quantity, as shown in the formula: Where Hm(k) is the transfer function of the filter bank and m represents the number of Mel filters.
Citation Information
Patent Citations
Transformer substation equipment sound fault detection and positioning method based on deep learning
CN112183647A