Abnormal sound detection method and system for industrial equipment
By combining depthwise separable convolutional blocks and bottleneck layers in the abnormal sound detection model, the problems of limited computing resources and datasets for abnormal sound detection in industrial equipment are solved, and efficient and accurate abnormal sound recognition is achieved.
Patent Information
- Application Number
- CN202310987478.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-07
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2043-08-07
AI Technical Summary
Existing technologies have difficulty efficiently detecting abnormal sounds in industrial equipment, especially when computing resources and data sets are limited. The model has difficulty training decision boundaries, resulting in frequent false positives.
The first detection network and the second detection network are combined to build an abnormal sound detection model through depth-wise separable convolution blocks and bottleneck layers. The model accuracy is evaluated by AUC score, and the model with the highest score is selected for training.
The number of parameters and calculations is significantly reduced, the efficiency and accuracy of abnormal sound detection are improved, and the final recognition effect is ensured.
Smart Images

Figure CN116935888B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of abnormal sound signal detection, and in particular to a method and system for detecting abnormal sound of industrial equipment. Background Art
[0002] For industrial machinery, abnormal sound detection is used to determine whether the sound emitted by the target machine is abnormal. Abnormal sounds may indicate a machine malfunction, and timely detection can reduce risks and losses. Monitoring the machine's operating status by listening to its acoustic signals has been widely used in anomaly detection and predictive maintenance.
[0003] Industrial equipment operates under stable conditions, and the characteristics of its sound signals are relatively stable, with very few anomalies or failures. This makes it difficult to obtain truly abnormal sound samples. Furthermore, industrial audio signals are diverse and have complex background noise components.
[0004] Therefore, the superiority of unsupervised methods based on deep learning is particularly prominent in the field of abnormal sound detection.
[0005] Advanced deep learning networks often require extensive computing resources and datasets for training, which exceeds the computing power of many mobile and embedded devices. In existing deep learning methods, when normal samples from different machines are very similar, it is difficult for the model to train the decision boundary, which can lead to frequent false positives in the detection system. Summary of the Invention
[0006] To solve the above technical problems, the present invention provides a method for detecting abnormal sound in industrial equipment, comprising the steps of:
[0007] S1: Build an abnormal sound detection model, which includes: a first detection network and a second detection network;
[0008] S2: Train the first detection network using the sample data to obtain a first normal score and a first abnormal score;
[0009] S3: Training the second detection network using the sample data to obtain a second normal score and a second abnormal score;
[0010] S4: Calculate the AUC score using the first normal score, the first abnormal score, the second normal score, and the second abnormal score;
[0011] S5: Repeat steps S2-S4 until the maximum number of iterations is reached, and the abnormal sound detection model with the highest AUC score is used as the trained abnormal sound detection model;
[0012] S6: Identify the sound to be detected through the trained abnormal sound detection model to obtain the abnormality detection results of the industrial equipment.
[0013] Preferred:
[0014] The first detection network includes a first two-dimensional convolutional layer, multiple bottleneck layers, a second two-dimensional convolutional layer, an average pooling layer and a third two-dimensional convolutional layer, which are connected in sequence.
[0015] Preferably, the bottleneck layer is constructed as follows:
[0016] Get a 1×1 ordinary convolutional layer, upgrade the ordinary convolutional layer to a high-dimensional space, replace each ordinary convolutional block in the ordinary convolutional layer with a depth-wise separable convolutional block in the high-dimensional space, and then reduce the dimension to 1×1 space to obtain a depth-wise separable convolutional layer;
[0017] In the depthwise separable convolutional layer, the nonlinear activation function is replaced by a linear activation function, and the two low-dimensional tensors after the input and the linear activation function are connected by an inverted residual block to obtain the bottleneck layer.
[0018] Preferred:
[0019] The second detection network includes an encoder, a latent space module, a decoder, and a reconstruction module connected in sequence.
[0020] Preferably, step S2 is specifically as follows:
[0021] Input the sample data into the first detection network for identification to obtain the first normal sample set and the first abnormal sample set Get the first normal score of each normal sample in the first normal sample set Get the first anomaly score of each anomaly sample in the first anomaly sample set
[0022] in, is the i-th normal sample in the first normal sample set, N1 - is the total number of samples in the first normal sample set, is the jth abnormal sample in the first abnormal sample set, N1 + is the total number of samples in the first abnormal sample set.
[0023] Preferably, step S3 is specifically as follows:
[0024] The sample data is input into the second detection network for identification to obtain the second normal sample set and the second abnormal sample set Get the second normal score of each normal sample in the second normal sample set Get the second anomaly score of each abnormal sample in the second abnormal sample set
[0025] in, is the uth normal sample in the second normal sample set, N2 - is the total number of samples in the second normal sample set, is the vth abnormal sample in the second abnormal sample set, N2 + is the total number of samples in the second abnormal sample set.
[0026] Preferably, the calculation formula of the AUC score in step S4 is:
[0027]
[0028]
[0029]
[0030]
[0031] Among them, if Then K - =N2 - ,like Then K - =N1-, if Then K + =N2 + ,like Then K + =N1 + ;
[0032] is the first normal fraction, is the first anomaly score, is the second normal fraction, is the second anomaly score;
[0033] N1- is the total number of samples in the first normal sample set, N1 + is the total number of samples in the first abnormal sample set, N2- is the total number of samples in the second normal sample set, N2 + is the total number of samples in the second abnormal sample set.
[0034] An abnormal sound detection system for industrial equipment, including modules:
[0035] A model building module is used to build an abnormal sound detection model, the abnormal sound detection model includes: a first detection network V2 and a second detection network AE;
[0036] A first detection network training module, configured to train the first detection network using sample data to obtain a first normal score and a first abnormal score;
[0037] A second detection network training module is used to train the second detection network using sample data to obtain a second normal score and a second abnormal score;
[0038] An AUC score calculation module is configured to calculate an AUC score using the first normal score, the first abnormal score, the second normal score, and the second abnormal score;
[0039] An iterative training module is used to repeat the training until the maximum number of iterations is reached, and the abnormal sound detection model with the highest AUC score is used as the trained abnormal sound detection model;
[0040] The anomaly detection module is used to identify the sound to be detected through the trained abnormal sound detection model to obtain anomaly detection results for industrial equipment.
[0041] The present invention has the following beneficial effects:
[0042] The first detection network of the present invention adopts a bottleneck layer with a depthwise separable convolution block. The abnormal sound detection model is constructed by combining the first detection network and the second detection network. This model design can significantly reduce the number of parameters and computational complexity, thereby improving the efficiency of abnormal sound detection. The recognition accuracy of the abnormal sound detection model is accurately evaluated by calculating the AUC score, and the abnormal sound detection model with the highest AUC score is selected as the trained abnormal sound detection model to ensure the final recognition effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 This is a flow chart of a method according to an embodiment of the present invention;
[0044] Figure 2 Schematic diagram of depth-separable convolution block in three-dimensional space;
[0045] Figure 3 This is a schematic diagram of the bottleneck layer in three-dimensional space;
[0046] Figure 4 is a structural diagram of the second detection network;
[0047] Figure 5 This is a schematic diagram of the judgment results of the abnormal sound detection model;
[0048] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0049] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0050] Reference Figure 1 The present invention provides a method for detecting abnormal sound of industrial equipment, comprising the steps of:
[0051] S1: Build an abnormal sound detection model, which includes: a first detection network and a second detection network;
[0052] S2: Train the first detection network using the sample data to obtain a first normal score and a first abnormal score;
[0053] S3: Training the second detection network using the sample data to obtain a second normal score and a second abnormal score;
[0054] S4: Calculate the AUC score using the first normal score, the first abnormal score, the second normal score, and the second abnormal score;
[0055] S5: Repeat steps S2-S4 until the maximum number of iterations is reached, and the abnormal sound detection model with the highest AUC score is used as the trained abnormal sound detection model;
[0056] S6: Identify the sound to be detected through the trained abnormal sound detection model to obtain the abnormality detection results of the industrial equipment.
[0057] Further:
[0058] The first detection network includes a first two-dimensional convolutional layer, multiple bottleneck layers, a second two-dimensional convolutional layer, an average pooling layer and a third two-dimensional convolutional layer, which are connected in sequence.
[0059] Specifically, the first detection network uses the MobileNetV2 network. Table 1 lists the detailed architecture of the MobileNetV2 network:
[0060] Table 1
[0061]
[0062] In Table 2, the first column represents the size of the input tensor of each layer of the MobileNetV2 network; n indicates that each row describes a sequence consisting of one or more identical (module stride) layers, that is, the number of repetitions of bottleneck; all layers in the same sequence have the same number of convolution blocks c; the first module of each sequence has a stride s, and the stride of other modules is 1; k is the number of output categories; all spatial convolutions use a 3×3 kernel; the expansion factor t is always applied to the input size described in Table 1; except for the first layer, a constant expansion rate is used throughout the network; the present invention trains the network to identify from which part the observed signal is generated, and the model outputs a softmax value, that is, the predicted probability of each part.
[0063] Furthermore, the bottleneck layer construction process is:
[0064] Get a 1×1 ordinary convolutional layer, upgrade the ordinary convolutional layer to a high-dimensional space, replace each ordinary convolutional block in the ordinary convolutional layer with a depth-wise separable convolutional block in the high-dimensional space, and then reduce the dimension to 1×1 space to obtain a depth-wise separable convolutional layer;
[0065] In the depthwise separable convolutional layer, the nonlinear activation function is replaced by a linear activation function, and the two low-dimensional tensors after the input and the linear activation function are connected by an inverted residual block to obtain the bottleneck layer.
[0066] Specifically, the MobileNetV2 network uses depthwise separable convolution as the building block of the network architecture, splitting the ordinary convolution into two independent "decoupled" versions to replace the complete convolution operator; the depthwise separable convolution block structure is shown in Table 2 below:
[0067] Table 2
[0068]
[0069] Ordinary convolution uses h i ×w i ×d i The input tensor L i , and apply the convolution kernel To generate h i ×w i ×d j The output tensor L j , whose computational cost is h i w i ·d i ·d j ·k·k; The computational cost of depth-wise separable convolution is calculated according to formula (1):
[0070] k 2 ×h i ×w i ×d i +1×1×d i ×h i ×w i ×d j (1)
[0071] That is hi·w i ·d i (k 2 +d j ) (2)
[0072] Compared with ordinary convolution, depth-wise separable convolution effectively reduces k 2 MobileNetV2 uses k=3 (3×3 depth-wise separable convolution), so the computational cost is 1 / 8 to 1 / 9 of that of ordinary convolution, significantly reducing the number of parameters and computational complexity with only a slight decrease in accuracy.
[0073] Figure 2 This is a schematic diagram of the depthwise separable convolution block in three-dimensional space. The standard ordinary convolution block is a convolution block that can process information from multiple channels at the same time. The depthwise separable convolution block includes two separate convolutions: the first layer is called depthwise convolution, which uses a single convolution block to process each input channel separately to achieve lightweight filtering; the second layer is pointwise convolution, which constructs new features by processing linear combinations of cross-channel information.
[0074] Another key to the lightweight MobileNetV2 network is the use of a bottleneck layer, which can be expressed as a combination of three operators, as shown in Equations (3) and (4):
[0075]
[0076] N=ReLU6odwiseoReLU6 (4)
[0077] Where A:R s×s×k →R s×s×n , represents linear transformation; N:R s×s×n →R s'×s'×n , represents nonlinear cross-channel transformation; B:R s'×s'×n → s'×s'×k' , which means a linear transformation of the output domain.
[0078] Assuming the input domain is |x| and the output domain is |y|, the computational cost of calculating F(x) is expressed as:
[0079] |s 2 k|+|s' 2 k'|+O(max(s 2 ,s' 2 )) (5)
[0080] Figure 3 This is a schematic diagram of the bottleneck layer in three-dimensional space. Compared with traditional layers, the effective depth-wise separable convolution reduces the amount of computation by almost k2 times. MobileNetV2 uses k=3 (3×3 depth-wise separable convolution), so the computational cost is 8 to 9 times smaller than the standard convolution, while the accuracy is only slightly reduced.
[0081] Further:
[0082] The second detection network includes an encoder, a latent space module, a decoder, and a reconstruction module connected in sequence.
[0083] Specifically, the structure of the second detection network AE is as follows Figure 4As shown in the figure, the encoder and decoder networks are composed of an input fully connected neural network (FCN) layer, four fully connected dense layers (Dense Layer) and an output FCN layer respectively; each DenseLayer layer has 512 hidden units, followed by batch normalization (Normalization) and ReLU activation function; the bottleneck layer (Bottleneck Layer) of the second detection network AE is set to a fully connected layer with 8 hidden units, thus generating an 8-dimensional potential space; except for the output layer of the decoder, the ReLU activation function is used after each FCN layer;
[0084] The encoder maps high-dimensional input samples to low-dimensional abstract representations to achieve sample compression and dimensionality reduction; the decoder converts the abstract representation into the desired output to obtain the reconstruction of the input sample. The difference between the original input vector and the network output vector is called the reconstruction error. AE backpropagates the error through the gradient descent algorithm to adjust the network parameters to minimize the reconstruction error.
[0085] Suppose that a training set x∈R is given T , for a D-dimensional input x, the input dimension x d The value range is [0,1]. The autoencoder obtains a reconstruction y that is as close to the input as possible by learning the feedforward hidden representation h(x) of its input x, as shown in Equations (6) and (7):
[0086] h(x)=g(b+w1x) (6)
[0087] y=sigm(c+w2h(x)) (7)
[0088] Where w1, w2 are matrices representing the weights of the encoding layer and decoding layer respectively; b, c are vectors; and g is a nonlinear activation function.
[0089] To train the AE network, we first need to specify the training loss function. For binary observations, the cross-entropy loss function is generally selected. The autoencoder is trained to optimize the parameters (w1, w2, b, c) to reduce the average loss of the training examples. Mini-batch stochastic gradient descent is usually used, which is specifically expressed as Equation (8):
[0090]
[0091] Furthermore, step S2 is specifically as follows:
[0092] Input the sample data into the first detection network for identification to obtain the first normal sample set and the first abnormal sample set Get the first normal score of each normal sample in the first normal sample set Get the first anomaly score of each anomaly sample in the first anomaly sample set
[0093] in, is the i-th normal sample in the first normal sample set, N1- is the total number of samples in the first normal sample set, is the jth abnormal sample in the first abnormal sample set, N1 + is the total number of samples in the first abnormal sample set.
[0094] Furthermore, step S3 is specifically as follows:
[0095] The sample data is input into the second detection network for identification to obtain the second normal sample set and the second abnormal sample set Get the second normal score of each normal sample in the second normal sample set Get the second anomaly score of each abnormal sample in the second abnormal sample set
[0096] in, is the u-th normal sample in the second normal sample set, N2- is the total number of samples in the second normal sample set, is the vth abnormal sample in the second abnormal sample set, N2 + is the total number of samples in the second abnormal sample set.
[0097] Furthermore, the goal of anomaly monitoring is to determine whether a given data point is normal. Taking a classification task as an example, if there are two labels of samples, 0 and 1, then for an input sample, the classification model may also produce two output results, 0 and 1. Let label 1 be a positive sample (Positive), label 0 be a negative sample (Negative), the correct judgment of the model is true (True), and the wrong judgment is false (False), then the judgment of the model can be summarized into four cases, such as Figure 5 The specific definitions are listed below:
[0098] True Positive (TP): The condition is true and the prediction is also true;
[0099] True Negative (TN): The condition is false and the prediction is also false;
[0100] False positive (FP): the condition is false and the prediction is true;
[0101] False Negative (FN): The condition is true but the prediction is false.
[0102] Putting the above definitions together forms a confusion matrix. Based on the confusion matrix, we can derive the values of accuracy, precision, and recall to further understand the performance of the model. The following figure shows the confusion matrix and all the formulas.
[0103] Accuracy (Acc) describes how many of the data sets the model correctly predicts as positive or negative.
[0104] Precision (Pre) describes how many of the true predictions the model makes are correct.
[0105] Recall (Rec) describes how many of the true data points in the dataset the model correctly predicts.
[0106] On this basis, more metrics can be derived to evaluate model performance. The True Positive Rate (TPR), like the Recall Rate, describes how many actual data points are predicted as true by the model. The False Positive Rate (FPR) describes how many actual false data points are predicted as positive by the model. It is the ratio of false positives to all false data points. The TPR and FPR form a Receiver Operating Characteristic (ROC) curve, the Area Under the Curve (AUC). A higher AUC score indicates a more accurate abnormal sound detection model, and the maximum AUC score is 1.
[0107] The calculation formula of the AUC score in step S4 is:
[0108]
[0109]
[0110]
[0111]
[0112] Among them, if Then K-=N2-, if Then K-=N1-, if Then K + =N2 + ,like Then K + =N1 + ;
[0113] is the first normal fraction, is the first anomaly score, is the second normal fraction, is the second anomaly score;
[0114] N1- is the total number of samples in the first normal sample set, N1 + is the total number of samples in the first abnormal sample set, N2- is the total number of samples in the second normal sample set, N2 + is the total number of samples in the second abnormal sample set.
[0115] An abnormal sound detection system for industrial equipment, including modules:
[0116] A model building module is used to build an abnormal sound detection model, the abnormal sound detection model includes: a first detection network V2 and a second detection network AE;
[0117] A first detection network training module, configured to train the first detection network using sample data to obtain a first normal score and a first abnormal score;
[0118] A second detection network training module is used to train the second detection network using sample data to obtain a second normal score and a second abnormal score;
[0119] An AUC score calculation module is configured to calculate an AUC score using the first normal score, the first abnormal score, the second normal score, and the second abnormal score;
[0120] An iterative training module is used to repeat the training until the maximum number of iterations is reached, and the abnormal sound detection model with the highest AUC score is used as the trained abnormal sound detection model;
[0121] The anomaly detection module is used to identify the sound to be detected through the trained abnormal sound detection model to obtain anomaly detection results for industrial equipment.
[0122] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.
[0123] The serial numbers of the embodiments of the present invention are for descriptive purposes only and do not represent superiority or inferiority of the embodiments. In a unit claim that lists several means, several of these means may be embodied by the same item of hardware. The use of the terms first, second, and third, etc., does not denote any order and should be construed as identifiers.
[0124] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for detecting abnormal sound in industrial equipment, characterized in that: Including steps: S1: Build an abnormal sound detection model, which includes: a first detection network and a second detection network; The first detection network includes a first two-dimensional convolutional layer, multiple bottleneck layers, a second two-dimensional convolutional layer, an average pooling layer, and a third two-dimensional convolutional layer connected in sequence. The first detection network adopts the MobileNetV2 network, which uses depthwise separable convolution as the building block of the network architecture. MobileNetV2 uses 3×3 depthwise separable convolution. The second detection network includes an encoder, a latent space module, a decoder, and a reconstruction module connected in sequence; S2: Train the first detection network using the sample data to obtain a first normal score and a first abnormal score; S3: Training the second detection network using the sample data to obtain a second normal score and a second abnormal score; S4: Calculate the AUC score using the first normal score, the first abnormal score, the second normal score, and the second abnormal score; S5: Repeat steps S2-S4 until the maximum number of iterations is reached, and the abnormal sound detection model with the highest AUC score is used as the trained abnormal sound detection model; S6: Identify the sound to be detected through the trained abnormal sound detection model to obtain the abnormality detection results of the industrial equipment.
2. The method for detecting abnormal sound in industrial equipment according to claim 1, wherein: The construction process of the bottleneck layer is: Get a 1×1 ordinary convolutional layer, upgrade the ordinary convolutional layer to a high-dimensional space, replace each ordinary convolutional block in the ordinary convolutional layer with a depth-wise separable convolutional block in the high-dimensional space, and then reduce the dimension to 1×1 space to obtain a depth-wise separable convolutional layer; In the depthwise separable convolutional layer, the nonlinear activation function is replaced by a linear activation function, and the two low-dimensional tensors after the input and the linear activation function are connected by an inverted residual block to obtain the bottleneck layer.
3. The method for detecting abnormal sound in industrial equipment according to claim 1, wherein: Step S2 is specifically as follows: Input the sample data into the first detection network for identification to obtain the first normal sample set and the first abnormal sample set Get the first normal score of each normal sample in the first normal sample set Get the first anomaly score of each anomaly sample in the first anomaly sample set in, is the i-th normal sample in the first normal sample set, N1 - is the total number of samples in the first normal sample set, is the jth abnormal sample in the first abnormal sample set, N1 + is the total number of samples in the first abnormal sample set.
4. The method for detecting abnormal sound in industrial equipment according to claim 1, wherein: Step S3 is specifically as follows: The sample data is input into the second detection network for identification to obtain the second normal sample set and the second abnormal sample set Get the second normal score of each normal sample in the second normal sample set Get the second anomaly score of each abnormal sample in the second abnormal sample set in, is the uth normal sample in the second normal sample set, N2 - is the total number of samples in the second normal sample set, is the vth abnormal sample in the second abnormal sample set, N2 + is the total number of samples in the second abnormal sample set.
5. The method for detecting abnormal sound in industrial equipment according to claim 1, wherein: The calculation formula of the AUC score in step S4 is: Among them, if Then K - =N2 - ,like Then K - =N1 - ,like Then K + =N2 + ,like Then K + =N1 + ; is the first normal fraction, is the first anomaly score, is the second normal fraction, is the second anomaly score, x a - To pass The normal sample when the extreme value of the normal score calculated by the calculation formula is x b + To pass Abnormal samples when the anomaly score calculated by the calculation formula reaches an extreme value; N1 - is the total number of samples in the first normal sample set, N1 + is the total number of samples in the first abnormal sample set, N2 - is the total number of samples in the second normal sample set, N2 + is the total number of samples in the second abnormal sample set.
6. An abnormal sound detection system for industrial equipment, characterized in that: Included modules: A model building module is used to build an abnormal sound detection model, the abnormal sound detection model includes: a first detection network V2 and a second detection network AE; A first detection network training module is configured to train the first detection network using sample data to obtain a first normal score and a first abnormal score; wherein the first detection network includes a first two-dimensional convolutional layer, multiple bottleneck layers, a second two-dimensional convolutional layer, an average pooling layer, and a third two-dimensional convolutional layer connected in sequence; the first detection network adopts a MobileNetV2 network, which uses depthwise separable convolution as a building block of the network architecture, and MobileNetV2 uses 3×3 depthwise separable convolution; A second detection network training module is used to train the second detection network using sample data to obtain a second normal score and a second abnormal score; wherein the second detection network includes an encoder, a latent space module, a decoder, and a reconstruction module connected in sequence; An AUC score calculation module is configured to calculate an AUC score using the first normal score, the first abnormal score, the second normal score, and the second abnormal score; An iterative training module is used to repeat the training until the maximum number of iterations is reached, and the abnormal sound detection model with the highest AUC score is used as the trained abnormal sound detection model; The anomaly detection module is used to identify the sound to be detected through the trained abnormal sound detection model to obtain anomaly detection results for industrial equipment.
Citation Information
Patent Citations
Air conditioner indoor unit noise anomaly detection method based on time-frequency domain deep learning algorithm
CN112669879A
Abnormal sound detection device, abnormal sound generation device, and abnormal sound generation method
CN113899577A