Equipment abnormal sound detection method based on adaptive filtering and deep learning

Through the combination of adaptive filtering and deep learning, a joint training framework and feature fusion method are built, which solves the problems of high false alarm rate and model deployment of equipment antonyms detection in complex noise environments, and realizes accurate retention and efficient detection of different audio bands, which are suitable for real-time monitoring of industrial environments.

CN120340530APending Publication Date: 2025-07-18HEFEI TAIZE TURBINE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510701240.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing equipment antonym detection technology is difficult to achieve accurate retention of different audio bands in complex noise environments, with high false alarm rates, and deep models are difficult to efficiently deploy on edge computing devices.

Method used

The combined method of adaptive filtering and deep learning is adopted to build a joint training framework, and through the dynamic fusion of adaptive filtering and residual convolution network, the lightweight residual network is enhanced, the adversarial samples are generated, the filter frequency response curve is dynamically adjusted, and the model parameters are compressed in combination with the channel attention mechanism to achieve accurate fusion of features and the key frequency band retention of the abnormal sound signals.

Benefits of technology

In complex noise environments, the retention rate of the different audio band is >95%, and the false alarm rate is reduced to below 2%, which significantly improves detection accuracy and reliability, meets the low latency and low power consumption requirements of edge devices, and is suitable for real-time monitoring of industrial environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340530A_ABST
    Figure CN120340530A_ABST
Patent Text Reader

Abstract

The invention provides an equipment abnormal sound detection method based on adaptive filtering and deep learning, and the method comprises the following steps: S1, collecting a large number of normal working audio signals of equipment, and carrying out the filtering processing of the obtained normal working audio signals; s2, constructing a joint training framework of adaptive filtering and a residual convolutional network; s3, constructing a three-channel feature fusion framework; s4, adding an adversarial enhanced lightweight residual network; s5, constructing an adaptive frequency band selection algorithm, and dynamically adjusting the frequency response curve of the filter based on the importance of the features output by the residual network; s6, filtering processing is carried out on the equipment audio signals collected in real time; and after being processed by the steps S2 to S5, the feature library is compared with the feature library during training, and abnormality is detected. According to the method, accurate retention of the abnormal audio frequency segments is realized, and the retention rate of the abnormal audio frequency segments is gt; and the false alarm rate is reduced from 25% to below 2%, so that the accuracy and the reliability of abnormal sound detection are remarkably improved, and the abnormal sound of the equipment can be stably detected in a complex noise environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of abnormal sound detection of equipment, and particularly relates to a method for detecting abnormal sound of equipment based on adaptive filtering and deep learning. Background Art

[0002] The abnormal sound detection technology of industrial equipment is one of the core research directions in the field of equipment health management (PHM). Its development has experienced a technological transformation from signal processing dominance to data-driven evolution. Early research focused on time-frequency analysis and pattern recognition of acoustic signals, achieving noise suppression through signal processing techniques such as LMS algorithm adaptive filtering and wavelet transform, and completing anomaly detection by combining artificial feature extraction means such as MFCC (Mel Frequency Cepstral Coefficient) and spectral kurtosis. Such methods rely on physical prior knowledge to construct feature engineering, and rely on artificial experience to design features and filtering parameters. They are suitable for simple scenarios with single noise type and stable working conditions, and have a certain reliability, but it is difficult to cope with complex noise interference and the dynamic characteristics of non-stationary signals.

[0003] With the penetration of machine learning technology, shallow learning models (such as SVM, random forest) and time series modeling methods (such as HMM, GRU) have gradually been applied to this field, improving the mapping ability between features and fault categories through statistical learning. However, such methods are still limited by the characterization dimension of artificial features, and have insufficient feature fusion ability for multi-source heterogeneous acoustic signals. Generally speaking, the current equipment abnormal sound detection technology has the following main problems: traditional filtering algorithms are difficult to form parameter co-optimization with neural networks, resulting in the separation of noise suppression and feature learning objectives; existing methods for three-dimensional feature fusion of time-frequency-space have not yet formed a unified framework for joint modeling of the time dynamics, frequency domain locality, and spatial correlation of acoustic signals; in the balance design of lightweight and accuracy, the contradiction between the number of parameters of deep models and the edge computing ability of industrial equipment is becoming increasingly prominent, and model compression often sacrifices the time series modeling ability. Summary of the Invention

[0004] The main purpose of the present invention is to provide a method for detecting abnormal sound of equipment based on adaptive filtering and deep learning, which can effectively solve the problems in the background art, accurately retain the abnormal audio segment, with the retention rate of the abnormal audio segment > 95%, and reduce the false alarm rate from 25% to less than 2%, significantly improving the accuracy and reliability of abnormal sound detection, and being able to stably detect the abnormal sound of equipment in a complex noise environment, realizing the efficient deployment and real-time monitoring of the model in the actual industrial environment.

[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0006] A method for detecting abnormal sound of equipment based on adaptive filtering and deep learning, the steps are as follows:

[0007] S1. Collect a large number of audio signals when the device is working properly, and filter the obtained audio signals when the device is working properly;

[0008] S2. Construct a joint training framework for adaptive filtering and residual convolutional network;

[0009] S3. Construct a three-channel feature fusion architecture: Dynamically fuse the time-domain waveform after adaptive filtering, MFCC time-frequency features, and semantic features extracted by the deep residual network;

[0010] S4. Add an adversarial enhanced lightweight residual network, generate adversarial samples through controllable noise injection, and expand the dataset while enhancing the generalization ability of the model;

[0011] S5. Construct an adaptive frequency band selection algorithm, and dynamically adjust the filter frequency response curve based on the feature importance output by the residual network;

[0012] S6. Filter the audio signals of the device collected in real time; and after being processed through steps S2 to S5, compare with the feature base during training to detect abnormalities.

[0013] Further, in S2, in the joint training framework, the output of the filter is used as the input of the deep learning model, and the loss function of the deep learning model is differentiated with respect to the filter parameters to achieve dynamic adjustment of the parameters.

[0014] Further, the update of the filter parameters follows the following formula:

[0015]

[0016] Among them, θ represents the filter parameters, α is the learning rate, and L is the loss function.

[0017] Further, in S3, during dynamic fusion, a weighted fusion strategy is adopted, and the fusion process follows the following formula:

[0018] F_fuse = w1·F_time + w2·F_mfcc + w3·F_semantics

[0019] Among them, F_fuse represents the fused features, F_time, F_mfcc, and F_semantics represent the time-domain waveform features, MFCC time-frequency features, and semantic features extracted by the deep residual network respectively, and w1, w2, and w3 are the corresponding weight coefficients, which are dynamically adjusted through the attention mechanism.

[0020] Furthermore, the dynamic adjustment of the attention mechanism is to embed a channel attention module in the residual block. Combining pruning and quantization techniques, the number of model parameters is compressed to 20% of the traditional residual network while maintaining a detection accuracy of >95%, meeting the requirements of edge devices with a latency <50 ms and power consumption <5 W.

[0021] Furthermore, in S4, the generation of adversarial samples follows the following formula:

[0022]

[0023] where \(x_{adv}\) represents the adversarial sample, \(x\) is the original sample, \(\epsilon\) is the noise intensity, is the gradient of the loss function with respect to the sample.

[0024] Furthermore, when generating adversarial samples, a channel attention compression mechanism is combined to remove redundant feature channels and reduce the number of model parameters.

[0025] Furthermore, the compression of the number of model parameters in the attention compression mechanism follows the following formula:

[0026] \(N_{compressed}=N\cdot(1 - r)\)

[0027] where \(N_{compressed}\) represents the number of model parameters after compression, \(N\) is the number of original model parameters, and \(r\) is the compression rate.

[0028] Furthermore, the strategy for generating adversarial samples adopts controllable noise injection and time-frequency masking techniques, achieving an F1-score of 94.3% with only 200 abnormal samples.

[0029] Furthermore, in S5, by analyzing the weight distribution of frequency information in the deep feature map, the key frequency band of the abnormal sound signal is determined, and then the frequency response curve of the filter is accurately adjusted; the adjustment of the filter frequency response curve follows the following formula;

[0030] \(H(f)=\frac{1}{1 + e^{-\beta(f - f_c)}}\)

[0031] where \(H(f)\) represents the filter frequency response curve, \(f\) is the frequency, \(\beta\) is the steepness coefficient, and \(f_c\) is the characteristic frequency.

[0032] Compared with the prior art, the present invention is a method for detecting abnormal sounds of equipment based on adaptive filtering and deep learning, and has the following beneficial effects: It solves the problem of signal distortion in complex scenarios of non-stationary noise, avoids the misfiltering of effective signals, and at the same time ensures the integrity of the key abnormal sound components of the signal after noise reduction. The signal-to-noise ratio is increased by ≥6dB. By using skip connections and channel attention mechanisms to strengthen key features, the advantages of each feature can be fully utilized, enabling the model to have more accurate perception and understanding of abnormal sound signals of different types and characteristics, effectively meeting the requirements for detecting abnormal sounds of equipment under complex working conditions. Furthermore, the deep learning model can more accurately learn the effective features of abnormal sound signals, effectively solving the misfiltering defect of traditional frequency-domain filtering when the abnormal sound-noise frequency bands overlap, accurately retaining the key frequency bands of abnormal sounds (200 - 8000Hz), with the retention rate of abnormal sound frequency bands >95%, and reducing the false alarm rate from 25% to below 2%, significantly improving the accuracy and reliability of abnormal sound detection, and being able to stably detect abnormal sounds of equipment even in complex noise environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a schematic diagram of the detection process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] To make the objectives, technical means for implementation, advantages, and the achievements of the objectives and effects of the present invention easily understood, the present invention will be further described below in conjunction with specific embodiments.

[0035] As Figure 1 shown, a method for detecting abnormal sounds of equipment based on adaptive filtering and deep learning is as follows:

[0036] S1. Collect a large number of audio signals of the equipment during normal operation, and perform filtering processing on the obtained audio signals during normal operation;

[0037] S2. Construct a joint training framework of adaptive filtering and residual convolutional network;

[0038] S3. Construct a three-channel feature fusion architecture: Dynamically fuse the time-domain waveform after adaptive filtering, MFCC time-frequency features, and semantic features extracted by the deep residual network;

[0039] S4. Add an adversarial enhanced lightweight residual network, generate adversarial samples through controllable noise injection, expand the dataset while enhancing the generalization ability of the model;

[0040] S5. Construct an adaptive frequency band selection algorithm, and dynamically adjust the filter frequency response curve based on the feature importance output by the residual network;

[0041] S6. Perform filtering processing on the audio signals of the equipment collected in real time; and after being processed through the steps of S2 to S5, compare with the feature database during training to detect abnormalities.

[0042] Further, in S2, in the joint training framework, the output of the filter is used as the input of the deep learning model, and the loss function of the deep learning model is differentiated with respect to the filter parameters to achieve dynamic adjustment of the parameters.

[0043] Further, the update of the filter parameters follows the following formula:

[0044]

[0045] Where θ represents the filter parameters, α is the learning rate, and L is the loss function. The present invention constructs a joint training framework of adaptive filtering and residual convolutional network, and through the backpropagation mechanism, synchronously optimizes the filter step parameters and network weights. In the joint training framework, the output of the filter is used as the input of the deep learning model, and the loss function of the deep learning model is differentiated with respect to the filter parameters to achieve dynamic adjustment of the parameters. This mechanism enables the filter to dynamically adjust the parameters according to the needs of the deep learning model, and at the same time, the deep learning model can also perform targeted learning based on the filtered signal characteristics. This solves the problem of signal distortion in non-stationary noise scenarios, avoids the misfiltering of effective signals, and at the same time ensures the integrity of the key abnormal sound components of the denoised signal, with the signal-to-noise ratio increased by ≥6dB. Furthermore, the deep learning model can more accurately learn the effective characteristics of the abnormal sound signal, improving the accuracy of abnormal sound detection.

[0046] Further, in S3, during dynamic fusion, a weighted fusion strategy is adopted, and the fusion process follows the following formula:

[0047] F_fuse = w1·F_time + w2·F_mfcc + w3·F_semantics

[0048] Where F_fuse represents the fused features, F_time, F_mfcc, and F_semantics respectively represent the time-domain waveform features, MFCC time-frequency features, and semantic features extracted by the deep residual network, and w1, w2, and w3 are the corresponding weight coefficients, which are dynamically adjusted through the attention mechanism. This architecture synchronously retains the filtered original time-domain waveform (capturing instantaneous impact sounds) and MFCC time-frequency features (characterizing steady-state frequency-domain patterns), constructs complementary inputs, and strengthens key features through skip connections and channel attention mechanisms, which can give full play to the advantages of each feature, enabling the model to have more accurate perception and understanding of different types and characteristics of abnormal sound signals, and effectively meeting the requirements of equipment abnormal sound detection under complex working conditions.

[0049] Furthermore, the dynamic adjustment of the attention mechanism is to embed a channel attention module (SE Block) in the residual block. Combining pruning and quantization techniques, the number of model parameters is compressed to 20% of the traditional residual network while maintaining a detection accuracy of >95%, meeting the requirements of edge devices with a latency <50ms and power consumption <5W. Combining the attention dynamic fusion strategy, complementary characterization of transient impact sounds (such as metal collisions) and steady-state frequency domain patterns (such as bearing friction) is achieved, and this design increases the complex abnormal sound detection rate by more than 15%.

[0050] Furthermore, in S4, the generation of adversarial samples follows the following formula:

[0051]

[0052] where \(x_{adv}\) represents the adversarial sample, \(x\) is the original sample, \(\epsilon\) is the noise intensity, is the gradient of the loss function with respect to the sample. Furthermore, when generating adversarial samples, a channel attention compression mechanism is combined to remove redundant feature channels and reduce the number of model parameters. Furthermore, the compression of the number of model parameters of the attention compression mechanism follows the following formula:

[0053] \(N_{compressed}=N\cdot(1 - r)\)

[0054] where \(N_{compressed}\) represents the number of model parameters after compression, \(N\) is the number of original model parameters, and \(r\) is the compression rate. Verified by experiments, this method can reduce the number of model parameters to less than 20% of the conventional residual network while ensuring a detection accuracy of >95%, meeting the requirements of edge devices for low power consumption and low latency, and realizing the efficient deployment and real-time monitoring of the model in the actual industrial environment.

[0055] Furthermore, the strategy for generating adversarial samples adopts controllable noise injection and time-frequency masking techniques, achieving a 94.3% F1-score with only 200 abnormal samples, breaking through the limitation of scarce data in industrial scenarios, while traditional methods require ≥1000 abnormal samples; moreover, the model volume is compressed to 18MB (30% of the traditional residual network), the inference latency of edge devices is ≤50ms, and the power consumption is reduced by 70%.

[0056] Furthermore, in S5, by analyzing the weight distribution of frequency information in the deep feature map, the key frequency band of the abnormal sound signal is determined, and then the frequency response curve of the filter is accurately adjusted; the adjustment of the filter frequency response curve follows the following formula;

[0057] \(H(f)=\frac{1}{1 + e^{-\beta(f - f_c)}}\)

[0058] Among them, H(f) represents the filter frequency response curve, f is the frequency, β is the steepness coefficient, and f_c is the characteristic frequency. This algorithm effectively solves the mis-filtering defect of traditional frequency-domain filtering when the abnormal sound-noise frequency bands overlap, accurately retains the key frequency bands of abnormal sound (200 - 8000 Hz), realizes the accurate retention of the abnormal sound frequency band, the retention rate of the abnormal sound frequency band > 95%, and reduces the false alarm rate from 25% to below 2%, significantly improving the accuracy and reliability of abnormal sound detection, and can also stably detect the abnormal sound of the device in a complex noise environment.

[0059] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes, equivalent substitutions, improvements, etc., which should all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. An abnormal sound detection method for devices based on adaptive filtering and deep learning, characterized in that, The steps are as follows: S1. Collect a large number of audio signals when the device is working properly, and filter the obtained audio signals when the device is working properly; S2. Construct a joint training framework for adaptive filtering and residual convolutional network; S3. Construct a three-channel feature fusion architecture: dynamically fuse the time-domain waveform after adaptive filtering, MFCC time-frequency features, and semantic features extracted by the deep residual network; S4. Add an adversarial enhanced lightweight residual network, generate adversarial samples through controllable noise injection, expand the dataset, and enhance the generalization ability of the model; S5. Construct an adaptive frequency band selection algorithm, and dynamically adjust the filter frequency response curve based on the feature importance output by the residual network; S6. Filter the audio signals of the device collected in real time; and after being processed through steps S2 to S5, compare with the feature database during training to detect anomalies.

2. The method for detecting abnormal sound of a device based on adaptive filtering and deep learning according to claim 1, wherein, In S2, in the joint training framework, the output of the filter is used as the input of the deep learning model, and the loss function of the deep learning model is differentiated with respect to the filter parameters to achieve dynamic adjustment of the parameters.

3. The method for detecting abnormal sound of a device based on adaptive filtering and deep learning according to claim 2, wherein The update of the filter parameters follows the following formula: where θ represents the filter parameters, α is the learning rate, and L is the loss function.

4. A method for detecting abnormal sounds of a device based on adaptive filtering and deep learning according to claim 1, characterized in that, In S3, during dynamic fusion, a weighted fusion strategy is adopted, and the fusion process follows the following formula: F_fuse = w1·F_time + w2·F_mfcc + w3·F_semantics where F_fuse represents the fused features, F_time, F_mfcc, and F_semantics represent the time-domain waveform features, MFCC time-frequency features, and semantic features extracted by the deep residual network respectively, and w1, w2, and w3 are the corresponding weight coefficients, which are dynamically adjusted through the attention mechanism.

5. The method for detecting abnormal sound of a device based on adaptive filtering and deep learning according to claim 4, wherein The dynamic adjustment of the attention mechanism is to embed a channel attention module in the residual block, and combine pruning and quantization techniques to compress the model parameter amount to 20% of the traditional residual network, while maintaining a detection accuracy of >95%, meeting the requirements of edge devices with a latency <50ms and power consumption <5W.

6. The method for detecting abnormal sound of a device based on adaptive filtering and deep learning according to claim 1, characterized in that, In S4, the generation of adversarial samples follows the following formula: Among them, \(x_{adv}\) represents the adversarial example, \(x\) is the original example, and \(\epsilon\) is the noise intensity. is the gradient of the loss function with respect to the example.

7. A method for detecting abnormal sounds of a device based on adaptive filtering and deep learning according to claim 6, characterized in that, When generating adversarial samples, a channel attention compression mechanism is combined to remove redundant feature channels and reduce the model parameter amount.

8. A method for detecting abnormal sounds of a device based on adaptive filtering and deep learning according to claim 6, characterized in that, The compression of the model parameter amount of the attention compression mechanism follows the following formula: N_cp,[ressed = N·(1 - r) where N_compressed represents the compressed model parameter amount, N is the original model parameter amount, and r is the compression rate.

9. A method for detecting abnormal sounds of a device based on adaptive filtering and deep learning according to claim 6, characterized in that, The strategy for generating adversarial samples adopts controllable noise injection and time-frequency masking techniques, and achieves a 94.3% F1-score with only 200 abnormal samples.

10. The method for detecting abnormal sound of a device based on adaptive filtering and deep learning according to claim 1, characterized in that, In S5, by analyzing the weight distribution of frequency information in the deep feature map, the key frequency band of the abnormal sound signal is determined, and then the filter frequency response curve is accurately adjusted; the adjustment of the filter frequency response curve follows the following formula; H(f) = 1 / (1 + e^{-β(f - f_c)}) where H(f) represents the filter frequency response curve, f is the frequency, β is the steepness coefficient, and f_c is the characteristic frequency.