Abnormal Sound Detection Method, Device, Equipment and Storage Medium

Through the feature capture unit and densely connected delay neural network layer in the object detection model, the existing antonym detection technology has solved the problems of high accuracy, low cost and strong generalization in complex acoustic environments and diversified product lines, realizing high-precision antonym detection and providing interpretability support.

CN120089161BActive Publication Date: 2025-07-22GOERTEK INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510560155.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-22
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

Existing antonym detection technology is difficult to achieve high accuracy, low cost and strong generalization in complex acoustic environments and diversified product lines. Traditional methods have problems such as strong feature dependence, high computational cost, poor interpretability and translation invariance.

Method used

Feature capture units in the object detection model are used for feature extraction, and densely connected time-delay neural network layer and compression excitation module are used, combined with a multi-grained context perception mechanism to perform different sound detection, establish timing feature modeling and enhance time-frequency domain features.

Benefits of technology

It significantly improves detection accuracy and robustness, enables high-precision antonym detection in complex acoustic environments and diverse product lines, and provides interpretability support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089161B_ABST
    Figure CN120089161B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, device and storage medium for abnormal sound detection. The abnormal sound detection method is disclosed, including: extracting features from the two-dimensional feature representation of the target detection audio through the feature capture unit in the target detection model to obtain the audio features to be detected of the target detection audio. The feature capture unit is composed of multiple feature extraction modules and at least one squeeze-and-excitation module, and the feature extraction module is composed of multiple densely connected delay neural network layers; inputting the audio features to be detected into the detection and classification unit in the target detection model for abnormal sound detection to obtain the audio category to which the target detection audio belongs; obtaining the audio detection result of the target detection audio according to the audio category to which the target detection audio belongs. Through the above method, the modeling of time series features is optimized, a multi-granularity context awareness mechanism is established, the enhancement of time-frequency domain features is strengthened, the detection accuracy and robustness of the model are significantly improved, and at the same time its generalization ability is greatly enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of acoustic technologies, and in particular to a method, apparatus, device, and storage medium for detecting abnormal sounds. Background Art

[0002] In the field of acoustic product manufacturing, abnormal sound detection is a core link to ensure the acoustic quality of products, directly determining the user experience and market competitiveness. Traditional detection methods rely on the subjective judgment of human listeners, with inherent defects such as low efficiency, high cost, and poor consistency. Although current automated detection technologies attempt to improve efficiency through two-dimensional spectrogram analysis and convolutional neural networks (CNNs), they still face significant bottlenecks: two-dimensional spectrogram analysis methods rely on manual feature engineering, are prone to ignoring abnormal sound-sensitive information, have limited resolution ability in complex noise scenarios, and high-resolution time-frequency analysis leads to a sharp increase in computational burden; traditional CNN methods, although having the ability to extract local features, are insufficient in modeling the long-term dependence of sound signals, and their translational invariance may weaken the detection accuracy for abnormal sound time-series sensitive scenarios. Existing technologies are difficult to achieve high-precision, low-computation-cost, strong generalization, and interpretable abnormal sound detection in complex acoustic environments and diverse product lines, restricting large-scale deployment and real-time diagnosis capabilities in industrial scenarios.

[0003] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of this application is to provide a method, apparatus, device, and storage medium for detecting abnormal sounds, aiming to solve the technical problem that existing abnormal sound detection technologies cannot achieve high-precision, low-cost, and strong generalization abnormal sound detection in complex acoustic environments and diverse product lines.

[0005] To achieve the above object, this application proposes an abnormal sound detection method, and the method includes:

[0006] Performing feature extraction on a two-dimensional feature representation corresponding to a target detection audio through a feature capture unit in a target detection model to obtain a to-be-detected audio feature of the target detection audio, where the feature capture unit is composed of a plurality of feature extraction modules and at least one squeeze-and-excitation module, and the feature extraction module is composed of a plurality of densely connected delay neural network layers;

[0007] Inputting the to-be-detected audio feature into a detection and classification unit in the target detection model to perform abnormal sound detection, and obtaining a category to which the audio of the target detection audio belongs;

[0008] Obtaining an audio detection result of the target detection audio according to the category to which the audio of the target detection audio belongs.

[0009] In one embodiment, the step of extracting the to-be-detected audio features of the target detection audio by using a feature capture unit in the target detection model from the two-dimensional feature representation corresponding to the target detection audio includes:

[0010] Input the two-dimensional feature representation corresponding to the target detection audio into the feature capture unit in the target detection model, and perform temporal feature extraction on the two-dimensional feature representation through a first feature extraction module of the feature capture unit to obtain the target temporal features corresponding to the target detection audio;

[0011] Input the target temporal features into a squeeze-and-excitation module in the feature capture unit, and perform weight adjustment on the target temporal features through the squeeze-and-excitation module to obtain the weighted audio features of the target detection audio;

[0012] According to the weighted audio features and a second feature extraction module in a plurality of feature extraction modules, perform feature extraction on the target detection audio to obtain the to-be-detected audio features corresponding to the target detection audio.

[0013] In one embodiment, the densely connected delay neural network layer includes a feedforward neural network and at least one delay neural network;

[0014] The step of performing temporal feature extraction on the two-dimensional feature representation through a first feature extraction module of the feature capture unit to obtain the target temporal features corresponding to the target detection audio includes:

[0015] Input the two-dimensional feature representation into the first feature extraction module of the feature capture unit, and perform feature transformation on the two-dimensional feature representation through a first feedforward neural network in the first feature extraction module to obtain the first audio features of the target detection audio, where the first feedforward neural network is a basic unit of a first densely connected delay neural network layer in a plurality of densely connected delay neural network layers;

[0016] Input the first audio features into the delay neural networks in the first densely connected delay neural network layer respectively, perform temporal feature extraction on the first audio features through the delay neural networks, and obtain the target temporal features corresponding to the target detection audio according to the feature extraction results of the delay neural networks and the hierarchical structure of the first feature extraction module.

[0017] In one embodiment, before the step of extracting the to-be-detected audio features of the target detection audio by using a feature capture unit in the target detection model from the two-dimensional feature representation corresponding to the target detection audio, it further includes:

[0018] In response to a heterophony detection instruction, perform signal processing on the target detection audio to obtain the filter bank features of the target detection audio;

[0019] Input the filter bank features into the front-end convolutional module in the target detection model, perform local feature extraction on the filter bank features through the front-end convolutional module, and obtain the two-dimensional feature representation corresponding to the target detection audio according to the extraction result.

[0020] In one embodiment, the step of performing signal processing on the target detection audio in response to a heterophony detection instruction to obtain the filter bank features of the target detection audio includes:

[0021] In response to a heterophony detection instruction, perform frame segmentation processing on the target detection audio according to a preset window to obtain multiple frames of audio signals corresponding to the target detection audio;

[0022] Perform windowing processing on each audio signal respectively to obtain the windowed signals corresponding to each audio signal;

[0023] Perform power spectrum calculation on each windowed signal respectively to determine the linear frequency power spectrum of each windowed signal;

[0024] Perform Mel filter bank processing on the linear frequency power spectrum of each windowed signal respectively, and obtain the filter bank features of each windowed signal according to the processing result;

[0025] Determine the filter bank features of the target detection audio according to the filter bank features of each windowed signal.

[0026] In one embodiment, before the step of performing feature extraction on the two-dimensional feature representation corresponding to the target detection audio through the feature capture unit in the target detection model to obtain the audio features to be detected of the target detection audio, it further includes:

[0027] Perform data augmentation according to multiple sample audios and the heterophony label information of each sample audio to obtain multiple model training samples and the heterophony label information of each model training sample;

[0028] Perform signal processing on each model training sample to obtain the filter bank features of each model training sample;

[0029] Set the parameters of the loss function according to a preset interval factor and a preset scaling factor to determine the target loss function;

[0030] The initial detection model is trained according to the target loss function, the target gradient descent optimizer, the target scheduler, the filter bank features of each model training sample, and the mispronunciation label information of each model training sample to obtain a target detection model. The initial detection model is composed of an initial front-end convolution module, a plurality of initial feature extraction modules, at least one initial squeeze-and-excitation module, and an initial detection classification unit.

[0031] In one embodiment, the step of performing data augmentation on multiple sample audios and the mispronunciation label information of each sample audio to obtain multiple model training samples and the mispronunciation label information of each model training sample includes:

[0032] Randomly select in a preset set of rate ratios to determine a target rate ratio;

[0033] Audio sampling is performed on each sample audio according to the target rate ratio to obtain multiple extended audios and the mispronunciation label information of each extended audio;

[0034] Multiple model training samples and the mispronunciation label information of each model training sample are obtained according to the extended audios and the mispronunciation label information of each extended audio, and the sample audios and the mispronunciation label information of each sample audio.

[0035] In addition, to achieve the above object, the present application also proposes a mispronunciation detection device. The mispronunciation detection device includes: an extraction module, configured to extract features of a two-dimensional feature representation corresponding to a target detection audio through a feature capture unit in a target detection model to obtain a to-be-detected audio feature of the target detection audio. The feature capture unit is composed of a plurality of feature extraction modules and at least one squeeze-and-excitation module, and the feature extraction module is composed of a plurality of densely connected delay neural network layers;

[0036] A detection module, configured to input the to-be-detected audio feature into a detection classification unit in the target detection model for mispronunciation detection to obtain the audio category to which the target detection audio belongs;

[0037] A processing module, configured to obtain an audio detection result of the target detection audio according to the audio category to which the target detection audio belongs.

[0038] In addition, to achieve the above object, the present application also proposes a mispronunciation detection device. The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The computer program is configured to implement the steps of the mispronunciation detection method as described above.

[0039] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the abnormal sound detection method described above are implemented.

[0040] In addition, to achieve the above object, the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the abnormal sound detection method described above are implemented.

[0041] The present application provides an abnormal sound detection method. In the present application, a feature capture unit in a target detection model extracts features from a two-dimensional feature representation corresponding to a target detection audio to obtain a to-be-detected audio feature of the target detection audio. The feature capture unit is composed of multiple feature extraction modules and at least one squeeze-and-excitation module, and the feature extraction module is composed of multiple densely connected delay neural network layers. The to-be-detected audio feature is input into a detection and classification unit in the target detection model for abnormal sound detection to obtain the audio category to which the target detection audio belongs. An audio detection result of the target detection audio is obtained according to the audio category to which the target detection audio belongs. By the above method, the time-series feature modeling is optimized, a multi-granularity context awareness mechanism is established, the time-frequency domain feature enhancement is strengthened, the detection accuracy and robustness of the model are significantly improved, and at the same time, its generalization ability is greatly enhanced. It can achieve high-precision abnormal sound detection at low computational cost in complex acoustic environments and diverse product lines, and provide interpretability support for the detection results. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0043] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0044] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the abnormal sound detection method of the present application;

[0045] Figure 2 It is a schematic diagram of the FCM network structure provided for Embodiment 1 of the present application;

[0046] Figure 3 It is a schematic flowchart provided for Embodiment 2 of the abnormal sound detection method of the present application;

[0047] Figure 4 Schematic diagram of the network structure of the feature extraction module provided in the second embodiment of the present application;

[0048] Figure 5 Schematic diagram of the network structure of the SE module provided in the second embodiment of the present application;

[0049] Figure 6 Schematic flow chart provided in the third embodiment of the abnormal sound detection method of the present application;

[0050] Figure 7 Brief schematic flow chart of the abnormal sound detection method provided in the third embodiment of the present application;

[0051] Figure 8 Schematic diagram of the module structure of the abnormal sound detection device in the embodiment of the present application;

[0052] Figure 9 Schematic diagram of the device structure of the hardware operating environment involved in the abnormal sound detection method in the embodiment of the present application.

[0053] The realization of the purpose, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners

[0054] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0055] In order to better understand the technical solutions of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0056] The main solution of the embodiment of the present application is: the feature extraction unit in the target detection model extracts features from the two-dimensional feature representation corresponding to the target detection audio to obtain the audio features to be detected of the target detection audio. The feature extraction unit is composed of multiple feature extraction modules and at least one squeeze-and-excitation module, and the feature extraction module is composed of multiple densely connected delay neural network layers; the audio features to be detected are input into the detection and classification unit in the target detection model for abnormal sound detection to obtain the audio category to which the target detection audio belongs; the audio detection result of the target detection audio is obtained according to the audio category to which the target detection audio belongs.

[0057] In the field of acoustic product manufacturing, the acoustic quality of products is one of the core factors determining user experience and market competitiveness. Acoustic products such as headphones, speakers, and microphones need to undergo a strict abnormal sound detection process before leaving the factory to ensure that there are no defects such as noise, distortion, abnormal background noise, or structural resonance during operation. Traditional detection methods highly rely on the subjective judgment of human listeners. By repeatedly playing standardized audio samples and relying on the human ear to identify abnormalities. However, with the exponential growth of the shipment volume of acoustic products and the accelerating speed of product iteration, the disadvantages of manual detection have become increasingly prominent: low detection efficiency, rising labor costs, and the human ear is vulnerable to fatigue, individual hearing differences, and environmental noise interference, resulting in difficulty in ensuring detection consistency. Although some enterprises have introduced automated detection equipment, the existing technology is still difficult to achieve high-precision and strong generalization of abnormal sound determination in complex acoustic scenarios and diverse product lines.

[0058] Existing abnormal sound detection technologies for acoustic products mainly focus on two-dimensional spectrogram analysis and traditional CNN. Two-dimensional spectrogram analysis has the following disadvantages: strong dependence on features, relying on manually designed feature extraction (such as Mel spectrogram, wavelet transform), which may ignore key information sensitive to abnormal sounds, and feature selection requires domain experience; loss of dynamic information, static spectrograms are difficult to capture the temporal dynamic characteristics of sound signals; sensitive to noise, the resolution ability of artificial features may significantly decrease under complex background noise; high computational cost: high-resolution time-frequency analysis will greatly increase the computational burden. CNN methods have the following disadvantages: insufficient temporal modeling, traditional CNN is good at local spatial feature extraction, but has limited ability to model long-term dependence relationships of sound signals; large data requirements, a large number of labeled abnormal samples are needed to train a robust model, while abnormal samples are usually scarce in actual industrial scenarios; poor interpretability, the black-box nature makes it difficult to locate the specific frequency band / time domain features of abnormal sounds, which is not conducive to fault diagnosis and analysis; interference of translational invariance, the inherent translational invariance of CNN may weaken the detection ability for scenarios sensitive to the timing of abnormal sound occurrence.

[0059] This application optimizes the temporal feature modeling and establishes a multi-granularity context awareness mechanism, strengthens the enhancement of time-frequency domain features, significantly improves the detection accuracy and robustness of the model, and at the same time greatly enhances its generalization ability. It can achieve high-precision abnormal sound detection at low computational cost in complex acoustic environments and diverse product lines, and provide interpretability support for the detection results.

[0060] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, an abnormal sound detection device, etc. that can implement the above functions. Hereinafter, taking the abnormal sound detection device as an example, this embodiment and the following embodiments will be described.

[0061] Based on this, an embodiment of the present application provides a heterophony detection method. Referring to Figure 1 , Figure 1 is a schematic flowchart of the first embodiment of the heterophony detection method of the present application.

[0062] In this embodiment, the heterophony detection method includes steps S10 to S30:

[0063] Step S10, extracting features from the two-dimensional feature representation corresponding to the target detection audio through the feature capture unit in the target detection model to obtain the audio features to be detected of the target detection audio. The feature capture unit is composed of multiple feature extraction modules and at least one squeeze-and-excitation module, and each feature extraction module is composed of multiple densely connected time-delay neural network layers.

[0064] It should be noted that the target detection audio refers to the audio to be detected generated by the acoustic product that needs to perform heterophony detection. Heterophony is an abnormal sound signal that is unexpected and not allowed by design during the normal operation of the acoustic product. Its manifestation is significantly different from the normal acoustic output. Heterophony includes, but is not limited to, types such as noise, crackling sound, abnormal background noise, and structural resonance.

[0065] It can be understood that the target detection model is composed of an FCM (Front-end Convolutional Module), a feature capture unit, and a detection and classification unit. The feature capture unit uses D-TDNN as the backbone network, which is composed of multiple feature extraction modules and at least one SE module (Squeeze-and-Excitation Module). Each feature extraction module is composed of a series of D-TDNN (Densely Connected Time-Delay Neural Network) layers. The basic unit of each D-DTNN layer is composed of an FNN (Feedforward Neural Network) and a TDNN (Time-Delay Neural Network) to form a dense connection mechanism. The core components of the detection and classification unit include, but are not limited to, fully connected layers and activation functions, and are used to output the heterophony detection classification result.

[0066] In a specific implementation, an SE module may exist in each D-TDNN layer of each feature extraction module. The SE module may be located before the FNN or after the TDNN; or, an SE module exists between each feature extraction module; or, an SE module is added after the last feature extraction module; or, the SE may be located at other positions in the feature capture unit, that is, the SE module can be embedded in any intermediate layer such as a CNN or a TDNN. Without modifying the backbone network structure, by reducing the channel weights in the low signal-to-noise ratio frequency band, the stability of the model in a complex environment is improved. Therefore, the specific location and specific number of the SE module in this embodiment are not limited.

[0067] It should be noted that this embodiment is committed to solving the problems of local feature capture limitations and insufficient long-term acoustic context modeling existing in two-dimensional spectrograms and traditional CNN methods in the detection of abnormal sounds in acoustic products. Due to the fixed frequency resolution of traditional two-dimensional spectrograms, the details of high-frequency weak abnormal sounds are lost, and the local convolution kernels of CNNs are difficult to model the dynamic propagation characteristics of abnormal sounds on the time axis. In addition, the differences in the fundamental frequencies of acoustic products across product lines easily cause the traditional CNN model to have spectrum pattern mismatches due to insufficient generalization of artificial features, especially in low signal-to-noise ratio scenarios, the sensitivity to abnormal harmonic components significantly decreases.

[0068] It can be understood that after obtaining the target detection audio of the acoustic product, the target detection audio is processed to extract the FBank (Mel Filter Bank) features corresponding to the target detection audio, and the FBank features of the target detection audio are input into the target detection model. After the FCM in the target detection model extracts the features, the feature map output by the FCM is flattened in the channel dimension and the frequency dimension to obtain the two-dimensional feature representation of the target detection audio.

[0069] In a specific implementation, the two-dimensional feature representation is sent to the feature capture unit, and the feature extraction module and the SE module with a series of D-TDNN layers in the feature capture unit are used to further process the two-dimensional feature representation, so as to obtain the audio feature to be detected of the target detection audio.

[0070] It should be noted that during feature extraction, a D-TDNN using a dense connection mechanism is adopted. Dense connection cascades the outputs of all previous layers as the input of the current layer. The chain-like feature fusion of dense connection enables the subsequent layers to dynamically select features at different abstraction levels, and the reduction in the number of parameters directly reduces the memory occupation during inference. The SE module extracts channel-level statistics through global average pooling to capture the global distribution characteristics of the feature map. In the Excitation stage, a fully connected layer is used to model the non-linear relationship between channels.

[0071] In a feasible implementation, before step S10, steps A11 to A13 may further be included:

[0072] Step A11, in response to a strange sound detection instruction, perform signal processing on the target detection audio to obtain the filter bank features of the target detection audio.

[0073] It should be noted that the strange sound detection instruction is used to trigger the strange sound detection process of the acoustic product to determine whether there is a strange sound in the audio emitted by the acoustic product. The strange sound detection instruction can be actively initiated by the strange sound detection device when detecting the acoustic product, or can be initiated by the user to the strange sound detection device. This embodiment does not limit this.

[0074] It can be understood that when receiving the strange sound detection instruction, obtain the target detection audio of the acoustic product, and perform processing such as frame division - fast Fourier transform - Mel filtering on the target detection audio, so as to obtain the filter bank features of the target detection audio. In this embodiment, the filter bank features of the target detection audio include the FBank features corresponding to each frame.

[0075] Step A12, input the filter bank features into the front-end convolutional module in the target detection model, and perform local feature extraction on the filter bank features through the front-end convolutional module, and obtain the two-dimensional feature representation corresponding to the target detection audio according to the extraction result.

[0076] It should be noted that in this embodiment, input the filter bank features into the FCM. The FCM integrates residual connections and is composed of multiple two-dimensional convolutional blocks containing residual connections, encodes the acoustic features in the time-frequency domain, and utilizes the high-resolution time-frequency details. Its specific architecture is as Figure 2 shown. The front-end convolutional module FCM integrates 4 residual modules (residual modules 1 to 4). Each residual module passes the original input through a skip connection to alleviate gradient disappearance and support the module to expand to deeper layers. Use the FCM to achieve high-frequency time-frequency detail capture. The local receptive field characteristic of the two-dimensional convolution enables it to slide simultaneously in the time domain and the frequency domain, and directly model the local pattern of the acoustic features; the residual connection optimizes the training. In the deep convolutional structure, the residual skip connection alleviates the gradient disappearance problem through the identity mapping, making the network depth scalable. Figure 2 Among them, [1, F, T] are the filter bank features of the target detection audio, and [D, 1, T] is the feature map output by the FCM.

[0077] It can be understood that there is a feature map output by the FCM in the extraction result, and the feature map is flattened along the channel dimension and the frequency dimension, so as to obtain the two-dimensional feature representation of the target detection audio.

[0078] In a feasible implementation, step A11 may further include steps B11 to B15:

[0079] Step B11, in response to the abnormal sound detection instruction, perform frame segmentation on the target detection audio according to a preset window to obtain multiple frames of audio signals corresponding to the target detection audio.

[0080] It should be noted that since the target detection audio is a continuous speech signal, in order to improve the processing efficiency, when the abnormal sound detection instruction is received, the target detection audio is segmented into short-time segments according to a preset window, so as to obtain multiple frames of audio signals corresponding to the target detection audio. In this embodiment, the length of the preset window is set to 25 ms to ensure capturing the short-time characteristics of the audio; the frame shift is 10 ms, and the adjacent frames overlap by 15 ms.

[0081] Step B12, perform windowing processing on each audio signal respectively to obtain windowed signals corresponding to each audio signal respectively.

[0082] It should be noted that a Hamming window is applied to each frame of audio signal, and the windowed audio signal is the windowed signal, so as to reduce the spectral leakage of the audio signal and make the two ends of the frame smoothly attenuate to zero.

[0083] Step B13, calculate the power spectrum of each windowed signal respectively to determine the linear frequency power spectrum of each windowed signal.

[0084] It should be noted that after windowing is completed, perform a fast Fourier transform (FFT) on each frame of windowed signal to obtain a complex spectrum, and calculate the square of the amplitude spectrum, so as to obtain the linear frequency power spectrum of each frame of windowed signal.

[0085] Step B14, perform Mel filter bank processing on the linear frequency power spectrum of each windowed signal respectively, and obtain the filter bank features of each windowed signal according to the processing results.

[0086] It should be noted that perform Mel filter bank processing on the linear frequency power spectrum of each frame of windowed signal respectively. Pass the linear frequency power spectrum through a Mel filter bank composed of a preset number (for example, 80) of triangular filters, so as to obtain the Mel energy value of each frame of windowed signal. Take the natural logarithm of the Mel energy value of each frame of windowed signal, so as to output the filter bank features (i.e., FBank features) of each frame of windowed signal. When the number of triangular filters is 80, 80-dimensional FBank features will be obtained.

[0087] Step B15, determine the filter bank features of the target detection audio according to the filter bank features of each windowed signal.

[0088] It should be noted that the FBank features of each framed windowed signal are aggregated, and the aggregated result is the filter bank feature of the target detection audio.

[0089] Step S20: Input the audio features to be detected into the detection and classification unit in the target detection model for abnormal sound detection, and obtain the audio category to which the target detection audio belongs.

[0090] It should be noted that the audio category refers to the category label to which the target detection audio is classified, including but not limited to normal audio, noise, broken sound, etc. The detection and classification unit is the decision-making layer in the abnormal sound detection model, which is responsible for mapping the audio features to be detected into the category mapping space to complete the abnormal sound detection task. In this embodiment, the fully connected layer in the detection and classification unit is responsible for mapping the features to the category space to generate the scores of each category, and the activation function in the detection and classification unit is used to convert the scores into a probability distribution to support classification decisions.

[0091] Step S30: Obtain the audio detection result of the target detection audio according to the audio category to which the target detection audio belongs.

[0092] It should be noted that after identifying the audio category to which the target detection audio belongs through the target detection model, it is possible to clarify whether there is an abnormal sound in the audio to be detected, and when there is an abnormal sound, the specific category of the abnormal sound, so as to obtain the final audio detection result. When there is an abnormal sound in the target detection audio, the acoustic product can be specifically improved and optimized according to the specific audio category to ensure the product quality.

[0093] This embodiment provides an abnormal sound detection method. In this embodiment, the feature capture unit in the target detection model extracts features from the two-dimensional feature representation corresponding to the target detection audio to obtain the audio features to be detected of the target detection audio. The feature capture unit is composed of multiple feature extraction modules and at least one squeeze-and-excitation module, and the feature extraction module is composed of multiple densely connected time-delay neural network layers; the audio features to be detected are input into the detection and classification unit in the target detection model for abnormal sound detection to obtain the audio category to which the target detection audio belongs; the audio detection result of the target detection audio is obtained according to the audio category to which the target detection audio belongs. By the above method, the temporal features are modeled and optimized, and a multi-granularity context awareness mechanism is established to strengthen the enhancement of time-frequency domain features, significantly improving the detection accuracy and robustness of the model. At the same time, its generalization ability is greatly enhanced, and high-precision abnormal sound detection can be achieved at low computational cost in complex acoustic environments and diverse product lines, and provide interpretability support for the detection results.

[0094] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3 , step S10, the abnormal sound detection method further includes steps S11 to S13:

[0095] Step S11, input the two-dimensional feature representation corresponding to the target detection audio into the feature capture unit in the target detection model, and perform temporal feature extraction on the two-dimensional feature representation through the first feature extraction module of the feature capture unit to obtain the target temporal feature corresponding to the target detection audio.

[0096] It should be noted that the first feature extraction module refers to the first feature extraction module for receiving the two-dimensional feature representation of the target detection audio. Temporal feature extraction is performed on the two-dimensional feature representation through a series of D-TDNN layers in the first feature extraction module, and the feature output by the TDNN layer in the last D-TDNN layer of the first feature extraction module is the target temporal feature corresponding to the target detection audio. When there are multiple TDNN layers in each D-TDNN layer, the target temporal feature is the result obtained by multiplying the outputs of multiple TDNN layers in the last D-TDNN layer.

[0097] It can be understood that in this embodiment, there are three feature extraction modules in the feature capture unit, and the number of layers of each module is 12 / 24 / 16 layers. The growth rate k of each module is reduced from 64 to 32 to construct a narrower D-TDNN layer, and the increase in the number of parameters caused by the increase in depth is offset by reducing the number of channels per layer; an input TDNN layer with a downsampling rate of 1 / 2 is added in front of the D-TDNN backbone to accelerate the calculation. The network structure of the feature capture unit with D-TDNN as the backbone network is as Figure 4 shown, Figure 4 The dense connection blocks _1 to 3 in represent three feature extraction modules. Each D-TDNN layer is composed of a feedforward neural network FNN and two time-delay neural networks TDNN. At the same time, a dense connection method is adopted. After the results of the two TDNN outputs in each layer are multiplied, they are sent to the next layer as the input of the next layer. The chain feature fusion of dense connection enables the backend layer to dynamically select features at different abstraction levels, and the reduction in the number of parameters directly reduces the memory occupancy during inference. Compared with using two feature extraction modules with 6 and 12 layers respectively for each feature extraction module, the network structure of this embodiment can form deeper feature abstractions.

[0098] It can be understood that since the D-TDNN backbone network in this embodiment adopts a dense connection mechanism, it has parameter efficiency, and while requiring fewer parameters than traditional TDNN, it can achieve better performance.

[0099] Step S12: Input the target temporal feature into the squeeze-and-excitation module in the feature capture unit, and adjust the weight of the target temporal feature through the squeeze-and-excitation module to obtain the weighted audio feature of the target detection audio.

[0100] It should be noted that in this embodiment, the case where there is an SE module in each D-TDNN layer is taken as an example for illustration, and the SE module is located after each TDNN layer. Input the target temporal feature output from the TDNN layer of the last D-TDNN layer of the first feature extraction module into the SE module. Use the SE module to learn the feature weights according to the loss function through Squeeze, Excitation, and Scale operations, so that the effective feature maps have larger weights, enhance the channel features of the input feature maps, and combine the learned channel attention information with the input feature maps. At this time, the result output by the SE module is the weighted audio feature of the target detection audio.

[0101] Step S13: Extract the features of the target detection audio according to the weighted audio feature and the second feature extraction module in the multiple feature extraction modules to obtain the audio feature to be detected corresponding to the target detection audio.

[0102] It should be noted that the second feature extraction module refers to all other feature extraction modules except the first feature extraction module among the multiple feature extraction modules. In this embodiment, there are two feature extraction modules in the second feature extraction module.

[0103] It can be understood that when the weighted audio feature is input into the second feature extraction module, in the case where there is an SE module in each D-TDNN layer and the SE module is located after the TDNN layer, the result output by the TDNN layer of the last D-TDNN layer of this feature extraction module will be sent to the SE module. At this time, the result output by the SE module will enter the second feature extraction module in the second feature extraction module, and the above steps are repeated until all modules in the feature capture unit have completed their functions. The result output by the last module in the feature capture unit is the audio feature to be detected. In this embodiment, if there is an SE module in each D-TDNN layer and the SE module is located after the TDNN layer, then the result generated by the last SE module is the audio feature to be detected corresponding to the target detection audio.

[0104] In a feasible implementation manner, the dense connection time-delay neural network layer includes a feedforward neural network and at least one time-delay neural network. In step S11: performing temporal feature extraction on the two-dimensional feature representation through the first feature extraction module of the feature capture unit to obtain the target temporal feature corresponding to the target detection audio, which may include steps C11~C12:

[0105] Step C11, inputting the two-dimensional feature representation into the first feature extraction module of the feature capture unit, and performing feature transformation on the two-dimensional feature representation through the first feedforward neural network in the first feature extraction module to obtain the first audio feature of the target detection audio. The first feedforward neural network is the basic unit of the first dense connection time-delay neural network layer in multiple dense connection time-delay neural network layers.

[0106] It should be noted that the first dense connection time-delay neural network layer refers to the D-TDNN layer located at the first position in the first feature extraction module for receiving the two-dimensional feature representation. The first feedforward neural network is the FNN in the first D-TDNN layer. Inputting the two-dimensional feature representation into the first FNN, the first FNN extracts global features and enhances the non-linear expression ability, thereby obtaining the first audio feature of the target detection audio and providing a high-dimensional representation for the subsequent TDNN layer.

[0107] Step C12, inputting the first audio feature into the time-delay neural network in the first dense connection time-delay neural network layer respectively, performing temporal feature extraction on the first audio feature through the time-delay neural network, and obtaining the target temporal feature corresponding to the target detection audio according to the feature extraction result of the time-delay neural network and the hierarchical structure of the first feature extraction module.

[0108] It should be noted that inputting the first audio feature output by the first FNN into the TDNN in the first D-TDNN layer, the TDNN slides the convolutional kernel in the time dimension to capture local context (such as the dependence relationship between adjacent frames), and expands the receptive field through dilated convolution. When there is only one TDNN, the result output by the TDNN is sent to the SE module (in the case where there is an SE module in each D-TDNN layer and the SE module is located after the TDNN layer), or directly input into the next D-TDNN layer in the first feature extraction module; repeat the above steps until all D-TDNN layers in the first feature extraction module have completed their operations.

[0109] It can be understood that when there are multiple TDNN layers in each D-TDNN layer, the result output by the FNN is input into multiple TDNNs respectively, and the results output by each TDNN are multiplied and then sent to the next structure.

[0110] In a specific implementation, taking the example that an SE module is constructed in each D-TDNN layer for illustration, as Figure 5 shown, the SE module first performs spatial feature compression on the feature map, realizes global average pooling in the spatial dimension to obtain a feature map of 1×1×C, learns through the FC fully connected layer to obtain a feature map with channel attention, with a dimension of 1×1×C, multiplies the feature map with channel attention of 1×1×C and the original input feature map of H×W×C channel by channel with weight coefficients, and finally outputs a feature map with channel attention, where : Parameter-free compression and preliminary excitation to generate channel statistics; : Parameterized learning of channel weights to dynamically enhance key features; : Apply the weights to the original features to complete the adaptive adjustment.

[0111] It should be noted that in this embodiment, a direct connection is established between the inputs of two consecutive D-TDNN layers. The mathematical expression of the l-th layer D-TDNN is: . Among them, represents the input of the feature extraction module, represents the output of the l-th layer, represents the non-linear transformation of this layer. In this embodiment, the depth of the D-TDNN network is increased. Through the direct connection between layers, the network depth is allowed to increase significantly, improving the ability to model long-term acoustic context; the direct connection between layers enables the bypassing of deep non-linear transformations during backpropagation, alleviating the problem of gradient disappearance, and at the same time controlling the complexity by reducing the number of filter channels in each layer.

[0112] This embodiment provides a method for detecting abnormal sounds. In this embodiment, the two-dimensional feature representation corresponding to the target detection audio is input into the feature capture unit in the target detection model. The first feature extraction module in the feature capture unit performs temporal feature extraction on the two-dimensional feature representation to obtain the target temporal feature corresponding to the target detection audio; the target temporal feature is input into the compression excitation module in the feature capture unit, and the compression excitation module adjusts the weights of the target temporal feature to obtain the weighted audio feature of the target detection audio; according to the weighted audio feature and the second feature extraction module in multiple feature extraction modules, feature extraction is performed on the target detection audio to obtain the audio feature to be detected corresponding to the target detection audio. Through the above method, feature extraction is performed through densely connected D-TDNN layers and SE modules, realizing accurate multi-scale temporal-frequency joint modeling, capable of dynamically adaptive feature enhancement and noise suppression, not only improving the learning ability and efficiency of the model, but also improving the generalization ability of the model and the recognition ability of key features.

[0113] Based on the first embodiment and / or the second embodiment of the present application, in the third embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment and second embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 6 , before step S10 of the abnormal sound detection method, the method further includes steps S01 to S04:

[0114] Step S01, perform data augmentation on multiple sample audio and the abnormal sound label information of each sample audio to obtain multiple model training samples and the abnormal sound label information of each model training sample.

[0115] It should be noted that the sample audio is the original audio file used for training, including various types of abnormal sounds and normal audio; the abnormal sound label information is the category identification information of each sample audio, which is used to indicate whether there is an abnormal sound in the audio and the specific abnormal sound type when there is an abnormal sound.

[0116] It can be understood that in order to ensure the diversity of samples during model training, it is necessary to process the sample audio through data augmentation techniques (such as speed perturbation techniques, etc.) to generate more training audio, and each generated training audio has its corresponding abnormal sound label information. In this embodiment, the model training samples are composed of a large number of sample audio and the generated training audio.

[0117] In a specific implementation, the sample audio can be obtained by sending a swept-frequency signal containing multiple test frequencies to the target speaker and collecting the audio signal emitted by the target speaker, or can be obtained by other means of collection, and this embodiment does not limit this.

[0118] Step S02, perform signal processing on each model training sample to obtain the filter bank features of each model training sample.

[0119] It should be noted that after obtaining multiple model training samples, perform operations such as frame division - fast Fourier transform and power spectrum calculation - Mel filtering on the model training samples, so as to obtain the FBank features of multiple frames of signals corresponding to each model training sample. The FBank features of multiple frames of signals corresponding to each model training sample together constitute the filter bank features of each model training sample.

[0120] It can be understood that, for a further understanding of the signal processing process, the signal processing process is now illustrated by way of example, and the parameters involved in the illustration process do not limit the processing process: The continuous speech of the module training samples is segmented into short-time segments with a 25 ms window (moving 10 ms each time, generating a 15 ms overlap). After applying a Hamming window to each frame and performing FFT on each frame to obtain the power spectrum, through a Mel filter bank composed of 80 triangular filters, the linear frequency is compressed into 80-dimensional Mel-scale energy values. After taking the logarithmic operation on the output of the filter bank, 80-dimensional Fbank features of each frame are extracted.

[0121] Step S03, set the parameters of the loss function according to the preset interval factor and the preset scaling factor to determine the target loss function.

[0122] It should be noted that the preset interval factor and the preset scaling factor are parameters set according to requirements. In this embodiment, the preset interval factor and the preset scaling factor are set to 0.2 and 32 respectively. At the same time, in this embodiment, the AAM-Softmax loss function with an angular additive margin is selected, and the AAM-Softmax loss function is set using the preset interval factor and the preset scaling factor to obtain the target loss function. The AAM-Softmax (Additive Angular Margin Softmax) loss function is an improved softmax loss function used in deep learning models. It introduces an angular margin on the basis of the traditional softmax loss to enhance the inter-class difference and reduce the intra-class variation, thereby improving the discrimination ability of the model.

[0123] Step S04, train the initial detection model according to the target loss function, the target gradient descent optimizer, the target scheduler, the filter bank features of each model training sample, and the outlier label information of each model training sample to obtain the target detection model. The initial detection model is composed of an initial front-end convolutional module, multiple initial feature extraction modules, at least one initial squeeze-and-excitation module, and an initial detection classification unit.

[0124] It should be noted that in this embodiment, the target gradient descent optimizer is illustrated by taking the Stochastic Gradient Descent (SGD) optimizer as an example; the target scheduler in this embodiment includes a cosine annealing scheduler and a linear warm-up scheduler; the learning rate is dynamically adjusted between 0.1 and 1e-4; the momentum parameter is set to 0.9; and the weight decay is 1e-4. The above-mentioned scheduler, optimizer, learning rate, momentum parameter, and weight decay can all be adjusted according to requirements, and this embodiment does not limit this. The initial front-end convolution module is an untrained FCM; each initial feature extraction module is composed of multiple densely connected untrained D-TDNN layers; the initial squeeze-and-excitation module is an untrained SE module; the initial detection and classification unit is an untrained detection and classification unit, and its core components include a fully connected layer and an activation function.

[0125] It can be understood that by dividing multiple model training samples, a training set and a validation set are obtained. Using the audio in the training set and its corresponding abnormal sound label information, the initial detection model is trained in combination with the target gradient optimizer, the target scheduler, the learning rate, the target loss function, the momentum parameter, and the weight decay, so as to obtain a preliminarily trained detection model. During the training process, the target loss function is used to guide the model to optimize parameters and improve the classification accuracy.

[0126] In a specific implementation, the preliminarily trained detection model is tested with the validation set, and various performance indicators are calculated, including but not limited to accuracy, recall, and F1 score, etc. If all the performance indicators of the detection model meet the expected standards, it can be used as the final target detection model; otherwise, the hyperparameters or structure of the preliminarily trained detection model are adjusted and retrained to obtain the target detection model.

[0127] In a feasible implementation manner, step S01 may further include steps D11 to D13:

[0128] Step D11, randomly select in a preset rate ratio set to determine the target rate ratio.

[0129] It should be noted that the preset rate ratio set is composed of multiple rate ratios. The rate ratio can be preset or randomly generated, and this embodiment does not limit this. Randomly extract a rate ratio from the preset rate ratios and use it as the target rate ratio. For example, the preset rate ratio set is {0.9, 1.0, 1.1}.

[0130] Step D12, perform audio sampling on each sample audio according to the target rate ratio to obtain multiple extended audios and the abnormal sound label information of each extended audio.

[0131] It should be noted that multiple sample audio to be processed are randomly selected from multiple sample audios. Each sample audio to be processed can correspond to the same target speed ratio, or each sample audio to be processed can respectively correspond to different target speed ratios.

[0132] It can be understood that audio sampling is performed on each sample audio to be processed through the target speed ratio corresponding to each sample audio to be processed. The specific process is as follows: speed perturbation is performed on the audio through the target speed ratio, and the playback speed of the audio is changed through a speed change algorithm (such as resampling or waveform stretching) to generate an extended audio. In this embodiment, it is ensured that the abnormal sound label information of the extended audio is consistent with the abnormal sound label information of the original sample audio.

[0133] Step D13: Obtain a plurality of model training samples and the abnormal sound label information of each model training sample according to each extended audio, the abnormal sound label information of each extended audio, each sample audio, and the abnormal sound label information of each sample audio.

[0134] It should be noted that all newly generated extended audios and all original sample audios are aggregated to obtain a plurality of model training samples; and based on the abnormal sound label information of the newly generated extended audios and the abnormal sound label information of each sample audio, the abnormal sound label information of each model training sample is obtained.

[0135] This embodiment provides a method for detecting abnormal sounds. In this embodiment, data augmentation is performed according to multiple sample audios and the abnormal sound label information of each sample audio to obtain a plurality of model training samples and the abnormal sound label information of each model training sample; signal processing is performed on each model training sample to obtain the filter bank features of each model training sample; parameter settings are performed on the loss function according to a preset interval factor and a preset scaling factor to determine the target loss function; the initial detection model is trained according to the target loss function, the target gradient descent optimizer, the target scheduler, the filter bank features of each model training sample, and the abnormal sound label information of each model training sample to obtain the target detection model. The initial detection model is composed of an initial front-end convolutional module, a plurality of initial feature extraction modules, at least one initial squeeze-and-excitation module, and an initial detection classification unit. Through the above method, by adjusting the parameters of the loss function and combining optimization strategies, the trained target detection model significantly improves the accuracy and robustness of abnormal sound detection, realizes efficient feature discrimination ability and strong generalization ability for complex acoustic scenes, and at the same time optimizes the training convergence speed and stability.

[0136] Exemplarily, to help understand the implementation process of the abnormal sound detection method obtained by combining the above Embodiment 1 and Embodiment 2 in this embodiment, please refer to Figure 7 , Figure 7A brief process schematic diagram of a heterophony detection method is provided. Specifically: S1, a swept-frequency signal containing multiple test frequencies is sent to the target speaker, and the audio signal emitted by the target speaker (i.e., the sample audio) is collected. S2, the continuous speech corresponding to the audio signal is segmented into short-time segments with a 25-ms window (moving 10 ms each time, resulting in a 15-ms overlap). A Hamming window is applied to each frame. After performing FFT on each frame to obtain the power spectrum, through a Mel filter bank composed of 80 triangular filters, the linear frequency is compressed into 80-dimensional Mel-scale energy values. A log operation is taken on the output of the filter bank to extract 80-dimensional Fbank features. Speed perturbation augmentation is performed by randomly sampling the speed ratio from {0.9, 1.0, 1.1}. S3, the network first inputs the acoustic features into the FCM network. The generated feature map is then flattened along the channel dimension and the frequency dimension and passed as input to the D-TDNN backbone network. The D-TDNN backbone network contains three modules, and each module consists of a series of D-TDNN layers. In each D-TDNN layer, an SE module is constructed. The number of layers in each module is adjusted to 12 / 24 / 16 layers to form deeper feature abstractions; the growth rate k of each module is reduced from 64 to 32 to construct narrower D-TDNN layers, and the increase in the number of parameters caused by the increase in depth is offset by reducing the number of channels in a single layer; an input TDNN layer with a downsampling rate of 1 / 2 is added before the D-TDNN backbone to accelerate the calculation. S4, set the training parameters of the model: select the AAM-Softmax loss function with angular additive margin, and its margin factor and scaling factor are set to 0.2 and 32 respectively. During the training process, the stochastic gradient descent (SGD) optimizer is used, in conjunction with a cosine annealing scheduler and a linear warm-up scheduler, and the learning rate is dynamically adjusted between 0.1 and 1e-4. The momentum parameter is set to 0.9, and the weight decay is 1e-4. S5, train the model: input the processed data set obtained in S2 into the improved network in S3, and train the network according to the parameters in S4 to obtain the target detection model. S6, model inference: during inference, first perform the processing in S2 on the wav corresponding to the target detection audio to obtain the Fbank features, and input them into the model trained in S5 for inference to obtain the result.

[0137] This embodiment solves the problems of limited local feature capture and insufficient long-term acoustic context modeling in two-dimensional spectrograms and traditional CNN methods for abnormal sound detection in acoustic products. The processing methods mainly include the following aspects: (1) The front-end convolutional module (FCM) consists of multiple two-dimensional convolutional blocks with residual connections, encodes acoustic features in the time-frequency domain, and utilizes high-resolution time-frequency details. (2) D-TDNN adopts a dense connection mechanism, which is parameter-efficient and achieves better performance while requiring fewer parameters than traditional TDNN. (3) The SE module is used to learn feature weights according to the loss function through Squeeze, Excitation, and Scale operations, making the effective feature maps have larger weights. (4) A direct connection is established between the inputs of two consecutive D-TDNN layers, significantly increasing the depth of the D-TDNN network, and controlling the complexity by reducing the number of filter channels in each layer.

[0138] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the abnormal sound detection method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0139] This application also provides an abnormal sound detection device. Please refer to Figure 8 , and the abnormal sound detection device includes:

[0140] The extraction module 10 is used to extract features from the two-dimensional feature representation corresponding to the target detection audio through the feature capture unit in the target detection model, and obtain the audio features to be detected of the target detection audio. The feature capture unit consists of multiple feature extraction modules and at least one squeeze-and-excitation module, and the feature extraction module consists of multiple densely connected time-delay neural network layers.

[0141] The detection module 20 is used to input the audio features to be detected into the detection and classification unit in the target detection model for abnormal sound detection, and obtain the category to which the audio of the target detection audio belongs.

[0142] The processing module 30 is used to obtain the audio detection result of the target detection audio according to the category to which the audio of the target detection audio belongs.

[0143] Optionally, the extraction module 10 is further used for:

[0144] Input the two-dimensional feature representation corresponding to the target detection audio into the feature capture unit in the target detection model. Through the first feature extraction module of the feature capture unit, perform temporal feature extraction on the two-dimensional feature representation to obtain the target temporal feature corresponding to the target detection audio; input the target temporal feature into the squeeze-and-excitation module in the feature capture unit, and through the squeeze-and-excitation module, perform weight adjustment on the target temporal feature to obtain the weighted audio feature of the target detection audio; according to the weighted audio feature and the second feature extraction module in multiple feature extraction modules, perform feature extraction on the target detection audio to obtain the audio feature to be detected corresponding to the target detection audio.

[0145] Optionally, the extraction module 10 is further configured to:

[0146] Input the two-dimensional feature representation into the first feature extraction module of the feature capture unit, and through the first feedforward neural network in the first feature extraction module, perform feature transformation on the two-dimensional feature representation to obtain the first audio feature of the target detection audio. The first feedforward neural network is the basic unit of the first densely connected temporal neural network layer in multiple densely connected temporal neural network layers; input the first audio feature into the temporal neural network in the first densely connected temporal neural network layer respectively, and through the temporal neural network, perform temporal feature extraction on the first audio feature, and obtain the target temporal feature corresponding to the target detection audio according to the feature extraction result of the temporal neural network and the hierarchical structure of the first feature extraction module.

[0147] Optionally, the extraction module 10 is further configured to:

[0148] In response to the abnormal sound detection instruction, perform signal processing on the target detection audio to obtain the filter bank feature of the target detection audio; input the filter bank feature into the front-end convolution module in the target detection model, and through the front-end convolution module, perform local feature extraction on the filter bank feature, and obtain the two-dimensional feature representation corresponding to the target detection audio according to the extraction result.

[0149] Optionally, the extraction module 10 is further configured to:

[0150] In response to the abnormal sound detection instruction, frame the target detection audio according to a preset window to obtain multiple frames of audio signals corresponding to the target detection audio; perform windowing processing on each audio signal respectively to obtain windowed signals corresponding to each audio signal respectively; calculate the power spectrum of each windowed signal respectively to determine the linear frequency power spectrum of each windowed signal; perform Mel filter bank processing on the linear frequency power spectrum of each windowed signal respectively, and obtain the filter bank features of each windowed signal according to the processing results; determine the filter bank features of the target detection audio according to the filter bank features of each windowed signal.

[0151] Optionally, the extraction module 10 is further configured to:

[0152] Perform data augmentation according to multiple sample audios and the abnormal sound label information of each sample audio to obtain multiple model training samples and the abnormal sound label information of each model training sample; perform signal processing on each model training sample to obtain the filter bank features of each model training sample; set the parameters of the loss function according to a preset interval factor and a preset scaling factor to determine the target loss function; train the initial detection model according to the target loss function, the target gradient descent optimizer, the target scheduler, the filter bank features of each model training sample, and the abnormal sound label information of each model training sample to obtain the target detection model, where the initial detection model is composed of an initial front-end convolution module, multiple initial feature extraction modules, at least one initial squeeze-and-excitation module, and an initial detection classification unit.

[0153] Optionally, the extraction module 10 is further configured to:

[0154] Randomly select in a preset rate ratio set to determine the target rate ratio; perform audio sampling on each sample audio according to the target rate ratio respectively to obtain multiple extended audios and the abnormal sound label information of each extended audio; obtain multiple model training samples and the abnormal sound label information of each model training sample according to the extended audios and the abnormal sound label information of each extended audio, and the sample audios and the abnormal sound label information of each sample audio.

[0155] The abnormal sound detection device provided by the present application adopts the abnormal sound detection method in the above-mentioned embodiment, and can solve the technical problem that the existing abnormal sound detection technology cannot achieve high-precision, low-cost, and strong generalization abnormal sound detection in complex acoustic environments and diverse product lines. Compared with the existing technology, the beneficial effects of the abnormal sound detection device provided by the present application are the same as those of the abnormal sound detection method provided by the above-mentioned embodiment, and other technical features in the abnormal sound detection device are the same as the features disclosed in the method of the above-mentioned embodiment, and will not be elaborated here.

[0156] The present application provides a abnormal sound detection device, and the abnormal sound detection device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the abnormal sound detection method in Embodiment 1 above.

[0157] Reference is made below to Figure 9 , which shows a schematic structural diagram of an abnormal sound detection device suitable for implementing the embodiments of the present application. The abnormal sound detection device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description), PMPs (Portable Media Player), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The abnormal sound detection device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0158] As Figure 9 shown, the abnormal sound detection device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can execute various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the abnormal sound detection device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the abnormal sound detection device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an abnormal sound detection device having various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be alternatively implemented or had.

[0159] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by a processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0160] The abnormal sound detection device provided by the present application adopts the abnormal sound detection method in the above embodiments, and can solve the technical problem that the existing abnormal sound detection technology cannot achieve high-precision, low-cost, and strong generalization abnormal sound detection in complex acoustic environments and diverse product lines. Compared with the prior art, the beneficial effects of the abnormal sound detection device provided by the present application are the same as those of the abnormal sound detection method provided by the above embodiments, and other technical features in the abnormal sound detection device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.

[0161] It should be understood that each part disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0162] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0163] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the abnormal sound detection method in the above embodiments.

[0164] The computer-readable storage medium provided by the present application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0165] The above computer-readable storage medium can be included in the abnormal sound detection device; it can also exist separately without being assembled into the abnormal sound detection device.

[0166] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by the abnormal sound detection device, the abnormal sound detection device is caused to: extract features from the two-dimensional feature representation corresponding to the target detection audio through the feature capture unit in the target detection model to obtain the audio features to be detected of the target detection audio, the feature capture unit is composed of multiple feature extraction modules and at least one squeeze-and-excitation module, and the feature extraction module is composed of multiple densely connected delay neural network layers; input the audio features to be detected into the detection and classification unit in the target detection model for abnormal sound detection to obtain the audio category to which the target detection audio belongs; and obtain the audio detection result of the target detection audio according to the audio category to which the target detection audio belongs.

[0167] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN: Local Area Network) or a wide area network (WAN: Wide Area Network), or it can be connected to an external computer (for example, by connecting through an Internet service provider using the Internet).

[0168] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and this module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0169] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0170] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned abnormal sound detection method, and can solve the technical problem that the existing abnormal sound detection technology cannot achieve high-precision, low-cost, and strong generalization abnormal sound detection in complex acoustic environments and diverse product lines. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the abnormal sound detection method provided in the above embodiments, and will not be elaborated here.

[0171] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the abnormal sound detection method as described above.

[0172] The computer program product provided by the present application can solve the technical problem that the existing abnormal sound detection technology cannot achieve high-precision, low-cost and strong generalization abnormal sound detection in complex acoustic environments and diverse product lines. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the abnormal sound detection method provided by the above embodiments, and will not be elaborated herein.

[0173] The foregoing are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A method for detecting abnormal noise, characterized in that, The method includes: Performing feature extraction on the two-dimensional feature representation corresponding to the target detection audio through a feature capture unit in the target detection model to obtain the to-be-detected audio features of the target detection audio. The feature capture unit is composed of multiple feature extraction modules and at least one squeeze-and-excitation module. The feature extraction module is composed of multiple densely connected time-delay neural network layers. The multiple D-DTNN layers in the feature extraction module adopt a dense connection mechanism, and the outputs of all previous D-TDNN layers are used as the input of the current D-TDNN layer to form a densely connected chain-like feature fusion. Inputting the to-be-detected audio features into a detection and classification unit in the target detection model to perform abnormal sound detection, and obtaining the audio category to which the target detection audio belongs. Obtaining the audio detection result of the target detection audio according to the audio category to which the target detection audio belongs.

2. The method according to claim 1, wherein The step of performing feature extraction on the two-dimensional feature representation corresponding to the target detection audio through a feature capture unit in the target detection model to obtain the to-be-detected audio features of the target detection audio includes: Inputting the two-dimensional feature representation corresponding to the target detection audio into the feature capture unit in the target detection model, and performing temporal feature extraction on the two-dimensional feature representation through a first feature extraction module of the feature capture unit to obtain the target temporal features corresponding to the target detection audio. Inputting the target temporal features into the squeeze-and-excitation module in the feature capture unit, and performing weight adjustment on the target temporal features through the squeeze-and-excitation module to obtain the weighted audio features of the target detection audio. Performing feature extraction on the target detection audio according to the weighted audio features and a second feature extraction module in multiple feature extraction modules to obtain the to-be-detected audio features corresponding to the target detection audio.

3. The method according to claim 2, characterized in that The densely connected time-delay neural network layer includes a feedforward neural network and at least one time-delay neural network. The step of performing temporal feature extraction on the two-dimensional feature representation through a first feature extraction module of the feature capture unit to obtain the target temporal features corresponding to the target detection audio includes: Inputting the two-dimensional feature representation into the first feature extraction module of the feature capture unit, and performing feature transformation on the two-dimensional feature representation through a first feedforward neural network in the first feature extraction module to obtain the first audio features of the target detection audio. The first feedforward neural network is the basic unit of the first densely connected time-delay neural network layer in multiple densely connected time-delay neural network layers. Inputting the first audio features into the time-delay neural network in the first densely connected time-delay neural network layer respectively, performing temporal feature extraction on the first audio features through the time-delay neural network, and obtaining the target temporal features corresponding to the target detection audio according to the feature extraction result of the time-delay neural network and the hierarchical structure of the first feature extraction module.

4. The method according to any one of claims 1 to 3, characterized in that, Before the step of performing feature extraction on the two-dimensional feature representation corresponding to the target detection audio through a feature capture unit in the target detection model to obtain the to-be-detected audio features of the target detection audio, it further includes: In response to the abnormal sound detection instruction, signal processing is performed on the target detection audio to obtain the filter bank features of the target detection audio; The filter bank features are input into the front-end convolution module in the target detection model, and the front-end convolution module performs local feature extraction on the filter bank features, and a two-dimensional feature representation corresponding to the target detection audio is obtained according to the extraction result.

5. The method according to claim 4, wherein The step of performing signal processing on the target detection audio in response to the abnormal sound detection instruction to obtain the filter bank features of the target detection audio includes: In response to the abnormal sound detection instruction, frame division processing is performed on the target detection audio according to a preset window to obtain multiple frames of audio signals corresponding to the target detection audio; Windowing processing is respectively performed on each audio signal to obtain windowed signals corresponding to each audio signal; Power spectrum calculation is respectively performed on each windowed signal to determine the linear frequency power spectrum of each windowed signal; Mel filter bank processing is respectively performed on the linear frequency power spectra of each windowed signal, and filter bank features of each windowed signal are obtained according to the processing results; The filter bank features of the target detection audio are determined according to the filter bank features of each windowed signal.

6. The method according to any one of claims 1 to 3, characterized in that, Before the step of performing feature extraction on the two-dimensional feature representation corresponding to the target detection audio through the feature capture unit in the target detection model to obtain the audio features to be detected of the target detection audio, it further includes: Data augmentation is performed according to multiple sample audios and the abnormal sound label information of each sample audio to obtain multiple model training samples and the abnormal sound label information of each model training sample; Signal processing is performed on each model training sample to obtain the filter bank features of each model training sample; Parameter settings are performed on the loss function according to a preset interval factor and a preset scaling factor to determine the target loss function; The initial detection model is trained according to the target loss function, the target gradient descent optimizer, the target scheduler, the filter bank features of each model training sample, and the abnormal sound label information of each model training sample to obtain the target detection model. The initial detection model is composed of an initial front-end convolution module, multiple initial feature extraction modules, at least one initial squeeze-and-excitation module, and an initial detection classification unit.

7. The method according to claim 6, wherein The step of performing data augmentation according to multiple sample audios and the abnormal sound label information of each sample audio to obtain multiple model training samples and the abnormal sound label information of each model training sample includes: Random selection is performed in the preset rate ratio set to determine the target rate ratio; Audio sampling is respectively performed on each sample audio according to the target rate ratio to obtain multiple extended audios and the abnormal sound label information of each extended audio; According to the extended audios and the abnormal sound label information of each extended audio, and the sample audios and the abnormal sound label information of each sample audio, multiple model training samples and the abnormal sound label information of each model training sample are obtained.

8. A abnormal sound detection device, characterized in that, The abnormal sound detection device includes: An extraction module, configured to extract features from the two-dimensional feature representation corresponding to the target detection audio through a feature capture unit in the target detection model, to obtain the audio features to be detected of the target detection audio. The feature capture unit is composed of multiple feature extraction modules and at least one squeeze-and-excitation module. The feature extraction module is composed of multiple densely connected time-delay neural network layers. The multiple D-DTNN layers in the feature extraction module adopt a dense connection mechanism, and the outputs of all previous D-TDNN layers are used as the input of the current D-TDNN layer to form a dense-connected chain-like feature fusion; A detection module, configured to input the audio features to be detected into a detection and classification unit in the target detection model for abnormal sound detection, to obtain the audio category to which the target detection audio belongs; A processing module, configured to obtain the audio detection result of the target detection audio according to the audio category to which the target detection audio belongs.

9. An abnormal sound detection device, characterized in that, The device includes: a memory, a processor, and an abnormal sound detection program stored on the memory and executable on the processor. The abnormal sound detection program is configured to implement the steps of the abnormal sound detection method according to any one of claims 1 to 7.

10. A storage medium, characterized in that, An abnormal sound detection program is stored on the storage medium. When the abnormal sound detection program is executed by a processor, it implements the steps of the abnormal sound detection method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Abnormal sound detection method and system of neural network based on multiple receptive fields

    CN119091909A

  • Voiceprint coding method, voiceprint coding network training method and electronic equipment

    CN119400187A