Abnormal sound detection method and device, equipment and storage medium

By using the feature capture unit of the object detection model and the densely connected delay neural network layer in the antonym detection technology, the problem of difficult to achieve high-precision antonym detection in complex acoustic environments in the prior art is solved, and the detection effects of high precision, low cost, strong generalization and interpretability are achieved.

CN120089161AActive Publication Date: 2025-06-03GOERTEK INC

Patent Information

Application Number
CN202510560155.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

Existing antonal detection technology is difficult to achieve high-precision, low-cost, strong generalization and interpretability antonal detection in complex acoustic environments and diversified product lines.

Method used

The feature capture unit in the object detection model is used to extract the audio features, and combine the densely connected delay neural network layer and the compression excitation module to perform different sound detection.

Benefits of technology

It significantly improves the detection accuracy and robustness of the model, enhances generalization, realizes high-precision antonym detection with low computing cost, and provides interpretable detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089161A_ABST
    Figure CN120089161A_ABST
Patent Text Reader

Abstract

The invention discloses an abnormal sound detection method and device, equipment and a storage medium, and the method comprises the steps: carrying out the feature extraction of a two-dimensional feature representation of a target detection audio through a feature capture unit in a target detection model, and obtaining a to-be-detected audio feature of the target detection audio, the feature capture unit is composed of a plurality of feature extraction modules and at least one compression excitation module, and each feature extraction module is composed of a plurality of densely connected time delay neural network layers; inputting the to-be-detected audio features into a detection classification unit in a target detection model for abnormal sound detection to obtain an audio category of the target detection audio; and obtaining an audio detection result of the target detection audio according to the audio category of the target detection audio. By means of the mode, modeling optimization is conducted on the time sequence features, a multi-granularity context sensing mechanism is established, time-frequency domain feature enhancement is enhanced, the detection precision and robustness of the model are remarkably improved, and meanwhile generalization of the model is greatly enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of acoustic technologies, and particularly to a method, device, equipment, and storage medium for detecting abnormal sounds. Background Art

[0002] In the field of acoustic product manufacturing, abnormal sound detection is a core link in ensuring the acoustic quality of products, directly determining the user experience and market competitiveness. Traditional detection methods rely on the subjective judgment of human listeners, with inherent defects such as low efficiency, high cost, and poor consistency. Although current automated detection technologies attempt to improve efficiency through two-dimensional spectrogram analysis and Convolutional Neural Network (CNN), they still face significant bottlenecks: the two-dimensional spectrogram analysis method relies on manual feature engineering, easily ignores abnormal sound-sensitive information, has limited resolution ability in complex noise scenarios, and high-resolution time-frequency analysis leads to a sharp increase in computational burden; traditional CNN methods, although having the ability to extract local features, are insufficient in modeling the long-term dependence of sound signals, and at the same time, their translational invariance may weaken the detection accuracy for abnormal sound time-series sensitive scenarios. Existing technologies are difficult to achieve high-precision, low-computation-cost, strong generalization, and interpretable abnormal sound detection in complex acoustic environments and diverse product lines, restricting large-scale deployment and real-time diagnosis capabilities in industrial scenarios.

[0003] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of this application is to provide a method, device, equipment, and storage medium for detecting abnormal sounds, aiming to solve the technical problem that existing abnormal sound detection technologies cannot achieve high-precision, low-cost, and strong generalization abnormal sound detection in complex acoustic environments and diverse product lines.

[0005] To achieve the above purpose, this application proposes an abnormal sound detection method, and the method includes: Performing feature extraction on the two-dimensional feature representation corresponding to the target detection audio through a feature capture unit in the target detection model to obtain the audio feature to be detected of the target detection audio, where the feature capture unit is composed of multiple feature extraction modules and at least one squeeze-and-excitation module, and the feature extraction module is composed of multiple densely connected delay neural network layers; Inputting the audio feature to be detected into a detection and classification unit in the target detection model for abnormal sound detection to obtain the audio category to which the target detection audio belongs; Obtaining the audio detection result of the target detection audio according to the audio category to which the target detection audio belongs.

[0006] In one embodiment, the step of extracting the audio feature to be detected of the target detection audio by the feature capture unit in the target detection model includes: Input the two-dimensional feature representation corresponding to the target detection audio into the feature capture unit in the target detection model, and perform temporal feature extraction on the two-dimensional feature representation through the first feature extraction module of the feature capture unit to obtain the target temporal feature corresponding to the target detection audio; Input the target temporal feature into the squeeze-and-excitation module in the feature capture unit, and perform weight adjustment on the target temporal feature through the squeeze-and-excitation module to obtain the weighted audio feature of the target detection audio; Extract the feature of the target detection audio according to the weighted audio feature and the second feature extraction module in the multiple feature extraction modules to obtain the audio feature to be detected corresponding to the target detection audio.

[0007] In one embodiment, the densely connected delay neural network layer includes a feedforward neural network and at least one delay neural network; The step of performing temporal feature extraction on the two-dimensional feature representation through the first feature extraction module of the feature capture unit to obtain the target temporal feature corresponding to the target detection audio includes: Input the two-dimensional feature representation into the first feature extraction module of the feature capture unit, and perform feature transformation on the two-dimensional feature representation through the first feedforward neural network in the first feature extraction module to obtain the first audio feature of the target detection audio, where the first feedforward neural network is the basic unit of the first densely connected delay neural network layer in the multiple densely connected delay neural network layers; Input the first audio feature into the delay neural network in the first densely connected delay neural network layer respectively, perform temporal feature extraction on the first audio feature through the delay neural network, and obtain the target temporal feature corresponding to the target detection audio according to the feature extraction result of the delay neural network and the hierarchical structure of the first feature extraction module.

[0008] In one embodiment, before the step of extracting the audio feature to be detected of the target detection audio by the feature capture unit in the target detection model, it further includes: In response to the abnormal sound detection instruction, perform signal processing on the target detection audio to obtain the filter bank feature of the target detection audio; Input the filter bank features into the front-end convolutional module in the target detection model, extract local features from the filter bank features through the front-end convolutional module, and obtain a two-dimensional feature representation corresponding to the target detection audio according to the extraction results.

[0009] In one embodiment, the step of performing signal processing on the target detection audio in response to an abnormal sound detection instruction to obtain the filter bank features of the target detection audio includes: In response to an abnormal sound detection instruction, perform frame splitting on the target detection audio according to a preset window to obtain multiple frames of audio signals corresponding to the target detection audio; Perform windowing processing on each audio signal respectively to obtain windowed signals corresponding to each audio signal respectively; Calculate the power spectrum of each windowed signal respectively to determine the linear frequency power spectrum of each windowed signal; Perform Mel filter bank processing on the linear frequency power spectrum of each windowed signal respectively, and obtain the filter bank features of each windowed signal according to the processing results; Determine the filter bank features of the target detection audio according to the filter bank features of each windowed signal.

[0010] In one embodiment, before the step of extracting the to-be-detected audio features of the target detection audio by using the feature capture unit in the target detection model, it further includes: Perform data augmentation according to multiple sample audios and the abnormal sound label information of each sample audio to obtain multiple model training samples and the abnormal sound label information of each model training sample; Perform signal processing on each model training sample to obtain the filter bank features of each model training sample; Set the parameters of the loss function according to a preset interval factor and a preset scaling factor to determine the target loss function; Train the initial detection model according to the target loss function, the target gradient descent optimizer, the target scheduler, the filter bank features of each model training sample, and the abnormal sound label information of each model training sample to obtain the target detection model, where the initial detection model is composed of an initial front-end convolutional module, multiple initial feature extraction modules, at least one initial squeeze-and-excitation module, and an initial detection classification unit.

[0011] In one embodiment, the step of performing data augmentation according to multiple sample audios and the abnormal sound label information of each sample audio to obtain multiple model training samples and the abnormal sound label information of each model training sample includes: Randomly select in a preset rate ratio set to determine the target rate ratio; Audio sampling is performed on each sample audio according to the target rate ratio to obtain a plurality of extended audios and the abnormal sound label information of each extended audio; According to each extended audio and the abnormal sound label information of each extended audio, each sample audio and the abnormal sound label information of each sample audio, a plurality of model training samples and the abnormal sound label information of each model training sample are obtained.

[0012] In addition, to achieve the above object, the present application also proposes an abnormal sound detection device, where the abnormal sound detection device includes: an extraction module, configured to extract features from the two-dimensional feature representation corresponding to the target detection audio through the feature capture unit in the target detection model to obtain the to-be-detected audio features of the target detection audio, and the feature capture unit is composed of a plurality of feature extraction modules and at least one squeeze-and-excitation module, and the feature extraction module is composed of a plurality of densely connected delay neural network layers; A detection module, configured to input the to-be-detected audio features into the detection and classification unit in the target detection model for abnormal sound detection to obtain the audio category to which the target detection audio belongs; A processing module, configured to obtain the audio detection result of the target detection audio according to the audio category to which the target detection audio belongs.

[0013] In addition, to achieve the above object, the present application also proposes an abnormal sound detection device, where the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the abnormal sound detection method as described above.

[0014] In addition, to achieve the above object, the present application also proposes a storage medium, where the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the abnormal sound detection method as described above are implemented.

[0015] In addition, to achieve the above object, the present application also provides a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the abnormal sound detection method as described above are implemented.

[0016] The present application provides a heterophony detection method. In the present application, a feature capture unit in a target detection model extracts features from a two-dimensional feature representation corresponding to a target detection audio to obtain a to-be-detected audio feature of the target detection audio. The feature capture unit is composed of a plurality of feature extraction modules and at least one squeeze-and-excitation module, and the feature extraction module is composed of a plurality of densely connected delay neural network layers; the to-be-detected audio feature is input into a detection and classification unit in the target detection model for heterophony detection to obtain the audio category to which the target detection audio belongs; an audio detection result of the target detection audio is obtained according to the audio category to which the target detection audio belongs. By the above method, the modeling of time series features is optimized, a multi-granularity context awareness mechanism is established, the enhancement of time-frequency domain features is strengthened, the detection accuracy and robustness of the model are significantly improved, and at the same time, its generalization ability is greatly enhanced. It can achieve high-precision heterophony detection with low computational cost in complex acoustic environments and diverse product lines, and provide interpretability support for the detection results. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the heterophony detection method of the present application; Figure 2 It is a schematic diagram of the FCM network structure provided for Embodiment 1 of the present application; Figure 3 It is a schematic flowchart provided for Embodiment 2 of the heterophony detection method of the present application; Figure 4 It is a schematic diagram of the network structure of the feature extraction module provided for Embodiment 2 of the present application; Figure 5 It is a schematic diagram of the network structure of the SE module provided for Embodiment 2 of the present application; Figure 6 It is a schematic flowchart provided for Embodiment 3 of the heterophony detection method of the present application; Figure 7 It is a schematic flowchart of the brief process of the heterophony detection method provided for Embodiment 3 of the present application; Figure 8 It is a schematic diagram of the module structure of the heterophony detection device for the embodiments of the present application; Figure 9 This is a schematic diagram of the device structure of the hardware operating environment involved in the abnormal sound detection method in the embodiments of the present application.

[0020] The implementation, functional characteristics, and advantages of the present application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. Specific embodiments

[0021] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0022] To better understand the technical solutions of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific embodiments.

[0023] The main solution of the embodiments of the present application is as follows: Feature extraction is performed on the two-dimensional feature representation corresponding to the target detection audio through the feature capture unit in the target detection model to obtain the audio feature to be detected of the target detection audio. The feature capture unit is composed of multiple feature extraction modules and at least one squeeze-and-excitation module, and the feature extraction module is composed of multiple densely connected delay neural network layers; the audio feature to be detected is input into the detection and classification unit in the target detection model for abnormal sound detection to obtain the audio category to which the target detection audio belongs; the audio detection result of the target detection audio is obtained according to the audio category to which the target detection audio belongs.

[0024] In the field of acoustic product manufacturing, the acoustic quality of products is one of the core elements determining user experience and market competitiveness. Acoustic products such as headphones, speakers, and microphones need to undergo a strict abnormal sound detection process before leaving the factory to ensure that there are no defects such as noise, distortion, abnormal background noise, or structural resonance during their operation. Traditional detection methods highly rely on the subjective judgment of human listeners. By repeatedly playing standardized audio samples and relying on the human ear to identify abnormalities. However, with the exponential growth of the shipment volume of acoustic products and the accelerating product iteration speed, the disadvantages of manual detection have become increasingly prominent: low detection efficiency, rising labor costs, and the human ear is vulnerable to fatigue, individual hearing differences, and environmental noise interference, resulting in difficulty in ensuring detection consistency. Although some enterprises have introduced automated detection equipment, the existing technologies still have difficulty in achieving high-precision and strong generalization abnormal sound determination in complex acoustic scenarios and diverse product lines.

[0025] The existing abnormal sound detection technologies for acoustic products mainly focus on two-dimensional spectrogram analysis and traditional CNN. The two-dimensional spectrogram analysis has the following disadvantages: strong dependence on features, relying on manually designed feature extraction (such as Mel spectrogram, wavelet transform), which may ignore key information sensitive to abnormal sounds, and feature selection requires domain experience; loss of dynamic information, static spectrograms are difficult to capture the temporal dynamic characteristics of sound signals; sensitive to noise, in complex background noise, the resolution ability of artificial features may decrease significantly; high computational cost: high-resolution time-frequency analysis will greatly increase the computational burden. The CNN method has the following disadvantages: insufficient temporal modeling, traditional CNN is good at local spatial feature extraction, but has limited ability to model long-term dependence relationships of sound signals; large data requirements, a large number of labeled abnormal samples are needed to train a robust model, while abnormal samples are usually scarce in actual industrial scenarios; poor interpretability, the black-box characteristic makes it difficult to locate the specific frequency band / time-domain features of abnormal sounds, which is not conducive to fault diagnosis and analysis; interference of translational invariance, the inherent translational invariance of CNN may weaken the detection ability for scenarios sensitive to the occurrence time sequence of abnormal sounds.

[0026] In this application, by optimizing the temporal feature modeling and establishing a multi-granularity context awareness mechanism, the enhancement of time-frequency domain features is strengthened, significantly improving the detection accuracy and robustness of the model. At the same time, its generalization ability is greatly enhanced, enabling high-precision abnormal sound detection at low computational cost in complex acoustic environments and diverse product lines, and providing interpretability support for the detection results.

[0027] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, an abnormal sound detection device, etc. that can implement the above functions. Hereinafter, taking the abnormal sound detection device as an example, this embodiment and the following embodiments will be described.

[0028] Based on this, the embodiment of this application provides an abnormal sound detection method, referring to Figure 1 , Figure 1 is a schematic flowchart of the first embodiment of the abnormal sound detection method of this application.

[0029] In this embodiment, the abnormal sound detection method includes steps S10 to S30: Step S10, extracting features from the two-dimensional feature representation corresponding to the target detection audio through the feature capture unit in the target detection model to obtain the audio features to be detected of the target detection audio. The feature capture unit is composed of multiple feature extraction modules and at least one squeeze-and-excitation module, and the feature extraction module is composed of multiple densely connected delay neural network layers.

[0030] It should be noted that the target detection audio refers to the audio to be detected generated by acoustic products that require abnormal sound detection. Abnormal sounds are unexpected and non-design-permitted abnormal sound signals generated under the normal working state of acoustic products. Their manifestation forms are significantly different from normal acoustic outputs. Abnormal sounds include, but are not limited to, types such as noise, distortion, abnormal background noise, and structural resonance.

[0031] It can be understood that the target detection model is composed of an FCM (Front-end Convolutional Module), a feature capture unit, and a detection and classification unit. The feature capture unit uses D-TDNN as the backbone network, which is composed of multiple feature extraction modules and at least one SE module (Squeeze-and-Excitation Module). Each feature extraction module consists of a series of D-TDNN (Densely Connected Time-Delay Neural Network) layers. The basic unit of each D-DTNN layer is composed of an FNN (Feedforward Neural Network) and a TDNN (Time-Delay Neural Network) to form a dense connection mechanism. The core components of the detection and classification unit include, but are not limited to, fully connected layers and activation functions, which are used to output the abnormal sound detection and classification results.

[0032] In specific implementation, there may be an SE module in each D-TDNN layer of each feature extraction module. This SE module can be located before the FNN or after the TDNN; or, there is an SE module between each feature extraction module; or, an SE module is added after the last feature extraction module; or, the SE can be located at other positions in the feature capture unit, that is, the SE module can be embedded in any intermediate layer such as CNN and TDNN without modifying the backbone network structure. By reducing the channel weights in the low signal-to-noise ratio frequency band, the stability of the model in complex environments is improved. Therefore, the specific location and specific number of the SE module in this embodiment are not limited.

[0033] It should be noted that this embodiment is dedicated to solving the problems of local feature capture limitations and insufficient long-time acoustic context modeling existing in two-dimensional spectra and traditional CNN methods in abnormal sound detection of acoustic products. Due to the fixed frequency resolution of traditional two-dimensional spectrograms, details of high-frequency weak abnormal sounds are lost, and the local convolution kernels of CNNs are difficult to model the dynamic propagation characteristics of abnormal sounds on the time axis. In addition, the differences in acoustic fundamental frequencies across product lines easily cause traditional CNN models to have spectrum pattern mismatches due to insufficient generalization of artificial features, especially the sensitivity to abnormal harmonic components significantly decreases in low signal-to-noise ratio scenarios.

[0034] It can be understood that after obtaining the target detection audio of the acoustic product, the target detection audio is processed to extract the FBank (Mel Filter Bank) features corresponding to the target detection audio, and the FBank features of the target detection audio are input into the target detection model. After the FCM in the target detection model extracts the features, the feature map output by the FCM is flattened in the channel dimension and the frequency dimension to obtain the two-dimensional feature representation of the target detection audio.

[0035] In a specific implementation, the two-dimensional feature representation is sent to the feature capture unit, and the two-dimensional feature representation is further processed by using the feature extraction module with a series of D-TDNN layers and the SE module in the feature capture unit, so as to obtain the audio features to be detected of the target detection audio.

[0036] It should be noted that during feature extraction, D-TDNN with a dense connection mechanism is used. Dense connection cascades the outputs of all previous layers as the input of the current layer. The chain feature fusion of dense connection enables the subsequent layers to dynamically select features at different abstraction levels, and the reduction of the number of parameters directly reduces the memory occupation during inference. The SE module extracts channel-level statistics through global average pooling to capture the global distribution characteristics of the feature map. In the Excitation stage, a fully connected layer is used to model the non-linear relationship between channels.

[0037] In a feasible implementation manner, before step S10, steps A11 to A13 may further be included: Step A11: In response to the abnormal sound detection instruction, perform signal processing on the target detection audio to obtain the filter bank features of the target detection audio.

[0038] It should be noted that the abnormal sound detection instruction is used to trigger the abnormal sound detection process of the acoustic product to determine whether there is an abnormal sound in the audio emitted by the acoustic product. The abnormal sound detection instruction may be actively initiated by the abnormal sound detection device when detecting the acoustic product, or may be initiated by the user to the abnormal sound detection device. This embodiment does not limit this.

[0039] It can be understood that when receiving the abnormal sound detection instruction, the target detection audio of the acoustic product is obtained, and the target detection audio is subjected to processes such as frame division - fast Fourier transform - Mel filtering, etc., so as to obtain the filter bank features of the target detection audio. In this embodiment, the filter bank features of the target detection audio include the FBank features corresponding to each frame.

[0040] Step A12: Input the filter bank features into the front-end convolutional module in the target detection model. The front-end convolutional module extracts local features from the filter bank features, and based on the extraction results, obtain the two-dimensional feature representation corresponding to the target detection audio.

[0041] It should be noted that in this embodiment, the filter bank features are input into the FCM. The FCM integrates residual connections and consists of multiple two-dimensional convolutional blocks with residual connections, which encodes the acoustic features in the time-frequency domain and utilizes the high-resolution time-frequency details. Its specific architecture is as Figure 2 shown. The front-end convolutional module FCM integrates 4 residual modules (residual modules 1-4). Each residual module passes the original input through a skip connection to alleviate the vanishing gradient and support the module to expand to deeper layers. Using FCM to achieve high-frequency time-frequency detail capture, the local receptive field characteristic of two-dimensional convolution enables it to slide simultaneously in both the time domain and the frequency domain, directly modeling the local patterns of acoustic features; the residual connection optimizes the training. In the deep convolutional structure, the residual skip connection alleviates the vanishing gradient problem through the identity mapping, making the network depth scalable. Figure 2 In [1, F, T] is the filter bank feature of the target detection audio, and [D, 1, T] is the feature map output by the FCM.

[0042] It can be understood that there is a feature map output by the FCM in the extraction results. Flatten the feature map along the channel dimension and the frequency dimension to obtain the two-dimensional feature representation of the target detection audio.

[0043] In a feasible implementation manner, step A11 may further include steps B11-B15: Step B11: In response to the abnormal sound detection instruction, perform frame segmentation on the target detection audio according to a preset window to obtain multiple frame audio signals corresponding to the target detection audio.

[0044] It should be noted that since the target detection audio is a continuous speech signal, to improve the processing efficiency, when receiving the abnormal sound detection instruction, the target detection audio is segmented into short-time segments according to a preset window, thereby obtaining multiple frame audio signals corresponding to the target detection audio. In this embodiment, the length of the preset window is set to 25 ms to ensure capturing the short-time features of the audio; the frame shift is 10 ms, and the adjacent frames overlap by 15 ms.

[0045] Step B12: Perform windowing processing on each audio signal to obtain the windowed signals corresponding to each audio signal respectively.

[0046] It should be noted that apply a Hamming window to each frame of audio signal. The windowed audio signal is the windowed signal, thereby reducing the spectral leakage of the audio signal and making the two ends of the frame smoothly decay to zero.

[0047] Step B13: Calculate the power spectrum of each windowed signal respectively to determine the linear frequency power spectrum of each windowed signal.

[0048] It should be noted that after windowing, perform a fast Fourier transform (FFT) on each frame of the windowed signal to obtain a complex frequency spectrum, and calculate the square of the amplitude spectrum, so as to obtain the linear frequency power spectrum of each frame of the windowed signal.

[0049] Step B14: Perform Mel filter bank processing on the linear frequency power spectrum of each windowed signal respectively, and obtain the filter bank features of each windowed signal according to the processing results.

[0050] It should be noted that perform Mel filter bank processing on the linear frequency power spectrum of each frame of the windowed signal respectively. Pass the linear frequency power spectrum through a Mel filter bank composed of a preset number (for example, 80) of triangular filters, so as to obtain the Mel energy value of each frame of the windowed signal. Take the natural logarithm of the Mel energy value of each frame of the windowed signal, so as to output the filter bank features (i.e., FBank features) of each frame of the windowed signal. When the number of triangular filters is 80, 80-dimensional FBank features will be obtained.

[0051] Step B15: Determine the filter bank features of the target detection audio according to the filter bank features of each windowed signal.

[0052] It should be noted that summarize the FBank features of each frame of the windowed signal, and the summarized result is the filter bank features of the target detection audio.

[0053] Step S20: Input the audio features to be detected into the detection and classification unit in the target detection model for abnormal sound detection, and obtain the audio category to which the target detection audio belongs.

[0054] It should be noted that the audio category to which the audio belongs refers to the category label to which the target detection audio is classified, including but not limited to normal audio, noise, broken sound, etc. The detection and classification unit is the decision-making layer in the abnormal sound detection model, and is responsible for mapping the audio features to be detected into the category mapping space to complete the abnormal sound detection task. In this embodiment, the fully connected layer in the detection and classification unit is responsible for mapping the features to the category space to generate the scores of each category, and the activation function in the detection and classification unit is used to convert the scores into a probability distribution to support classification decisions.

[0055] Step S30: Obtain the audio detection result of the target detection audio according to the audio category to which the target detection audio belongs.

[0056] It should be noted that after the audio category of the target detection audio is identified by the target detection model, it can be determined whether there is abnormal sound in the audio to be detected, and when there is abnormal sound, the specific category of the abnormal sound, so as to obtain the final audio detection result. When there is abnormal sound in the target detection audio, the acoustic product can be specifically improved and optimized according to the specific audio category to ensure the product quality.

[0057] This embodiment provides a method for detecting abnormal sound. In this embodiment, the feature capture unit in the target detection model extracts features from the two-dimensional feature representation corresponding to the target detection audio to obtain the audio feature to be detected of the target detection audio. The feature capture unit is composed of multiple feature extraction modules and at least one squeeze-and-excitation module, and the feature extraction module is composed of multiple densely connected time-delay neural network layers; the audio feature to be detected is input into the detection and classification unit in the target detection model for abnormal sound detection to obtain the audio category of the target detection audio; the audio detection result of the target detection audio is obtained according to the audio category of the target detection audio. By the above method, the time-series feature modeling is optimized, a multi-granularity context awareness mechanism is established, the enhancement of time-frequency domain features is strengthened, the detection accuracy and robustness of the model are significantly improved, and at the same time its generalization ability is greatly enhanced. It can achieve high-precision abnormal sound detection at low computational cost in complex acoustic environments and diverse product lines, and provide interpretability support for the detection results.

[0058] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as the above-mentioned embodiment 1 can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3 , step S10, the method for detecting abnormal sound further includes steps S11 to S13: Step S11, input the two-dimensional feature representation corresponding to the target detection audio into the feature capture unit in the target detection model, and perform time-series feature extraction on the two-dimensional feature representation through the first feature extraction module of the feature capture unit to obtain the target time-series feature corresponding to the target detection audio.

[0059] It should be noted that the first feature extraction module refers to the first feature extraction module used to receive the two-dimensional feature representation of the target detection audio. Through a series of D-TDNN layers in the first feature extraction module, time-series feature extraction is performed on the two-dimensional feature representation. The feature output by the TDNN layer in the last D-TDNN layer of the first feature extraction module is the target time-series feature corresponding to the target detection audio. When there are multiple TDNN layers in each D-TDNN layer, the target time-series feature is the result obtained by multiplying the outputs of multiple TDNN layers in the last D-TDNN layer.

[0060] It can be understood that in this embodiment, there are three feature extraction modules in the feature capture unit. The number of layers of each module is 12 / 24 / 16 layers. The growth rate k of each module is reduced from 64 to 32 to construct a narrower D-TDNN layer, and the increase in the number of parameters caused by the increase in depth is offset by reducing the number of channels in a single layer; an input TDNN layer with a downsampling rate of 1 / 2 is added before the D-TDNN backbone to accelerate the calculation. The network structure of the feature capture unit with D-TDNN as the backbone network is as Figure 4 shown, Figure 4 The dense connection blocks _1 to 3 in Figure 4 represent three feature extraction modules. Each D-TDNN layer is composed of a feedforward neural network FNN and two time-delay neural networks TDNN. At the same time, the dense connection method is adopted. After the results of the two TDNN outputs in each layer are multiplied, they are sent to the next layer as the input of the next layer. The chain feature fusion of dense connection enables the backend layer to dynamically select features at different abstraction levels, and the reduction in the number of parameters directly reduces the memory occupancy during inference. Compared with using two feature extraction modules, with the number of layers of each feature extraction module being 6 and 12 layers respectively, the network structure of this embodiment can form deeper feature abstractions.

[0061] It can be understood that since the D-TDNN backbone network in this embodiment adopts a dense connection mechanism, it has parameter efficiency, and while requiring fewer parameters than traditional TDNN, it can achieve better performance.

[0062] Step S12: Input the target temporal feature into the compression excitation module in the feature capture unit, and adjust the weight of the target temporal feature through the compression excitation module to obtain the weighted audio feature of the target detection audio.

[0063] It should be noted that in this embodiment, taking the example that there is an SE module in each D-TDNN layer, and the SE module is located after each TDNN layer. The target temporal feature output by the TDNN layer of the last D-TDNN layer of the first feature extraction module is input into the SE module. Using the SE module, through Squeeze, Excitation, and Scale operations, the feature weights are learned according to the loss function, so that the effective feature maps have larger weights, the channel features of the input feature maps are enhanced, and the learned channel attention information is combined with the input feature maps. At this time, the result output by the SE module is the weighted audio feature of the target detection audio.

[0064] Step S13: Extract features from the target detection audio according to the weighted audio feature and the second feature extraction module in multiple feature extraction modules to obtain the to-be-detected audio feature corresponding to the target detection audio.

[0065] It should be noted that the second feature extraction module refers to all the other feature extraction modules among the multiple feature extraction modules except the first feature extraction module. In this embodiment, there are two feature extraction modules in the second feature extraction module.

[0066] It can be understood that when the weighted audio feature is input into the second feature extraction module, and for the feature extraction module at the front, when there is an SE module in each D-TDNN layer and the SE module is located after the TDNN layer, the result output by the TDNN layer of the last D-TDNN layer of this feature extraction module will be sent into the SE module. At this time, the result output by the SE module will enter the second feature extraction module in the second feature extraction module, and the above steps are repeated until all the modules in the feature capture unit have completed their operations. The result output by the last module in the feature capture unit is the audio feature to be detected. In this embodiment, if there is an SE module in each D-TDNN layer and the SE module is located after the TDNN layer, then the result generated by the last SE module is the audio feature to be detected corresponding to the target detection audio.

[0067] In a feasible implementation manner, the dense connection time-delay neural network layer includes a feedforward neural network and at least one time-delay neural network. In step S11: The first feature extraction module of the feature capture unit performs temporal feature extraction on the two-dimensional feature representation to obtain the target temporal feature corresponding to the target detection audio, which may include steps C11~C12: Step C11, input the two-dimensional feature representation into the first feature extraction module of the feature capture unit, and perform feature conversion on the two-dimensional feature representation through the first feedforward neural network in the first feature extraction module to obtain the first audio feature of the target detection audio. The first feedforward neural network is the basic unit of the first dense connection time-delay neural network layer among the multiple dense connection time-delay neural network layers.

[0068] It should be noted that the first dense connection time-delay neural network layer refers to the D-TDNN layer located at the first position in the first feature extraction module and used to receive the two-dimensional feature representation. The first feedforward neural network is the FNN in the first D-TDNN layer. Input the two-dimensional feature representation into the first FNN, and the first FNN extracts global features and enhances the non-linear expression ability, thereby obtaining the first audio feature of the target detection audio and providing a high-dimensional representation for the subsequent TDNN layer.

[0069] Step C12: Input the first audio feature into the time-delay neural network in the first densely connected time-delay neural network layer respectively. Extract the temporal features of the first audio feature through the time-delay neural network, and obtain the target temporal feature corresponding to the target detection audio according to the feature extraction result of the time-delay neural network and the hierarchical structure of the first feature extraction module.

[0070] It should be noted that the first audio feature output by the first FNN is input into the TDNN in the first D-TDNN layer. The TDNN slides the convolutional kernel in the time dimension to capture local context (such as the dependency relationship between adjacent frames), and expands the receptive field through dilated convolution. When there is only one TDNN, the result output by the TDNN is sent to the SE module (in the case where there is an SE module in each D-TDNN layer and the SE module is located after the TDNN layer), or directly input into the next D-TDNN layer in the first feature extraction module; repeat the above steps until all D-TDNN layers in the first feature extraction module have completed their operations.

[0071] It can be understood that when there are multiple TDNN layers in each D-TDNN layer, the result output by the FNN is input into multiple TDNNs respectively, and the results output by each TDNN are multiplied and then sent to the next structure.

[0072] In specific implementation, taking the example that one SE module is constructed in each D-TDNN layer for illustration, as Figure 5 shown, the SE module first performs spatial feature compression on the feature map, realizes global average pooling in the spatial dimension to obtain a 1×1×C feature map, learns through the FC fully connected layer to obtain a feature map with channel attention, and the dimension is 1×1×C. Multiply the feature map with channel attention of 1×1×C and the original input feature map of H×W×C channel by channel with weight coefficients, and finally output a feature map with channel attention, where : Parameter-free compression and preliminary excitation to generate channel statistics; : Parameterized learning of channel weights to dynamically enhance key features; : Apply the weight to the original feature to complete the adaptive adjustment.

[0073] It should be noted that in this embodiment, a direct connection is established between the inputs of two consecutive D-TDNN layers. The mathematical expression of the l-th layer D-TDNN is: . Among them, represents the input of the feature extraction module, represents the output of the l-th layer, Represents the non-linear transformation of this layer. In this embodiment, the depth of the D-TDNN network is increased. Through direct inter-layer connections, the network depth is allowed to increase significantly, enhancing the ability to model long-term acoustic contexts. The direct inter-layer connections enable backpropagation to bypass the deep non-linear transformations, alleviating the vanishing gradient problem. At the same time, the complexity is controlled by reducing the number of filter channels in each layer.

[0074] This embodiment provides a method for detecting abnormal sounds. In this embodiment, the two-dimensional feature representation corresponding to the target detection audio is input into the feature capture unit in the target detection model. The first feature extraction module of the feature capture unit extracts the temporal features from the two-dimensional feature representation to obtain the target temporal features corresponding to the target detection audio. The target temporal features are input into the squeeze-and-excitation module in the feature capture unit, and the squeeze-and-excitation module adjusts the weights of the target temporal features to obtain the weighted audio features of the target detection audio. The target detection audio is further processed by the second feature extraction module in the multiple feature extraction modules according to the weighted audio features to obtain the audio features to be detected corresponding to the target detection audio. Through the above method, by using densely connected D-TDNN layers and SE modules for feature extraction, accurate multi-scale temporal-frequency joint modeling is achieved, enabling dynamic and adaptive feature enhancement and noise suppression. This not only improves the learning ability and efficiency of the model but also enhances the generalization ability of the model and the ability to identify key features.

[0075] Based on the first embodiment and / or the second embodiment of the present application, in the third embodiment of the present application, the same or similar content as in the above-mentioned first and second embodiments can be referred to the above introduction and will not be elaborated hereinafter. On this basis, please refer to Figure 6 , before step S10 of the abnormal sound detection method, the method further includes steps S01 to S04: Step S01, data augmentation is performed according to multiple sample audios and the abnormal sound label information of each sample audio to obtain multiple model training samples and the abnormal sound label information of each model training sample.

[0076] It should be noted that the sample audio is the original audio file for training, including various types of abnormal sounds and normal audios. The abnormal sound label information is the category identification information of each sample audio, which is used to indicate whether there is an abnormal sound in the audio and the specific type of abnormal sound when there is an abnormal sound.

[0077] It can be understood that in order to ensure the diversity of samples during model training, data augmentation techniques (such as speed perturbation techniques, etc.) are used to process the sample audio to generate more training audios, and each generated training audio has its corresponding abnormal sound label information. In this embodiment, the model training samples are composed of a large number of sample audios and generated training audios.

[0078] In a specific implementation, the sample audio can be obtained by sending a swept-frequency signal containing multiple test frequencies to the target speaker and collecting the audio signal emitted by the target speaker, or can be obtained by other acquisition methods, and this embodiment does not limit this.

[0079] Step S02: Perform signal processing on each model training sample to obtain the filter bank features of each model training sample.

[0080] It should be noted that after obtaining multiple model training samples, frame splitting - fast Fourier transform and power spectrum calculation - Mel filtering and other processes are performed on the model training samples, so as to obtain the FBank features of multiple frames of signals corresponding to each model training sample. The FBank features of multiple frames of signals corresponding to each model training sample together constitute the filter bank features of each model training sample.

[0081] It can be understood that, for a further understanding of the signal processing process, the signal processing process is illustrated by way of example below. The parameters involved in the illustration process do not limit the processing process: the continuous speech of the module training sample is segmented into short-time segments with a 25 ms window (moving 10 ms each time, generating a 15 ms overlap), a Hamming window is applied to each frame, after obtaining the power spectrum by performing FFT on each frame, through a Mel filter bank composed of 80 triangular filters, the linear frequency is compressed into 80-dimensional Mel-scale energy values, a logarithmic operation is taken on the output of the filter bank, and 80-dimensional Fbank features of each frame are extracted.

[0082] Step S03: Set the parameters of the loss function according to a preset interval factor and a preset scaling factor to determine the target loss function.

[0083] It should be noted that the preset interval factor and the preset scaling factor are parameters set according to requirements. In this embodiment, the preset interval factor and the preset scaling factor are set to 0.2 and 32 respectively. At the same time, in this embodiment, the AAM-Softmax loss function with angular additive margin is selected, and the AAM-Softmax loss function is set using the preset interval factor and the preset scaling factor, so as to obtain the target loss function. The AAM-Softmax (Additive Angular Margin Softmax) loss function is an improved softmax loss function used in deep learning models. It introduces an angular margin on the basis of the traditional softmax loss to enhance the inter-class difference and reduce the intra-class variation, thereby improving the discrimination ability of the model.

[0084] Step S04: Train the initial detection model according to the target loss function, the target gradient descent optimizer, the target scheduler, the filter bank features of each model training sample, and the abnormal voice label information of each model training sample to obtain a target detection model. The initial detection model is composed of an initial front-end convolution module, multiple initial feature extraction modules, at least one initial squeeze-and-excitation module, and an initial detection classification unit.

[0085] It should be noted that, in this embodiment, the target gradient descent optimizer is exemplified by the stochastic gradient descent (SGD) optimizer; the target scheduler in this embodiment includes a cosine annealing scheduler and a linear warm-up scheduler; the learning rate is dynamically adjusted between 0.1 and 1e-4; the momentum parameter is set to 0.9; the weight decay is 1e-4. The above-mentioned scheduler, optimizer, learning rate, momentum parameter, and weight decay can all be adjusted according to requirements, and this embodiment does not limit this. The initial front-end convolution module is an untrained FCM; each initial feature extraction module is composed of multiple densely connected untrained D-TDNN layers; the initial squeeze-and-excitation module is an untrained SE module; the initial detection classification unit is an untrained detection classification unit, and its core components include a fully connected layer and an activation function.

[0086] It can be understood that by dividing multiple model training samples, a training set and a validation set are obtained. Use the audio in the training set and its corresponding abnormal voice label information, and combine the target gradient optimizer, the target scheduler, the learning rate, the target loss function, the momentum parameter, and the weight decay to train the initial detection model, so as to obtain a preliminarily trained detection model. During the training process, the target loss function is used to guide the model to optimize parameters and improve the classification accuracy.

[0087] In a specific implementation, use the validation set to test the preliminarily trained detection model, and calculate various performance indicators, including but not limited to accuracy, recall, and F1 score, etc. If the performance indicators of the detection model all meet the expected standards, it can be used as the final target detection model; otherwise, adjust the hyperparameters or structure of the preliminarily trained detection model and retrain to obtain the target detection model.

[0088] In a feasible implementation manner, step S01 may further include steps D11 to D13: Step D11: Randomly select in a preset rate ratio set to determine a target rate ratio.

[0089] It should be noted that the preset rate ratio set is composed of multiple rate ratios, and the rate ratio can be pre-set or randomly generated, which is not limited in this embodiment. A rate ratio is randomly selected from the preset rate ratios and used as the target rate ratio. For example, the preset rate ratio set is {0.9, 1.0, 1.1}.

[0090] Step D12: sampling each sample audio according to the target rate ratio to obtain multiple extended audios and the foreign sound label information of each extended audio.

[0091] It should be noted that, a plurality of sample audios to be processed are randomly selected from a plurality of sample audios, and each of the sample audios to be processed may correspond to the same target rate ratio, or each of the sample audios to be processed may correspond to different target rate ratios.

[0092] It can be understood that, by using the target rate ratio corresponding to each sample audio to be processed, audio sampling is performed on each sample audio to be processed. The specific process is: the speed of the audio is disturbed by the target rate ratio, and the playback speed of the audio is changed by a speed change algorithm (such as resampling or waveform stretching) to generate an extended audio. In this embodiment, it is ensured that the foreign sound label information of the extended audio is consistent with the foreign sound label information of the original sample audio.

[0093] Step D13, obtaining multiple model training samples and the foreign sound label information of each model training sample according to each extended audio and the foreign sound label information of each extended audio, each sample audio and the foreign sound label information of each sample audio.

[0094] It should be noted that all newly generated extended audios and all original sample audios are aggregated to obtain multiple model training samples; and based on the foreign sound label information of the newly generated extended audios and the foreign sound label information of each sample audio, the foreign sound label information of each model training sample is obtained.

[0095] This embodiment provides a heterophony detection method. In this embodiment, data augmentation is performed according to multiple sample audio and the heterophony label information of each sample audio to obtain multiple model training samples and the heterophony label information of each model training sample; signal processing is performed on each model training sample to obtain the filter bank features of each model training sample; the loss function is parameter - set according to a preset interval factor and a preset scaling factor to determine the target loss function; the initial detection model is trained according to the target loss function, the target gradient descent optimizer, the target scheduler, the filter bank features of each model training sample, and the heterophony label information of each model training sample to obtain the target detection model, where the initial detection model is composed of an initial front - end convolutional module, multiple initial feature extraction modules, at least one initial squeeze - and - excitation module, and an initial detection classification unit. In the above manner, by adjusting the loss function parameters and combining optimization strategies, the trained target detection model significantly improves the accuracy and robustness of heterophony detection, realizes efficient feature discrimination ability and strong generalization ability for complex acoustic scenes, and at the same time optimizes the training convergence speed and stability.

[0096] Exemplarily, to help understand the implementation process of the heterophony detection method obtained by combining the above - mentioned Embodiment 1 and Embodiment 2, please refer to Figure 7 , Figure 7A brief process schematic diagram of a heterophony detection method is provided. Specifically: S1, a swept-frequency signal containing multiple test frequencies is sent to the target speaker, and the audio signal (i.e., the sample audio) emitted by the target speaker is collected. S2, the continuous speech corresponding to the audio signal is segmented into short-time segments with a 25-ms window (moving 10 ms each time, generating a 15-ms overlap). A Hamming window is applied to each frame. After performing FFT on each frame to obtain the power spectrum, through a Mel filter bank composed of 80 triangular filters, the linear frequency is compressed into 80-dimensional Mel-scale energy values. The log operation is taken on the output of the filter bank to extract 80-dimensional Fbank features. Speed perturbation augmentation is performed by randomly sampling the speed ratio from {0.9, 1.0, 1.1}. S3, the network first inputs the acoustic features into the FCM network. The generated feature map is then flattened along the channel dimension and the frequency dimension and passed as input to the D-TDNN backbone network. The D-TDNN backbone network contains three modules, and each module consists of a series of D-TDNN layers. In each D-TDNN layer, an SE module is constructed. The number of layers in each module is adjusted to 12 / 24 / 16 layers to form deeper feature abstractions; the growth rate k of each module is reduced from 64 to 32 to construct narrower D-TDNN layers, and the increase in the number of parameters caused by the increase in depth is offset by reducing the number of channels in a single layer; an input TDNN layer with a downsampling rate of 1 / 2 is added before the D-TDNN backbone to accelerate the calculation. S4, set the training parameters of the model: select the AAM-Softmax loss function with angular additive margin, and its margin factor and scaling factor are set to 0.2 and 32 respectively. During the training process, the Stochastic Gradient Descent (SGD) optimizer is used, combined with a cosine annealing scheduler and a linear warm-up scheduler, and the learning rate is dynamically adjusted between 0.1 and 1e-4. The momentum parameter is set to 0.9, and the weight decay is 1e-4. S5, train the model: input the processed data set obtained in S2 into the improved network in S3, and train the network according to the parameters in S4 to obtain the target detection model. S6, model inference: during inference, first process the wav corresponding to the target detection audio through S2 to obtain Fbank features, and input them into the model trained in S5 for inference to obtain the result.

[0097] This embodiment solves the problems of limited capture of local features and insufficient modeling of long-term acoustic context in two-dimensional spectrograms and traditional CNN methods in abnormal sound detection of acoustic products. The processing methods mainly include the following aspects: (1) The front-end convolutional module (FCM) consists of multiple two-dimensional convolutional blocks with residual connections, encodes acoustic features in the time-frequency domain, and utilizes high-resolution time-frequency details. (2) D-TDNN adopts a dense connection mechanism, has parameter efficiency, and achieves better performance while requiring fewer parameters than traditional TDNN. (3) The SE module is used to learn feature weights according to the loss function through Squeeze, Excitation, and Scale operations, making the effective feature maps have larger weights. (4) A direct connection is established between the inputs of two consecutive D-TDNN layers, significantly increasing the depth of the D-TDNN network, and controlling the complexity by reducing the number of filter channels in each layer.

[0098] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the abnormal sound detection method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0099] This application also provides an abnormal sound detection device. Please refer to Figure 8 , and the abnormal sound detection device includes: An extraction module 10, configured to extract features from the two-dimensional feature representation corresponding to the target detection audio through a feature capture unit in the target detection model to obtain the audio features to be detected of the target detection audio. The feature capture unit consists of multiple feature extraction modules and at least one squeeze-and-excitation module, and the feature extraction module consists of multiple dense connection time-delay neural network layers.

[0100] A detection module 20, configured to input the audio features to be detected into a detection and classification unit in the target detection model for abnormal sound detection to obtain the category to which the audio of the target detection audio belongs.

[0101] A processing module 30, configured to obtain the audio detection result of the target detection audio according to the category to which the audio of the target detection audio belongs.

[0102] Optionally, the extraction module 10 is further configured to: Input the two-dimensional feature representation corresponding to the target detection audio into the feature capture unit in the target detection model. Through the first feature extraction module of the feature capture unit, perform temporal feature extraction on the two-dimensional feature representation to obtain the target temporal feature corresponding to the target detection audio; input the target temporal feature into the squeeze-and-excitation module in the feature capture unit, and through the squeeze-and-excitation module, perform weight adjustment on the target temporal feature to obtain the weighted audio feature of the target detection audio; according to the weighted audio feature and the second feature extraction module in multiple feature extraction modules, perform feature extraction on the target detection audio to obtain the audio feature to be detected corresponding to the target detection audio.

[0103] Optionally, the extraction module 10 is further configured to: Input the two-dimensional feature representation into the first feature extraction module of the feature capture unit, and through the first feedforward neural network in the first feature extraction module, perform feature transformation on the two-dimensional feature representation to obtain the first audio feature of the target detection audio. The first feedforward neural network is the basic unit of the first densely connected time-delay neural network layer in multiple densely connected time-delay neural network layers; input the first audio feature into the time-delay neural network in the first densely connected time-delay neural network layer respectively, perform temporal feature extraction on the first audio feature through the time-delay neural network, and obtain the target temporal feature corresponding to the target detection audio according to the feature extraction result of the time-delay neural network and the hierarchical structure of the first feature extraction module.

[0104] Optionally, the extraction module 10 is further configured to: In response to the abnormal sound detection instruction, perform signal processing on the target detection audio to obtain the filter bank feature of the target detection audio; input the filter bank feature into the front-end convolution module in the target detection model, perform local feature extraction on the filter bank feature through the front-end convolution module, and obtain the two-dimensional feature representation corresponding to the target detection audio according to the extraction result.

[0105] Optionally, the extraction module 10 is further configured to: In response to the abnormal sound detection instruction, perform frame division processing on the target detection audio according to a preset window to obtain multiple frame audio signals corresponding to the target detection audio; perform windowing processing on each audio signal respectively to obtain the windowed signals corresponding to each audio signal respectively; perform power spectrum calculation on each windowed signal respectively to determine the linear frequency power spectrum of each windowed signal; perform Mel filter bank processing on the linear frequency power spectrum of each windowed signal respectively, and obtain the filter bank feature of each windowed signal according to the processing result; determine the filter bank feature of the target detection audio according to the filter bank features of each windowed signal.

[0106] Optionally, the extraction module 10 is further configured to: Perform data augmentation based on multiple sample audio and the abnormal sound label information of each sample audio to obtain multiple model training samples and the abnormal sound label information of each model training sample; perform signal processing on each model training sample to obtain the filter bank features of each model training sample; set the parameters of the loss function according to a preset interval factor and a preset scaling factor to determine the target loss function; train the initial detection model according to the target loss function, the target gradient descent optimizer, the target scheduler, the filter bank features of each model training sample, and the abnormal sound label information of each model training sample to obtain the target detection model, where the initial detection model is composed of an initial front-end convolutional module, multiple initial feature extraction modules, at least one initial squeeze-and-excitation module, and an initial detection classification unit.

[0107] Optionally, the extraction module 10 is further configured to: Randomly select in a preset set of rate ratios to determine the target rate ratio; perform audio sampling on each sample audio according to the target rate ratio to obtain multiple extended audio and the abnormal sound label information of each extended audio; obtain multiple model training samples and the abnormal sound label information of each model training sample according to the extended audio and the abnormal sound label information of each extended audio, and the sample audio and the abnormal sound label information of each sample audio.

[0108] The abnormal sound detection device provided by this application adopts the abnormal sound detection method in the above embodiment, and can solve the technical problem that the existing abnormal sound detection technology cannot achieve high-precision, low-cost, and strong generalization abnormal sound detection in complex acoustic environments and diverse product lines. Compared with the existing technology, the beneficial effects of the abnormal sound detection device provided by this application are the same as those of the abnormal sound detection method provided by the above embodiment, and other technical features in the abnormal sound detection device are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.

[0109] This application provides an abnormal sound detection device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the abnormal sound detection method in the first embodiment above.

[0110] Next, refer to Figure 9, which shows a schematic structural diagram of a heterophonic sound detection device suitable for implementing the embodiments of the present application. The heterophonic sound detection device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The shown heterophonic sound detection device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0111] As Figure 9 shown, the heterophonic sound detection device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the heterophonic sound detection device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the heterophonic sound detection device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a heterophonic sound detection device having various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems may be implemented or had.

[0112] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by a processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0113] The abnormal sound detection device provided by the present application adopts the abnormal sound detection method in the above embodiment, and can solve the technical problem that the existing abnormal sound detection technology cannot achieve high-precision, low-cost, and strong generalization abnormal sound detection in complex acoustic environments and diverse product lines. Compared with the prior art, the beneficial effects of the abnormal sound detection device provided by the present application are the same as those of the abnormal sound detection method provided by the above embodiment, and other technical features in the abnormal sound detection device are the same as those disclosed in the method of the previous embodiment, which will not be elaborated here.

[0114] It should be understood that the various parts disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0115] As mentioned above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0116] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the abnormal sound detection method in the above embodiment.

[0117] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0118] The above computer-readable storage medium can be included in the abnormal sound detection device; it can also exist independently without being assembled into the abnormal sound detection device.

[0119] The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the abnormal sound detection device, the abnormal sound detection device is caused to: extract features from the two-dimensional feature representation corresponding to the target detection audio through the feature capture unit in the target detection model to obtain the audio features to be detected of the target detection audio, the feature capture unit is composed of multiple feature extraction modules and at least one squeeze-and-excitation module, and the feature extraction module is composed of multiple densely connected delay neural network layers; input the audio features to be detected into the detection and classification unit in the target detection model for abnormal sound detection to obtain the audio category to which the target detection audio belongs; and obtain the audio detection result of the target detection audio according to the audio category to which the target detection audio belongs.

[0120] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN: Local Area Network) or a wide area network (WAN: Wide Area Network), or it can be connected to an external computer (for example, by connecting through an Internet service provider using the Internet).

[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0122] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0123] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned abnormal sound detection method, and can solve the technical problem that the existing abnormal sound detection technology cannot achieve high-precision, low-cost, and strong generalization abnormal sound detection in complex acoustic environments and diverse product lines. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the abnormal sound detection method provided by the above embodiments, and will not be elaborated here.

[0124] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the abnormal sound detection method as described above.

[0125] The computer program product provided by the present application can solve the technical problem that the existing abnormal sound detection technology cannot achieve high-precision, low-cost and strong generalization abnormal sound detection in complex acoustic environments and diverse product lines. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the abnormal sound detection method provided by the above embodiments, and will not be elaborated here.

[0126] The above are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A method for detecting abnormal sound, characterized in that: The method comprises: Extracting features from the two-dimensional feature representation corresponding to the target detection audio through a feature capture unit in the target detection model to obtain the audio features to be detected of the target detection audio, wherein the feature capture unit is composed of a plurality of feature extraction modules and at least one compression excitation module, and the feature extraction module is composed of a plurality of densely connected time-delay neural network layers; Inputting the audio feature to be detected into the detection classification unit in the target detection model to perform abnormal sound detection, and obtaining the audio category to which the target detection audio belongs; The audio detection result of the target detection audio is obtained according to the audio category of the target detection audio.

2. The method according to claim 1, characterized in that The step of extracting features from the two-dimensional feature representation corresponding to the target detection audio by a feature capture unit in the target detection model to obtain the audio features to be detected of the target detection audio comprises: Inputting the two-dimensional feature representation corresponding to the target detection audio into a feature capture unit in the target detection model, and performing time series feature extraction on the two-dimensional feature representation through a first feature extraction module of the feature capture unit to obtain a target time series feature corresponding to the target detection audio; Inputting the target time series feature into the compression excitation module in the feature capture unit, and adjusting the weight of the target time series feature by the compression excitation module to obtain the weighted audio feature of the target detection audio; The target detection audio is subjected to feature extraction according to the weighted audio feature and a second feature extraction module among the plurality of feature extraction modules to obtain an audio feature to be detected corresponding to the target detection audio.

3. The method according to claim 2, characterized in that The densely connected time-delay neural network layer includes a feedforward neural network and at least one time-delay neural network; The step of extracting time series features from the two-dimensional feature representation by a first feature extraction module of a feature capture unit to obtain target time series features corresponding to the target detection audio includes: Inputting the two-dimensional feature representation into a first feature extraction module of the feature capture unit, performing feature conversion on the two-dimensional feature representation through a first feedforward neural network in the first feature extraction module to obtain a first audio feature of the target detection audio, wherein the first feedforward neural network is a basic unit of a first densely connected time-delay neural network layer among a plurality of densely connected time-delay neural network layers; The first audio features are respectively input into the time-delay neural network in the first densely connected time-delay neural network layer, and the time-series features of the first audio features are extracted by the time-delay neural network. The target time-series features corresponding to the target detection audio are obtained according to the feature extraction results of the time-delay neural network and the hierarchical structure of the first feature extraction module.

4. The method according to any one of claims 1 to 3, characterized in that Before the step of extracting features from the two-dimensional feature representation corresponding to the target detection audio by the feature capture unit in the target detection model to obtain the audio features to be detected of the target detection audio, the step further includes: In response to the abnormal sound detection instruction, signal processing is performed on the target detection audio to obtain a filter bank feature of the target detection audio; The filter group features are input into a front-end convolution module in a target detection model, local features of the filter group features are extracted by the front-end convolution module, and a two-dimensional feature representation corresponding to the target detection audio is obtained according to the extraction result.

5. The method according to claim 4, characterized in that The step of performing signal processing on the target detection audio in response to the abnormal sound detection instruction to obtain the filter bank characteristics of the target detection audio comprises: In response to the abnormal sound detection instruction, the target detection audio is framed according to a preset window to obtain a plurality of frames of audio signals corresponding to the target detection audio; Perform windowing processing on each audio signal respectively to obtain a windowed signal corresponding to each audio signal; Performing power spectrum calculation on each windowed signal respectively to determine the linear frequency power spectrum of each windowed signal; Performing Mel filter bank processing on the linear frequency power spectrum of each windowed signal respectively, and obtaining the filter bank characteristics of each windowed signal according to the processing results; The filter bank characteristics of the target detection audio are determined according to the filter bank characteristics of each windowed signal.

6. The method according to any one of claims 1 to 3, characterized in that Before the step of extracting features from the two-dimensional feature representation corresponding to the target detection audio by the feature capture unit in the target detection model to obtain the audio features to be detected of the target detection audio, the step further includes: Performing data enhancement based on multiple sample audios and the label information of the foreign voice of each sample audio, to obtain multiple model training samples and the label information of the foreign voice of each model training sample; Perform signal processing on each model training sample to obtain filter bank features of each model training sample; Setting parameters of the loss function according to a preset interval factor and a preset scaling factor to determine a target loss function; The initial detection model is trained according to the target loss function, the target gradient descent optimizer, the target scheduler, the filter group characteristics of each model training sample and the heterophonic label information of each model training sample to obtain a target detection model. The initial detection model is composed of an initial front-end convolution module, multiple initial feature extraction modules, at least one initial compression excitation module and an initial detection classification unit.

7. The method according to claim 6, characterized in that The step of performing data enhancement according to the plurality of sample audios and the heterophonic label information of each sample audio to obtain the plurality of model training samples and the heterophonic label information of each model training sample comprises: Randomly select from a preset rate ratio set to determine a target rate ratio; Performing audio sampling on each sample audio according to the target rate ratio to obtain multiple extended audios and the foreign sound label information of each extended audio; According to each extended audio and the foreign sound label information of each extended audio, each sample audio and the foreign sound label information of each sample audio, multiple model training samples and the foreign sound label information of each model training sample are obtained.

8. A device for detecting abnormal sound, characterized in that: The abnormal sound detection device comprises: An extraction module is used to extract features from a two-dimensional feature representation corresponding to the target detection audio through a feature capture unit in the target detection model to obtain an audio feature to be detected of the target detection audio, wherein the feature capture unit is composed of a plurality of feature extraction modules and at least one compression excitation module, and the feature extraction module is composed of a plurality of densely connected time-delay neural network layers; A detection module, used for inputting the audio feature to be detected into the detection classification unit in the target detection model to perform abnormal sound detection, and obtaining the audio category to which the target detection audio belongs; A processing module is used to obtain an audio detection result of the target detection audio according to the audio category of the target detection audio.

9. An abnormal sound detection device, characterized in that: The device comprises: a memory, a processor, and an abnormal sound detection program stored in the memory and executable on the processor, wherein the abnormal sound detection program is configured to implement the steps of the abnormal sound detection method according to any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium stores an abnormal sound detection program, and when the abnormal sound detection program is executed by the processor, the steps of the abnormal sound detection method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Abnormal event detection method and device and electronic equipment

    CN113838478A

  • Multi-modal fusion mechanical defect detection method based on attention map convolutional network

    CN114170477A

  • Sound event detection and positioning method based on combined convolutional neural network

    CN115631771A

  • Synthetic speech detection method based on neural network and feature fusion

    CN117393000A

  • Abnormal sound detection method and system of neural network based on multiple receptive fields

    CN119091909A

Cited By

  • Audio equipment fault detection method and device, electronic equipment and readable storage medium

    CN121331165A