Abnormal sound detection method and device, equipment and storage medium

By adopting a Transformer-based antonal detection framework in acoustic detection, the problem of difficult to separate abnormal acoustic characteristics and background interference in traditional methods is solved, and high-precision and robust antonal detection is achieved, and the acoustic characteristics differences are adapted to different models of acoustic products.

CN120089162AActive Publication Date: 2025-06-03GOERTEK INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510560161.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

Traditional acoustic detection methods are difficult to accurately separate abnormal acoustic features and background interference, resulting in low detection accuracy and generalization ability.

Method used

A Transformer-based anomaly detection framework is adopted to dynamically capture the global correlation between time-frequency domain depth characterization and self-attention mechanism in acoustic signals, and multi-head attention is used to extract multi-scale anomaly mode features in parallel, and combine lightweight encoder design to achieve efficient feature decoupling in complex noise environments.

Benefits of technology

The model's sensitivity to weak abnormal sounds is significantly improved, and the abnormal sound detection is achieved with high robust abnormal sound detection, and the parameter fine-tuning strategy is used to quickly adapt the acoustic characteristics of different models of acoustic products, solving the problem of high cost of migration across product lines and large fluctuations in the misjudgment rate of traditional models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089162A_ABST
    Figure CN120089162A_ABST
Patent Text Reader

Abstract

The invention discloses an abnormal sound detection method, device and equipment and a storage medium, and the abnormal sound detection method comprises the steps: carrying out the feature extraction of a to-be-detected audio through a feature extraction unit in an abnormal sound detection model in response to an abnormal sound detection instruction, and obtaining a target audio feature of the to-be-detected audio, the feature extraction unit is composed of a plurality of twin module groups with mixed attention mechanisms and a semantic feature extraction module, target audio features are input into the abnormal sound classification unit in the model for abnormal sound detection, the category of the audio is obtained, and therefore the audio detection result of the audio to be detected is obtained. Through the above mode, the time-frequency domain depth representation and the self-attention mechanism are fused, the global relevance across time steps in the acoustic signal is dynamically captured, and multi-scale abnormal mode features are extracted in parallel by using multi-head attention, so that efficient feature decoupling in a complex noise environment is realized. And the sensitivity of the model to weak abnormal sound and the detection robustness are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of acoustic technologies, and particularly to a method, device, equipment, and storage medium for detecting abnormal sounds. Background Art

[0002] The acoustic quality of acoustic products is one of the core elements determining user experience and market competitiveness. Acoustic products such as headphones, speakers, and microphones need to undergo a strict abnormal sound detection process before leaving the factory to ensure that there are no defects such as noise, distorted sound, abnormal background noise, or structural resonance during their operation. Traditional detection methods highly rely on the subjective judgment of human listeners. By repeatedly playing standardized audio samples and relying on the human ear to identify abnormalities. However, with the exponential growth of the shipment volume of acoustic products and the accelerating speed of product iteration, the drawbacks of manual detection have become increasingly prominent: low detection efficiency, rising labor costs, and the human ear is susceptible to fatigue, individual hearing differences, and environmental noise interference, resulting in difficulty in ensuring detection consistency. Although some enterprises have introduced automated detection equipment, the existing technologies still have difficulty in achieving high-precision and strong generalization abnormal sound determination in complex acoustic scenarios and diverse product lines.

[0003] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of this application is to provide a method, device, equipment, and storage medium for detecting abnormal sounds, aiming to solve the technical problem that traditional acoustic detection methods have low detection accuracy and generalization ability because it is difficult to accurately separate abnormal acoustic features from background interference.

[0005] To achieve the above purpose, this application proposes a method for detecting abnormal sounds, and the method includes: In response to an abnormal sound detection instruction, the feature extraction unit in the abnormal sound detection model extracts features from the audio to be detected to obtain the target audio features of the audio to be detected. The feature extraction unit is composed of multiple twin module groups and a semantic feature extraction module, and there is a hybrid attention mechanism in the twin module groups; Input the target audio features into the abnormal sound classification unit in the abnormal sound detection model for abnormal sound detection to obtain the audio category to which the audio to be detected belongs; Obtain the audio detection result of the audio to be detected according to the audio category to which the audio to be detected belongs.

[0006] In one embodiment, the step of extracting features from the audio to be detected by the feature extraction unit in the abnormal sound detection model to obtain the target audio features of the audio to be detected includes: Perform signal conversion on the audio to be detected to obtain the audio spectrogram of the audio to be detected; Performing a cutting process on multiple windows corresponding to the audio spectrogram through a block embedding neural network to obtain multiple block tags corresponding to the audio spectrogram; Inputting the multiple block tags corresponding to the audio spectrogram into a feature extraction unit in an outlier detection model, and performing feature extraction on the multiple block tags through multiple twin module groups in the feature extraction unit to obtain associated multi-level features corresponding to the audio spectrogram; Performing semantic information aggregation on the audio spectrogram according to the associated multi-level features and a semantic feature extraction module in the feature extraction unit to obtain target audio features corresponding to the audio spectrogram.

[0007] In one embodiment, the step of performing feature extraction on multiple block tags through multiple twin module groups in the feature extraction unit to obtain associated multi-level features corresponding to the audio spectrogram includes: Inputting the multiple block tags into multiple twin module groups in the feature extraction unit, and performing feature extraction on the multiple block tags through a first twin module group in the multiple twin module groups to obtain first-level features corresponding to each block tag; Merging the multiple block tags according to the first-level features corresponding to each block tag to obtain multiple merged blocks and second-level features corresponding to each merged block; Performing feature extraction on the audio spectrogram according to the second-level features corresponding to each merged block and a second twin module group in the multiple twin module groups to obtain third-level features corresponding to the audio spectrogram; Performing feature processing on the third-level features corresponding to the audio spectrogram according to a convolutional layer with a preset size to obtain associated multi-level features corresponding to the audio spectrogram.

[0008] In one embodiment, the step of performing feature extraction on multiple block tags through a first twin module group in the multiple twin module groups to obtain first-level features corresponding to each block tag includes: Inputting the multiple block tags into a first twin module group in the multiple twin module groups, and performing local feature extraction on the multiple block tags through a local attention mechanism in the first twin module group to obtain local feature representations of each block tag; Performing global feature extraction on the multiple block tags according to the local feature representations of each block tag through a global attention mechanism in the first twin module group to obtain global feature representations of each block tag; Determining first-level features corresponding to each block tag according to the global feature representations of each block tag and the hierarchical structure of the first twin module group.

[0009] In one embodiment, the step of inputting the target audio feature into the abnormal sound classification unit of the abnormal sound detection model for abnormal sound detection to obtain the audio category to which the audio to be detected belongs includes: Input the target audio feature into the abnormal sound classification unit of the abnormal sound detection model, perform feature mapping on the target audio feature through the fully connected layer in the abnormal sound classification unit, and obtain the category probability map of the audio to be detected according to the mapping result; Perform global average pooling operation on the category probability map through the abnormal sound classification unit to obtain the category probability vector of the audio to be detected; Perform category determination according to the category probability vector of the audio to be detected to determine the audio category to which the audio to be detected belongs.

[0010] In one embodiment, before the step of extracting the target audio feature of the audio to be detected through the feature extraction unit in the abnormal sound detection model, the following steps are further included: Generate samples according to multiple sample audios and the abnormal sound classification labels of each sample audio to obtain multiple model training audios and the abnormal sound classification labels of each model training audio; Divide multiple model training audios to obtain a model training set and a model validation set; Train the initial detection model according to the model training set, the abnormal sound classification labels of each audio in the model training set, the learning rate warm-up strategy, and the target optimizer to obtain the first detection model, where the initial detection model is composed of multiple initial twin module groups, an initial semantic feature extraction module, and an initial abnormal sound classification unit; Evaluate the performance of the first detection model according to the model validation set, and determine the abnormal sound detection model according to the performance evaluation result.

[0011] In one embodiment, the step of generating samples according to multiple sample audios and the abnormal sound classification labels of each sample audio to obtain multiple model training audios and the abnormal sound classification labels of each model training audio includes: Randomly extract multiple sample audios to obtain multiple sample audio pairs; Perform linear interpolation on multiple sample audio pairs, and generate the first generated audio of each sample audio pair and the abnormal sound classification label of each first generated audio according to the linear interpolation result; Randomly mask each first generated audio and each sample audio, and generate multiple second generated audios and the abnormal sound classification labels of each second generated audio according to the masking result; Based on each sample audio and its abnormal sound classification label, each first generated audio and its abnormal sound classification label, and each second generated audio and its abnormal sound classification label, multiple model training audios and their abnormal sound classification labels are obtained.

[0012] In addition, to achieve the above object, the present application also proposes an abnormal sound detection device, which includes: an extraction module, configured to respond to an abnormal sound detection instruction, extract features of the audio to be detected through a feature extraction unit in the abnormal sound detection model to obtain the target audio features of the audio to be detected. The feature extraction unit is composed of multiple twin module groups and a semantic feature extraction module, and there is a hybrid attention mechanism in the twin module groups; A detection module, configured to input the target audio features into an abnormal sound classification unit in the abnormal sound detection model for abnormal sound detection to obtain the category to which the audio of the audio to be detected belongs; A processing module, configured to obtain the audio detection result of the audio to be detected according to the category to which the audio of the audio to be detected belongs.

[0013] In addition, to achieve the above object, the present application also proposes an abnormal sound detection device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The computer program is configured to implement the steps of the abnormal sound detection method as described above.

[0014] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the abnormal sound detection method as described above are implemented.

[0015] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the abnormal sound detection method as described above are implemented.

[0016] The present application provides a method for detecting abnormal sounds. In response to an abnormal sound detection instruction, the present application extracts features of the audio to be detected through a feature extraction unit in an abnormal sound detection model, obtaining the target audio features of the audio to be detected. The feature extraction unit is composed of multiple twin module groups and a semantic feature extraction module, and there is a hybrid attention mechanism in the twin module groups; the target audio features are input into the abnormal sound classification unit in the abnormal sound detection model for abnormal sound detection, obtaining the audio category to which the audio to be detected belongs; the audio detection result of the audio to be detected is obtained according to the audio category to which the audio to be detected belongs. In the above manner, by fusing time-frequency domain depth representation and self-attention mechanism, dynamically capturing the global correlation across time steps in acoustic signals, using multi-head attention to parallelly extract multi-scale abnormal pattern features, and combining lightweight encoder design to achieve efficient feature decoupling in complex noise environments, the sensitivity of the model to weak abnormal sounds is significantly improved, and high-robustness abnormal sound detection is achieved. At the same time, it can quickly adapt to the acoustic characteristic differences of different models of acoustic products through parameter fine-tuning strategies, providing an innovative technical path for solving the problems of high cross-product line migration cost and large fluctuation of false positive rates in traditional models. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the abnormal sound detection method of the present application; Figure 2 It is a schematic diagram of the network architecture provided for Embodiment 1 of the present application; Figure 3 It is a schematic flowchart provided for Embodiment 2 of the abnormal sound detection method of the present application; Figure 4 It is a schematic diagram of windowed cutting provided for Embodiment 2 of the present application; Figure 5 It is a schematic diagram of the flow of the attention mechanism provided for Embodiment 2 of the present application Figure 6 It is a schematic flowchart provided for Embodiment 3 of the abnormal sound detection method of the present application; Figure 7 It is a schematic flowchart of the brief abnormal sound detection method provided for Embodiment 3 of the present application; Figure 8 Schematic diagram of the module structure of the abnormal sound detection device according to an embodiment of the present application; Figure 9 Schematic diagram of the device structure of the hardware operating environment involved in the abnormal sound detection method according to an embodiment of the present application.

[0020] The implementation, functional features and advantages of the present application will be further described with reference to the accompanying drawings in combination with embodiments. Detailed implementation manners

[0021] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0022] In order to better understand the technical solutions of the present application, the following will be described in detail in combination with the accompanying drawings of the specification and specific implementation manners.

[0023] The main solution of the embodiment of the present application is: in response to an abnormal sound detection instruction, a feature extraction unit in the abnormal sound detection model extracts features from the audio to be detected to obtain the target audio features of the audio to be detected. The feature extraction unit is composed of a plurality of siamese module groups and a semantic feature extraction module, and there is a hybrid attention mechanism in the siamese module groups; the target audio features are input into the abnormal sound classification unit in the abnormal sound detection model for abnormal sound detection to obtain the audio category to which the audio to be detected belongs; the audio detection result of the audio to be detected is obtained according to the audio category to which the audio to be detected belongs.

[0024] The existing abnormal sound detection technologies for acoustic products mainly focus on signal analysis and traditional machine learning. At the signal analysis level, common methods include frequency response curve comparison, total harmonic distortion (THD) test, signal-to-noise ratio (SNR) calculation, etc. Whether there is an abnormality is determined by setting a threshold. Although such methods can capture some obvious defects, they lack sensitivity to subtle abnormal sounds (such as high-frequency whistling and intermittent background noise) that are sensitive to the human ear, and are easily affected by the environmental noise of the production line or the individual differences of products. At the machine learning level, classification models based on algorithms such as support vector machine (SVM) and random forest are used to classify abnormalities by extracting time-domain features (such as zero-crossing rate, energy entropy) or frequency-domain features (such as Mel spectrum, cepstral coefficients) of acoustic signals. However, the abnormal sound patterns of acoustic products are highly non-linear and random, and it is difficult for traditional feature engineering to cover such complex scenarios. In recent years, deep learning solutions based on convolutional neural network (CNN) have been tried for end-to-end detection, but their modeling ability is limited, and their adaptability to product design changes is poor. Frequent retraining of the model restricts the feasibility of large-scale deployment.

[0025] Due to the characteristics of strong randomness in the time-frequency distribution and low proportion of abnormal energy of the abnormal sounds of acoustic products, it is difficult for traditional acoustic detection methods to accurately separate abnormal acoustic features from background interference under the influence of noise interference and product structure differences, resulting in high sensitivity of the detection results to working conditions and weak cross-model generalization ability. This makes the high-precision and low false positive rate abnormal sound detection of acoustic products in industrial scenarios an urgent technical bottleneck to be broken through.

[0026] In view of the deficiencies of traditional methods in modeling long-term dependencies and the large limitations of manual feature design, this application proposes an abnormal sound detection framework based on Transformer: by fusing time-frequency domain deep representations and self-attention mechanisms, dynamically capturing the global correlation across time steps in acoustic signals, using multi-head attention to parallelly extract multi-scale abnormal pattern features, and combining lightweight encoder design to achieve efficient feature decoupling in complex noise environments. This solution can significantly improve the sensitivity of the model to weak abnormal sounds, and at the same time quickly adapt to the acoustic characteristic differences of different models of acoustic products through the parameter fine-tuning strategy, providing an innovative technical path to solve the problems of high cross-product line migration cost and large fluctuation of false positive rate of traditional models.

[0027] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a heterophony detection device, etc. that can implement the above functions. Hereinafter, taking the heterophony detection device as an example, this embodiment and the following embodiments will be described.

[0028] Based on this, the embodiment of the present application provides a heterophony detection method, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the heterophony detection method of the present application.

[0029] In this embodiment, the heterophony detection method includes steps S10 to S30: Step S10, in response to a heterophony detection instruction, extract features of the audio to be detected through a feature extraction unit in the heterophony detection model to obtain the target audio features of the audio to be detected. The feature extraction unit is composed of multiple twin module groups and a semantic feature extraction module, and there is a hybrid attention mechanism in the twin module groups.

[0030] It should be noted that the heterophony detection instruction is used to trigger the heterophony detection process of the acoustic product to determine whether there is a heterophony in the audio emitted by the acoustic product. The heterophony detection instruction can be actively initiated by the heterophony detection device when it detects the acoustic product, or can be initiated by the user to the heterophony detection device. This embodiment does not limit this. Heterophony refers to an unexpected and non-design-permitted abnormal sound signal generated under the normal working state of the acoustic product, and its manifestation is significantly different from the normal acoustic output. Heterophony includes but is not limited to types such as noise, broken sound, abnormal background noise, and structural resonance.

[0031] It can be understood that the heterophony detection model uses Twins-SVT as the backbone network and is composed of multiple twin module groups, a semantic feature extraction module, and a heterophony classification unit. The Twins-SVT backbone network adopts the Spatially Separable Self-Attention (SVT) mechanism, which is a model architecture containing multiple modules, combining local and global spatial self-attention mechanisms to effectively process data. The Twins-SVT architecture optimizes the computational complexity problem encountered by the traditional Transformer in processing data in a structured manner. The heterophony detection model can gradually reduce the resolution and increase the number of channels, enabling it to efficiently process features of different scales.

[0032] It should be understood that the rise of the Transformer model has provided a brand-new technological breakthrough for the detection of abnormal sounds in acoustic products. The global context modeling ability initially demonstrated by this model in natural language processing happens to fit the long-time sequence dependence characteristics of acoustic signals. Therefore, this embodiment proposes an algorithm that can fuse the deep representations in the time-frequency domain and utilize the global attention mechanism of the Transformer to achieve highly robust abnormal sound detection, which has breakthrough significance for improving the cross-model adaptability, complex noise suppression ability, and weak abnormal signal capture accuracy of industrial acoustic product abnormal sound detection.

[0033] In the specific implementation, the abnormal sound detection model is specifically set as 4 twin module groups. The 4 twin module groups respectively contain 3, 3, 9, and 3 levels to adapt to the time sequence-frequency domain feature interaction. Each level's corresponding twin module (i.e., the Twins module) has a hybrid attention mechanism, namely the local attention mechanism (Local Spatial Attention, LSA) and the global attention mechanism (Global Spatial Attention, GSA). In this embodiment, the local attention mechanism is used to calculate self-attention within the window, capture fine-grained features, and reduce the computational complexity. The global attention mechanism is used to supplement global information (such as long-time event associations in audio) through cross-window sparse sampling or shifted windows. It should be noted that the semantic feature extraction module refers to the labeled semantic CNN layer, which is used to integrate all frequency interval information. After the last twin module group, a convolutional layer with a preset size is also added. The specific connection method of the feature extraction unit is as Figure 2 shown. Each twin-spatial vision transformer in the figure represents a twin module group. Each twin module in each twin module group has a hybrid attention mechanism (i.e., global attention and local attention). After the last twin module group, a new convolutional layer is added, and then the output of the convolutional layer is subjected to a dimension reorganization operation. Finally, the result of the reshaping operation is sent to the semantic feature extraction module (i.e., the labeled semantic convolutional neural network shown in the figure).

[0034] It can be understood that the abnormal sound detection model is obtained by training the initial detection model with a large number of audio with abnormal sound classification labels. The initial detection model consists of multiple untrained twin module groups, a semantic feature extraction module, and an abnormal sound classification unit.

[0035] After receiving the abnormal sound detection instruction, obtain the audio to be detected generated by the acoustic product that needs to be detected for abnormal sound. Extract features from multiple patch tokens corresponding to the audio to be detected through the feature extraction unit in the abnormal sound detection module, capture the global correlation across time in the audio to be detected, and use multi-head attention to parallelly extract multi-scale abnormal pattern features, so as to obtain the target audio features of the audio to be detected. The target audio features include time-frequency local details and global semantics.

[0036] Step S20: Input the target audio features into the abnormal sound classification unit in the abnormal sound detection model for abnormal sound detection, and obtain the category to which the audio to be detected belongs.

[0037] It should be noted that the abnormal sound classification unit is the decision-making layer in the abnormal sound detection model, responsible for mapping the target audio features to the category mapping space and completing the abnormal sound detection task. The core components of the abnormal sound classification unit include, but are not limited to, the fully connected layer, global average pooling, and activation function.

[0038] It can be understood that inputting the target audio features into the abnormal sound classification unit in the abnormal sound detection model, the abnormal sound classification unit maps the target audio features to the category mapping space, and obtains the category to which the audio to be detected belongs. In this embodiment, the category to which the audio belongs refers to the category label to which the audio to be detected is assigned, including but not limited to normal audio, noise, broken sound, etc.

[0039] In a feasible implementation manner, step S20 may include steps A11 to A13: Step A11: Input the target audio features into the abnormal sound classification unit in the abnormal sound detection model, perform feature mapping on the target audio features through the fully connected layer in the abnormal sound classification unit, and obtain the category probability map of the audio to be detected according to the mapping result.

[0040] It should be noted that when performing audio classification through the abnormal sound classification unit, first input the target audio features into the fully connected layer in the abnormal sound classification unit. The fully connected layer maps the channel dimension to the number of categories, generates the original category scores at each time point, and finally outputs the category probability map of the audio to be detected according to the original category scores at the time points generated after mapping. For example, if the size of the target audio features is (T / (8P), 8D), then the fully connected layer maps the feature channels from 8D to the number of categories C, and the output category probability map is a (T / (8P), C) matrix.

[0041] Step A12: Perform a global average pooling operation on the category probability map through the abnormal sound classification unit to obtain the category probability vector of the audio to be detected.

[0042] It should be noted that the category probability map is averaged along the time dimension by the abnormal sound classification unit, and the time resolution is compressed into a single vector, so as to obtain a global category probability vector, which is used to represent the comprehensive probability that the audio to be detected belongs to each category. For example, if the category probability map is a (T / (8P), C) matrix, then the category probability vector is (1, C).

[0043] Step A13, perform category determination according to the category probability vector of the audio to be detected, and determine the category to which the audio of the audio to be detected belongs.

[0044] It should be noted that an activation function (such as the Softmax function) is used to generate a normalized probability distribution based on the category probability vector, and the category with the highest probability is taken as the prediction result, so as to obtain the category to which the audio of the audio to be detected belongs.

[0045] Step S30, obtain the audio detection result of the audio to be detected according to the category to which the audio of the audio to be detected belongs.

[0046] It should be noted that after the abnormal sound detection model identifies the category to which the audio of the audio to be detected belongs, it is possible to clarify whether there is an abnormal sound in the audio to be detected, and when there is an abnormal sound, the specific category of the abnormal sound, so as to obtain the final audio detection result.

[0047] This embodiment provides an abnormal sound detection method. In this embodiment, in response to an abnormal sound detection instruction, the feature extraction unit in the abnormal sound detection model extracts features from the audio to be detected to obtain the target audio features of the audio to be detected. The feature extraction unit is composed of multiple twin module groups and a semantic feature extraction module, and there is a hybrid attention mechanism in the twin module groups; the target audio features are input into the abnormal sound classification unit in the abnormal sound detection model for abnormal sound detection to obtain the category to which the audio of the audio to be detected belongs; the audio detection result of the audio to be detected is obtained according to the category to which the audio of the audio to be detected belongs. By the above method, by fusing time-frequency domain depth representation and self-attention mechanism, dynamically capturing the global correlation across time steps in the acoustic signal, using multi-head attention to parallelly extract multi-scale abnormal pattern features, and combining lightweight encoder design to achieve efficient feature decoupling in complex noise environments, the sensitivity of the model to weak abnormal sounds is significantly improved, and high-robustness abnormal sound detection is realized. At the same time, it can quickly adapt to the acoustic characteristic differences of different models of acoustic products through parameter fine-tuning strategies, providing an innovative technical path for solving the problems of high cross-product line migration cost and large fluctuation of misjudgment rate of traditional models.

[0048] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3, Step S10, the abnormal sound detection method further includes steps S11 to S14: Step S11, in response to an abnormal sound detection instruction, perform signal conversion on the audio to be detected to obtain the audio spectrogram of the audio to be detected.

[0049] It should be noted that when receiving an abnormal sound detection instruction, the abnormal sound detection device will obtain the audio to be detected of the corresponding acoustic product. Since the audio to be detected is a wav file (Waveform Audio File Format), the short-time Fourier transform needs to be performed on the audio to be detected, and the Mel filter bank is applied, and then logarithmic compression and Mel spectrogram generation are performed to obtain the audio spectrogram of the audio to be detected. The width of the audio spectrogram represents the time dimension (duration), the height represents the frequency range, and the length of the time axis is usually much larger than the number of frequency ranges.

[0050] Step S12, perform cutting processing on multiple windows corresponding to the audio spectrogram through a patch embedding neural network to obtain multiple patch tokens corresponding to the audio spectrogram.

[0051] It should be noted that in order to effectively capture the relationship between different frequencies at the same time, windowing and patching processing are performed on the spectrogram. First, the audio spectrogram is divided into multiple windows w1, w2,..., wn, and then the patch embedding neural network is used to further divide each window into patches.

[0052] It can be understood that the patch embedding neural network is used to cut multiple windows corresponding to the audio spectrogram into different patch tokens. As Figure 4 shown, taking the patch embedding neural network with a kernel size of P x P as an example, the width of the audio spectrogram represents the time series T, the height represents the frequency F, w1, w2,..., wn are multiple windows corresponding to the audio spectrogram, and the patch convolution neural network is used to divide the multiple windows corresponding to the audio spectrogram into patches, thereby obtaining a series of patch tokens.

[0053] Step S13, input the multiple patch tokens corresponding to the audio spectrogram into the feature extraction unit in the abnormal sound detection model, and perform feature extraction on the multiple patch tokens through multiple siamese module groups in the feature extraction unit to obtain the associated multi-level features corresponding to the audio spectrogram.

[0054] It should be noted that multiple patch tokens of the audio spectrogram are input into the feature extraction unit in the outlier detection model in sequence. Specifically, they are input into the Siamese module group in the feature extraction unit according to the priority of the arrangement time, frequency domain, and window of each block token. Multiple Siamese module groups gradually extract features from multiple block tokens, and different levels of features are generated through a merging operation after each Siamese module group. The features output by the last layer of the last Siamese module group are processed by a convolutional layer with a preset size, and the obtained features after processing are the associated multi-level features corresponding to the audio spectrogram.

[0055] Step S14, aggregate semantic information of the audio spectrogram according to the associated multi-level features and the semantic feature extraction module in the feature extraction unit to obtain the target audio features corresponding to the audio spectrogram.

[0056] It should be noted that the associated multi-level features output by the last Siamese module group are input into the semantic feature extraction module (i.e., Token-Semantic CNN). The semantic feature extraction module integrates all frequency interval information along the frequency axis to generate time dimension features, and the output result of the semantic feature extraction module is the target audio features corresponding to the audio spectrogram.

[0057] It can be understood that in this embodiment, the Token-Semantic CNN layer is used, with enhanced semantic mapping, supporting multiple tasks simultaneously. Cross-band semantic integration is achieved through convolution in the frequency band dimension, and global classification results are generated using average pooling subsequently, while retaining the time domain localization ability.

[0058] In a feasible implementation manner, in step S13: extracting features of multiple block tokens through multiple Siamese module groups in the feature extraction unit to obtain the associated multi-level features corresponding to the audio spectrogram may include steps B11 to B14: Step B11, input multiple block tokens into multiple Siamese module groups in the feature extraction unit, and extract features of multiple block tokens through the first Siamese module group in the multiple Siamese module groups to obtain first-level features corresponding to each block token.

[0059] It should be noted that the first Siamese module group refers to the Siamese module group located at the first position and used to receive the patch tokens output by the Patch-Embed CNN. Features of multiple block tokens are extracted through the Siamese modules corresponding to multiple levels in the Siamese module group. The features of each block token output by the last-level Siamese module in the first Siamese module group are the first-level features corresponding to each block token.

[0060] In this embodiment, Patch-Embed CNN is used to implement local feature extraction. By dividing the image into local blocks of a fixed size, the model can more focus on learning local detailed features, reduce the computational complexity, and improve the parallel processing ability.

[0061] Step B12: Merge the multiple block tags according to the first-level features corresponding to each block tag to obtain multiple merged blocks and the second-level features corresponding to each merged block.

[0062] It should be noted that after each twin module group ends, a merging operation is performed on the multiple block tags based on the first-level features corresponding to each block tag to reduce the resolution and expand the channels, so as to obtain multiple merged blocks and the second-level features corresponding to each merged block. For example, if the window corresponding to the merging operation is 2×2 and the output of the first twin module group is (T / P × F / P, D), the output after merging is (T / 2P × F / 2P, 2D).

[0063] Step B13: Extract features from the audio spectrogram according to the second-level features corresponding to each merged block and the second twin module group in the multiple twin module groups to obtain the third-level features corresponding to the audio spectrogram.

[0064] It should be noted that the second twin module group refers to all the other twin module groups except the first twin module group in the multiple twin module groups. Input the second-level features corresponding to each merged block into the twin module groups in the front in the second twin module group, and then gradually perform feature extraction through the multiple twin module groups. The result output by the last layer of the last twin module group in the second twin module group is the third-level features corresponding to the audio spectrogram.

[0065] Step B14: Perform feature processing on the third-level features corresponding to the audio spectrogram according to a convolutional layer with a preset size to obtain the associated multi-level features corresponding to the audio spectrogram.

[0066] It should be noted that in this embodiment, the convolutional layer with a preset size is set as a convolutional layer with a convolutional kernel size of (3, F / 8P) and a padding size of (1, 0), and it can also be adjusted according to the actual situation. This embodiment does not limit this.

[0067] It can be understood that by introducing a convolutional layer with a preset size, the third-level features are further processed to generate a new feature map. Then, a reshape operation is performed on the feature map generated by the convolutional layer to obtain the associated multi-level features that can meet the operation requirements of the subsequent semantic feature extraction module.

[0068] It can be understood that after each twin module group, a global attention anchor is used to guide the merging position, and adjacent patches with semantic relevance are preferentially merged. After being processed by 4 network groups, the size of the patch tokens is reduced by 8 times from (T / P × F / P, D) to (T / 8P × F / 8P, 8D). In this embodiment, the semantic association region is identified through attention weights, and patches with high correlation are preferentially merged, enabling the features to gradually transition from local details to overall semantics, and avoiding semantic breaks caused by mechanical merging. It can guide cross-region semantic aggregation, enabling logically merging of patches that are not adjacent in physical position but are semantically related, breaking through the limitation of the local receptive field of traditional CNNs.

[0069] In a specific implementation, for example, there are four twin module groups 1 to 4. A merging operation is performed based on the output of group 1, and then the result of the merging operation is used as the input of group 2. A merging operation is performed based on the output of group 2, and then the result of the merging operation is used as the input of group 3. A merging operation is performed based on the output of group 3, and the result of the merging operation is used as the input of group 4. The output of group 4 is used as the input of a convolutional layer with a preset size. After further processing the output of group 4 through the convolutional layer, a reshape operation is performed on the feature map generated by the convolutional layer, and the result of the reshape operation (i.e., the associated multi-level features) is used as the input of the Token-Semantic CNN layer.

[0070] In a feasible implementation manner, in step B11: Extracting features from multiple block tokens through the first twin module group in multiple twin module groups to obtain first-level features corresponding to each block token may include steps C11 to C13: Step C11: Inputting multiple block tokens into the first twin module group in multiple twin module groups, and performing local feature extraction on the multiple block tokens through the local attention mechanism in the first twin module group to obtain local feature representations of each block token.

[0071] Step C12: Performing global feature extraction on the multiple block tokens through the global attention mechanism in the first twin module group according to the local feature representations of each block token to obtain global feature representations of each block token.

[0072] Step C13: Determining the first-level features corresponding to each block token according to the global feature representations of each block token and the hierarchical structure of the first twin module group.

[0073] It should be noted that since each twin module group includes multiple levels, there is a twin module in each level, and each twin module includes a local attention mechanism and a global attention mechanism. The patch tokens are input into the first twin module group in sequence. The twin modules at each level in the first twin module group first use the local attention mechanism LSA to extract features and capture local acoustic features, so as to obtain the local feature representations of each block token.

[0074] It can be understood that the twin modules at each level use the global attention mechanism GSA to perform global feature extraction on each block token by combining the local feature representations of each block token generated by the local attention mechanism at this level, and obtain the global feature representations of each block token.

[0075] In a specific implementation, after the first level in the first twin module group outputs the global feature representations of each block token, the global feature representations of each block token will be input into the second level in the first twin module group. The second level uses the local attention mechanism and the global attention mechanism to combine the global feature representations of each block token output by the first level for feature extraction, and inputs the result output by the second level into the third level, repeating the above steps until all levels in the first twin module group have completed feature extraction. In this embodiment, the first-level feature refers to the result generated after the twin module in the last level in the first twin module uses the global attention mechanism GSA for feature extraction.

[0076] It should be noted that although the levels of each twin module group are different, each level performs the above steps, that is, the output of the previous level - the LSA of the current level - the GSA of the current level - the LSA of the next level - until the end. In this embodiment, there are 4 twin module groups, corresponding to 3, 3, 9, and 3 levels respectively.

[0077] It can be understood that in this embodiment, a hybrid attention mechanism of LSA-GSA is adopted to reduce the computational amount. For audio patch tokens arranged in the order of time - frequency - window, each attention module will calculate the correlation relationship within a specific continuous frequency band and time range. As the network deepens, the global attention anchor points will merge adjacent windows, enabling the attention calculation to be performed in a larger spatial domain. The alternating use of LSA - GSA is as Figure 5 shown in the figure. The figure shows the downsampling process from high resolution to low resolution. The global attention acts on the original feature map on the right. The size of the original feature map is H×W. The global attention captures long-range dependencies by calculating the correlation weights between all positions (H×W); the square with size m×n in the middle on the right represents the process from global localization to local refinement; the local attention divides the global region into m×n sub-regions, and the size of each sub-region is H / m × W / n, as Figure 5For the bottommost feature map in the right part of the middle, local attention only focuses on short-range dependencies within a specific sub-region, significantly reducing the computational complexity. The local attention mechanism LSA enhances the local pattern sensitivity and focuses on capturing local acoustic features. The global attention mechanism GSA addresses the requirement for temporal continuity of audio signals. By alternately using LSA and GSA, the problem of semantic fragmentation caused by local windows is avoided.

[0078] This embodiment provides a method for detecting abnormal sounds. In this embodiment, the audio spectrogram of the audio to be detected is obtained by performing signal conversion on the audio to be detected; multiple windows corresponding to the audio spectrogram are cut through a block embedding neural network to obtain multiple block tokens corresponding to the audio spectrogram; the multiple block tokens corresponding to the audio spectrogram are input into a feature extraction unit in an abnormal sound detection model, and multiple block tokens are subjected to feature extraction through multiple twin module groups in the feature extraction unit to obtain associated multi-level features corresponding to the audio spectrogram; semantic information aggregation is performed on the audio spectrogram according to the associated multi-level features and a semantic feature extraction module in the feature extraction unit to obtain target audio features corresponding to the audio spectrogram. Through the above method, the global correlation across time steps in the acoustic signal can be dynamically captured, multi-scale abnormal pattern features can be extracted in parallel using multi-head attention, and efficient feature decoupling in a complex noise environment is achieved by combining a lightweight encoder design.

[0079] Based on the first embodiment and / or the second embodiment of the present application, in the third embodiment of the present application, the same or similar content as in the above-mentioned first embodiment and second embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 6 ., before step S10 of the abnormal sound detection method, steps S01 to S04 are further included: Step S01, sample generation is performed according to multiple sample audios and the abnormal sound classification labels of each sample audio to obtain multiple model training audios and the abnormal sound classification labels of each model training audio.

[0080] It should be noted that the sample audio is the original audio file used for training, including various types of abnormal sounds and normal audio; the abnormal sound classification label is the category identifier of each sample audio, which is used to indicate whether there is an abnormal sound in the audio and the specific abnormal sound type when there is an abnormal sound.

[0081] It can be understood that in order to ensure the balance and diversity of samples during model training, data augmentation techniques (such as Mix-up, spectral masking, etc.) are used to process the sample audio to generate more training audios, and each generated training audio has its corresponding abnormal sound classification label. In this embodiment, the model training audio is composed of a large number of sample audios and generated training audios.

[0082] Step S02: Divide multiple model training audios to obtain a model training set and a model validation set.

[0083] It should be noted that, in a randomly selected manner, among multiple model training audios, a preset proportion (e.g., 80%) of the model training audios are selected as the model training set, and the remaining model training audios are used as the model validation set.

[0084] It can be understood that, in this embodiment, all model training audios are converted into a single-channel format with a sampling rate of 32 kHz. The short-time Fourier transform (STFT) and Mel spectrogram are calculated using a window length of 1024, a hop size of 320, and 64 Mel bands. The shape of the Mel spectrogram is ultimately (1024, 64). For each 1000-frame (10 seconds) sample, 24 zero frames are padded (total number of frames T = 1024), and the patch size is set to 4×4, with a patch window length of 256 frames.

[0085] Step S03: Train an initial detection model according to the model training set, the abnormal sound classification labels of each audio in the model training set, a learning rate warm-up strategy, and a target optimizer to obtain a first detection model. The initial detection model is composed of multiple initial twin module groups, an initial semantic feature extraction module, and an initial abnormal sound classification unit.

[0086] It should be noted that the model training audios in the model training set are windowed, and after windowing, they are cut into Patch-tokens by Patch-Embed CNN for iterative training input into the network. The learning rate warm-up strategy in this embodiment means that the learning rates for the first three rounds are set to 0.05, 0.1, and 0.2 in sequence, and then the learning rate is halved every ten rounds until it returns to 0.05. The specific value of the learning rate can be adjusted according to the actual situation, and this embodiment does not limit it.

[0087] It can be understood that the target optimizer is set as the AdamW optimizer in this embodiment, where β1 = 0.9, β2 = 0.999, ε = 1e-8, and the weight decay rate is 0.05. It can also be set as other optimizers, and the parameters can be adjusted according to the actual situation. This embodiment does not limit it. The multiple initial twin module groups are multiple untrained twin module groups; the initial semantic feature extraction module is an untrained Token-Semantic CNN layer; the initial abnormal sound classification unit is an untrained abnormal sound classification unit, and its core components include a fully connected layer, global average pooling, and an activation function.

[0088] In a specific implementation, the initial detection model is trained using the audio in the model training set and its corresponding abnormal sound classification labels, in combination with a weighted average strategy, a learning rate warm-up strategy, an objective optimizer, and pre-trained weights, so as to obtain a preliminarily trained first detection model. During the training process, the binary cross-entropy loss is calculated based on the predicted vectors of each audio and the abnormal sound classification labels of each audio, thereby guiding the model to optimize parameters and improve classification accuracy. In this embodiment, the number of training iterations can be set to 200, or it can be adjusted according to requirements, and this embodiment does not limit this.

[0089] It should be noted that during the training process, in order to solve the problem of class imbalance, a balanced sampler is used to adjust the probability of different label audios being extracted, so that the samples of each category can participate in the training process in a relatively balanced manner. Specifically, a weighted sampling method can be used to assign higher weights to the samples of the minority classes to ensure that their probability of being selected in each iteration is higher.

[0090] Step S04, perform performance evaluation on the first detection model according to the model validation set, and determine the abnormal sound detection model according to the performance evaluation result.

[0091] It should be noted that the first detection model is tested using the audios in the model validation set and their abnormal sound classification labels, and various performance indicators are calculated, including but not limited to accuracy, recall, and F1 score, etc. If the performance indicators of the first detection model all meet the expected standards, it can be used as the final abnormal sound detection model; otherwise, the hyperparameters or structure of the first detection model are adjusted and retrained to obtain the abnormal sound detection model.

[0092] It can be understood that in this embodiment, an abnormal sound detection model based on transformers and time-frequency information is used to improve the calculation efficiency, reduce the training parameters, and has better compatibility than traditional CNNs.

[0093] In a feasible implementation manner, step S01 may further include steps D11 to D14: Step D11, randomly extract multiple sample audios to obtain multiple pairs of sample audios.

[0094] Step D12, perform linear interpolation on multiple pairs of sample audios, and generate the first generated audio of each pair of sample audios and the abnormal sound classification labels of each first generated audio according to the linear interpolation result.

[0095] It should be noted that, by using the Mix-up augmentation technique, multiple sample pairs are randomly selected from multiple sample audios for linear interpolation, so as to generate new audios corresponding to each sample audio pair, and use them as the first generated audios. Combining the abnormal sound classification labels of each sample audio, the abnormal sound classification label of the first generated audio is determined. In this embodiment, the parameter is set to 0.5, making the new samples more inclined to evenly mix the two sample audios.

[0096] Step D13: Randomly mask each first generated audio and each sample audio, and generate multiple second generated audios and the abnormal sound classification labels of each second generated audio according to the masking results.

[0097] It should be noted that, by using the spectral masking technique, for multiple first generated audios and multiple sample audios, multiple audios are randomly selected, and the corresponding spectrograms are spectrally masked. The new samples generated after the masking are the second generated audios. Combining the abnormal sound classification labels of each sample audio, the abnormal sound classification label of the second generated audio is determined.

[0098] It can be understood that the spectral masking technique includes time masking and frequency masking. The specific operation of time masking is: randomly select a continuous time period on the time axis of the spectrogram, and set all frequency components within this time period to zero, thus simulating the information loss situation in some time periods of the audio signal.

[0099] The specific operation of frequency masking is to randomly select a frequency band with a certain width on the frequency axis of the spectrogram, and cover the information of this frequency band over the entire time period, thus simulating the loss situation of some frequency components in the audio signal. In this embodiment, the time masking length is set to 128 frames, and the frequency masking width is set to 16 frequency bands, which can also be adjusted according to the actual situation, and this embodiment does not limit this.

[0100] Step D14: Obtain multiple model training audios and the abnormal sound classification labels of each model training audio according to each sample audio and the abnormal sound classification label of each sample audio, each first generated audio and the abnormal sound classification label of each first generated audio, and each second generated audio and the abnormal sound classification label of each second generated audio.

[0101] It should be noted that all the new generated audios and the original sample audios are summarized to obtain multiple model training audios; and based on the abnormal sound classification labels of each new generated audio and the abnormal sound classification labels of the sample audios, the abnormal sound classification label of each model training audio is obtained.

[0102] This embodiment provides a heterophony detection method. In this embodiment, multiple model training audio and the heterophony classification labels of each model training audio are obtained by generating samples according to multiple sample audios and the heterophony classification labels of each sample audio; the multiple model training audio are divided to obtain a model training set and a model verification set; the initial detection model is trained according to the model training set, the heterophony classification labels of the audios in the model training set, the learning rate warm-up strategy, and the target optimizer to obtain a first detection model, where the initial detection model is composed of multiple initial twin module groups, an initial semantic feature extraction module, and an initial heterophony classification unit; the performance of the first detection model is evaluated according to the model verification set, and the heterophony detection model is determined according to the performance evaluation result. In the above manner, the high robustness of the model is ensured, the generalization ability of the model is improved, the training efficiency and stability are improved, and the convergence process of the model is accelerated.

[0103] Exemplarily, to help understand the implementation process of the heterophony detection method obtained by combining this embodiment with the above Embodiment 1 and Embodiment 2, please refer to Figure 7 , Figure 7 A brief flow schematic diagram of a heterophony detection method is provided. Specifically: S1, a sweep signal containing multiple test frequencies is sent to the target speaker, and the audio signal emitted by the target speaker is collected. S2, the Fourier transform and Mel spectrogram of the audio signal are calculated to obtain the spectral signal of the audio varying with frequency. 80% of the samples in the dataset are randomly selected as the training set, and the remaining are used as the verification set. After windowing, it is cut into Patch-tokens by Patch-Embed CNN and prepared to be input into the network. S3, the detection model gradually reduces the resolution and increases the number of channels, enabling it to efficiently process features of different scales; Twins-SVT adopts the Spatially Separable Self-Attention (SVT) mechanism, which divides the calculation into two steps: Local Self-Attention (LSA): Similar to the local window attention of CNN convolution, reducing the computational complexity; Global Self-Attention (GSA): The cross-window attention mechanism, supplementing global information. S4, set the training parameters of the model: Select Twins-SVT as the backbone network, use the pre-trained weights, set the number of training epochs to 200, use the AdamW optimizer (β1 = 0.9, β2 = 0.999, ε = 1e-8, weight decay rate 0.05), and adopt the learning rate warm-up strategy: The learning rates in the first three rounds are set to 0.05, 0.1, and 0.2 in sequence, and then the learning rate is halved every ten rounds until it returns to 0.05.

[0104] S5. Training the model: Input the Patch-tokens corresponding to the spectrograms obtained in S2 into the improved network in S3, and train the network according to the parameters in S4 to obtain an abnormal sound detection model. S6. Model inference: During inference, first convert the wav file of the detected audio into a spectrogram, and input the patch-tokens corresponding to the spectrogram image to be classified into the classification model trained in S5 for inference to obtain the classification result.

[0105] Through the transformer abnormal sound detection algorithm with a multi-level structure proposed in this embodiment, it focuses on solving the problem of insufficient detection accuracy in abnormal sound detection of acoustic products due to the lack of correlation of long-time acoustic signals and the difficulty in capturing weak abnormal patterns. It can achieve the following beneficial effects including but not limited to: 1) Using an abnormal sound detection algorithm based on transformer and time-frequency information to improve the calculation efficiency, reduce the training parameters, and have better compatibility than traditional CNN. (2) Using Patch-Embed CNN to achieve local feature extraction. By dividing the image into local blocks of a fixed size, the model can focus more on learning local detail features, reduce the calculation complexity, and improve the parallel processing ability. (3) Identify the semantic correlation region through the attention weight, and preferentially merge the patches with high correlation, so that the features gradually transition from local details to overall semantics, avoiding semantic breaks caused by mechanical merging. It can guide the semantic aggregation across regions, enabling logical merging of patches that are not adjacent physically but are semantically related, breaking through the limitation of the local receptive field of traditional CNN. (4) Use the mechanism of alternating local-global attention (LSA-GSA). LSA enhances the sensitivity to local patterns and focuses on capturing local acoustic features, while GSA solves the requirement for the temporal continuity of audio signals. By alternately using LSA and GSA, it avoids the semantic fragmentation problem caused by local windows. (5) Use a labeled semantic CNN layer after the final transformer module (i.e., the twin module group). The semantic mapping is enhanced, and it supports multi-tasks at the same time. Cross-band semantic integration is achieved through convolution in the frequency band dimension, and finally global classification results are generated by average pooling, retaining the time-domain localization ability.

[0106] It should be noted that the above examples are only for understanding this application and do not constitute a limitation to the abnormal sound detection method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0107] This application also provides an abnormal sound detection device. Please refer to Figure 8 and the abnormal sound detection device includes: An extraction module 10, configured to, in response to a heterophony detection instruction, extract features of an audio to be detected through a feature extraction unit in a heterophony detection model, so as to obtain target audio features of the audio to be detected. The feature extraction unit is composed of a plurality of siamese module groups and a semantic feature extraction module, and a hybrid attention mechanism exists in the siamese module groups.

[0108] A detection module 20, configured to input the target audio features into a heterophony classification unit in the heterophony detection model for heterophony detection, so as to obtain the audio category to which the audio to be detected belongs.

[0109] A processing module 30, configured to obtain an audio detection result of the audio to be detected according to the audio category to which the audio to be detected belongs.

[0110] Optionally, the extraction module 10 is further configured to: Perform signal conversion on the audio to be detected to obtain an audio spectrogram of the audio to be detected; perform cutting processing on a plurality of windows corresponding to the audio spectrogram through a block embedding neural network to obtain a plurality of block tokens corresponding to the audio spectrogram; input the plurality of block tokens corresponding to the audio spectrogram into the feature extraction unit in the heterophony detection model, and extract features of the plurality of block tokens through a plurality of siamese module groups in the feature extraction unit to obtain associated multi-level features corresponding to the audio spectrogram; perform semantic information aggregation on the audio spectrogram according to the associated multi-level features and the semantic feature extraction module in the feature extraction unit to obtain target audio features corresponding to the audio spectrogram.

[0111] Optionally, the extraction module 10 is further configured to: Input the plurality of block tokens into a plurality of siamese module groups in the feature extraction unit, and extract features of the plurality of block tokens through a first siamese module group in the plurality of siamese module groups to obtain first-level features corresponding to each block token; merge the plurality of block tokens according to the first-level features corresponding to each block token to obtain a plurality of merged blocks and second-level features corresponding to each merged block; extract features of the audio spectrogram according to the second-level features corresponding to each merged block and a second siamese module group in the plurality of siamese module groups to obtain third-level features corresponding to the audio spectrogram; perform feature processing on the third-level features corresponding to the audio spectrogram through a convolutional layer with a preset size to obtain associated multi-level features corresponding to the audio spectrogram.

[0112] Optionally, the extraction module 10 is further configured to: Input multiple block tokens into the first twin module group among multiple twin module groups, perform local feature extraction on the multiple block tokens through the local attention mechanism in the first twin module group to obtain local feature representations of each block token; perform global feature extraction on the multiple block tokens according to the local feature representations of each block token through the global attention mechanism in the first twin module group to obtain global feature representations of each block token; determine the first-level features corresponding to each block token according to the global feature representations of each block token and the hierarchical structure of the first twin module group.

[0113] Optionally, the detection module 20 is further configured to: Input the target audio feature into the abnormal sound classification unit in the abnormal sound detection model, perform feature mapping on the target audio feature through the fully connected layer in the abnormal sound classification unit, and obtain the category probability map of the audio to be detected according to the mapping result; perform global average pooling operation on the category probability map through the abnormal sound classification unit to obtain the category probability vector of the audio to be detected; perform category determination according to the category probability vector of the audio to be detected to determine the audio category to which the audio to be detected belongs.

[0114] Optionally, the extraction module 10 is further configured to: Generate multiple model training audios and the abnormal sound classification labels of each model training audio according to multiple sample audios and the abnormal sound classification labels of each sample audio; divide the multiple model training audios to obtain a model training set and a model validation set; train an initial detection model according to the model training set, the abnormal sound classification labels of each audio in the model training set, the learning rate warm-up strategy, and the target optimizer, where the initial detection model is composed of multiple initial twin module groups, an initial semantic feature extraction module, and an initial abnormal sound classification unit; perform performance evaluation on the first detection model according to the model validation set, and determine the abnormal sound detection model according to the performance evaluation result.

[0115] Optionally, the extraction module 10 is further configured to: Randomly extract multiple sample audio pairs from multiple sample audios; perform linear interpolation on the multiple sample audio pairs, and generate the first generated audio of each sample audio pair and the abnormal sound classification label of each first generated audio according to the linear interpolation result; randomly mask each first generated audio and each sample audio, and generate multiple second generated audios and the abnormal sound classification labels of each second generated audio according to the masking result; obtain multiple model training audios and the abnormal sound classification labels of each model training audio according to each sample audio and the abnormal sound classification label of each sample audio, each first generated audio and the abnormal sound classification label of each first generated audio, and each second generated audio and the abnormal sound classification label of each second generated audio.

[0116] The abnormal sound detection device provided by this application adopts the abnormal sound detection method in the above embodiment, which can solve the technical problem that the traditional acoustic detection method has low detection accuracy and generalization ability because it is difficult to accurately separate abnormal acoustic features from background interference. Compared with the prior art, the beneficial effects of the abnormal sound detection device provided by this application are the same as those of the abnormal sound detection method provided by the above embodiment, and other technical features in the abnormal sound detection device are the same as those disclosed in the method of the above embodiment, which will not be elaborated here.

[0117] This application provides an abnormal sound detection device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the abnormal sound detection method in the first embodiment above.

[0118] Refer to the following Figure 9 , which shows a schematic structural diagram of an abnormal sound detection device suitable for implementing the embodiments of this application. The abnormal sound detection device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The abnormal sound detection device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of this application.

[0119] As Figure 9As shown in the figure, the abnormal sound detection device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the abnormal sound detection device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the abnormal sound detection device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an abnormal sound detection device with various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems can be alternatively implemented or had.

[0120] Specifically, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.

[0121] The abnormal sound detection device provided by the present application adopts the abnormal sound detection method in the above embodiments, and can solve the technical problem that the traditional acoustic detection method has low detection accuracy and generalization ability because it is difficult to accurately separate abnormal acoustic features from background interference. Compared with the prior art, the beneficial effects of the abnormal sound detection device provided by the present application are the same as those of the abnormal sound detection method provided by the above embodiments, and other technical features in the abnormal sound detection device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.

[0122] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0123] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0124] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the abnormal sound detection method in the above embodiments.

[0125] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM: Random Access Memory), read-only memory (ROM: Read Only Memory), erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0126] The above computer-readable storage medium can be included in the abnormal sound detection device; it can also exist separately and not be assembled into the abnormal sound detection device.

[0127] The above computer-readable storage medium carries one or more programs, which, when executed by the abnormal sound detection device, cause the abnormal sound detection device to: in response to an abnormal sound detection instruction, extract features of the audio to be detected through a feature extraction unit in the abnormal sound detection model to obtain target audio features of the audio to be detected, where the feature extraction unit is composed of a plurality of siamese module groups and a semantic feature extraction module, and there is a hybrid attention mechanism in the siamese module groups; input the target audio features into an abnormal sound classification unit in the abnormal sound detection model to perform abnormal sound detection to obtain the audio category to which the audio to be detected belongs; and obtain an audio detection result of the audio to be detected according to the audio category to which the audio to be detected belongs.

[0128] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0130] The modules involved in the embodiments of this application can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.

[0131] The readable storage medium provided in this application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned abnormal sound detection method, which can solve the technical problems that the traditional acoustic detection method has low detection accuracy and generalization ability due to the difficulty in accurately separating abnormal acoustic features from background interference. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the abnormal sound detection method provided in the above embodiments, and will not be elaborated here.

[0132] This application also provides a computer program product, including a computer program, and the steps of the above-mentioned abnormal sound detection method are implemented when the computer program is executed by a processor.

[0133] The computer program product provided in this application can solve the technical problems that the traditional acoustic detection method has low detection accuracy and generalization ability due to the difficulty in accurately separating abnormal acoustic features from background interference. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the abnormal sound detection method provided in the above embodiments, and will not be elaborated here.

[0134] The above are only some embodiments of this application, and do not limit the patent scope of this application. Any equivalent structural transformation made by using the content of the specification and drawings of this application under the technical concept of this application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of this application.

Claims

1. A method for detecting abnormal sound, characterized in that: The method comprises: In response to the foreign sound detection instruction, a feature extraction unit in the foreign sound detection model extracts features of the audio to be detected to obtain target audio features of the audio to be detected, wherein the feature extraction unit is composed of a plurality of twin module groups and a semantic feature extraction module, and a hybrid attention mechanism exists in the twin module group; Inputting the target audio feature into the foreign sound classification unit in the foreign sound detection model to perform foreign sound detection, and obtaining the audio category to which the audio to be detected belongs; An audio detection result of the audio to be detected is obtained according to the audio category to which the audio to be detected belongs.

2. The method according to claim 1, characterized in that The step of extracting features of the audio to be detected by the feature extraction unit in the abnormal sound detection model to obtain the target audio features of the audio to be detected includes: Performing signal conversion on the audio to be detected to obtain an audio spectrogram of the audio to be detected; Cutting multiple windows corresponding to the audio spectrogram through a block embedding neural network to obtain multiple block labels corresponding to the audio spectrogram; Inputting a plurality of block markers corresponding to the audio spectrogram into a feature extraction unit in a different sound detection model, extracting features from the plurality of block markers through a plurality of twin module groups in the feature extraction unit, and obtaining associated multi-level features corresponding to the audio spectrogram; The audio spectrogram is aggregated with semantic information according to the associated multi-level features and the semantic feature extraction module in the feature extraction unit to obtain the target audio features corresponding to the audio spectrogram.

3. The method according to claim 2, characterized in that The step of extracting features from multiple block markers through multiple twin module groups in the feature extraction unit to obtain associated multi-level features corresponding to the audio spectrogram includes: Inputting a plurality of block markers into a plurality of twin module groups in the feature extraction unit, performing feature extraction on the plurality of block markers through a first twin module group in the plurality of twin module groups, and obtaining first-level features corresponding to each block marker; Merge multiple block markers according to the first-level features corresponding to each block marker to obtain multiple merged blocks and second-level features corresponding to each merged block; Extracting features of the audio spectrogram according to the second-level features corresponding to each merging block and the second twin module group in the plurality of twin module groups, to obtain third-level features corresponding to the audio spectrogram; The third-level features corresponding to the audio spectrogram are processed according to a convolutional layer of a preset size to obtain associated multi-level features corresponding to the audio spectrogram.

4. The method according to claim 3, characterized in that The step of extracting features from a plurality of block markers by using a first twin module group among the plurality of twin module groups to obtain first-level features corresponding to each block marker comprises: Inputting the plurality of block markers into a first twin module group among the plurality of twin module groups, extracting local features of the plurality of block markers through a local attention mechanism in the first twin module group, and obtaining a local feature representation of each block marker; According to the local feature representation of each block mark, a global feature extraction is performed on multiple block marks through the global attention mechanism in the first twin module group to obtain the global feature representation of each block mark; The first-level features corresponding to each block mark are determined according to the global feature representation of each block mark and the hierarchical structure of the first twin module group.

5. The method according to claim 1, characterized in that The step of inputting the target audio feature into the foreign sound classification unit in the foreign sound detection model to perform foreign sound detection and obtaining the audio category to which the audio to be detected belongs comprises: Inputting the target audio feature into the foreign sound classification unit in the foreign sound detection model, performing feature mapping on the target audio feature through the fully connected layer in the foreign sound classification unit, and obtaining a category probability map of the audio to be detected according to the mapping result; Performing a global average pooling operation on the category probability map by the foreign sound classification unit to obtain a category probability vector of the audio to be detected; A category determination is performed according to the category probability vector of the audio to be detected to determine the category to which the audio to be detected belongs.

6. The method according to any one of claims 1 to 5, characterized in that Before the step of extracting features of the audio to be detected by the feature extraction unit in the abnormal sound detection model to obtain the target audio features of the audio to be detected, the method further includes: Generate samples based on multiple sample audios and the foreign sound classification labels of each sample audio, and obtain multiple model training audios and the foreign sound classification labels of each model training audio; Divide multiple model training audios into a model training set and a model verification set; The initial detection model is trained according to the model training set, the foreign sound classification label of each audio in the model training set, the learning rate warm-up strategy and the target optimizer to obtain a first detection model, wherein the initial detection model is composed of a plurality of initial twin module groups, an initial semantic feature extraction module and an initial foreign sound classification unit; A performance evaluation is performed on the first detection model according to the model verification set, and a foreign sound detection model is determined according to the performance evaluation result.

7. The method according to claim 6, characterized in that The step of generating samples according to a plurality of sample audios and the foreign sound classification labels of each sample audio to obtain a plurality of model training audios and the foreign sound classification labels of each model training audio comprises: Randomly extracting multiple sample audios to obtain multiple sample audio pairs; Performing linear interpolation on a plurality of sample audio pairs, and generating a first generated audio of each sample audio pair and a foreign sound classification label of each first generated audio according to the linear interpolation result; Randomly mask each first generated audio and each sample audio, and generate a plurality of second generated audios and a foreign sound classification label of each second generated audio according to the masking result; Based on each sample audio and the foreign sound classification label of each sample audio, each first generated audio and the foreign sound classification label of each first generated audio, each second generated audio and the foreign sound classification label of each second generated audio, multiple model training audios and the foreign sound classification label of each model training audio are obtained.

8. A device for detecting abnormal sound, characterized in that: The abnormal sound detection device comprises: An extraction module, configured to extract features of the audio to be detected by a feature extraction unit in the audio detection model in response to a foreign sound detection instruction, and obtain target audio features of the audio to be detected, wherein the feature extraction unit is composed of a plurality of twin module groups and a semantic feature extraction module, and a hybrid attention mechanism exists in the twin module group; A detection module, used for inputting the target audio feature into the foreign sound classification unit in the foreign sound detection model to perform foreign sound detection, and obtaining the audio category to which the audio to be detected belongs; The processing module is used to obtain the audio detection result of the audio to be detected according to the audio category of the audio to be detected.

9. An abnormal sound detection device, characterized in that: The device comprises: a memory, a processor, and an abnormal sound detection program stored in the memory and executable on the processor, wherein the abnormal sound detection program is configured to implement the steps of the abnormal sound detection method according to any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium stores an abnormal sound detection program, and when the abnormal sound detection program is executed by the processor, the steps of the abnormal sound detection method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Abnormal sound detection method and device, electronic equipment and medium

    CN114004996A

  • Techniques for classification with neural networks

    CN114600119A

  • Video classification method, method, device and equipment for constructing classification model

    CN114882399A

  • Wind turbine generator operation and maintenance decision-making system based on artificial intelligence

    CN118690313A

  • Audio data dynamic transmission processing method and audio playing device

    CN119132315A