Abnormal Sound Detection Method, Apparatus, Device, and Storage Medium
Through the Transformer-based antonym detection framework, combined with time-frequency domain depth characterization and self-attention mechanism, the problem of insufficient accuracy and generalization capabilities of traditional acoustic detection methods in complex acoustic scenes is solved, and efficient and robust antonym detection is achieved.
Patent Information
- Application Number
- CN202510560161.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-30
AI Technical Summary
Traditional acoustic detection methods are difficult to achieve high-precision and generalizational heterophonic judgments in complex acoustic scenarios and diversified product lines, making it difficult to ensure detection consistency.
Using a Transformer-based anomaly detection framework, the time-frequency domain depth representation and self-attention mechanism are integrated, and the global correlation between the acoustic signals is dynamically captured, multi-head attention is used to extract multi-scale anomaly mode features in parallel, and combined with lightweight encoder design, it realizes efficient feature decoupling in complex noise environments.
It significantly improves the model's sensitivity to weak abnormal sounds, realizes highly robust abnormal sound detection, quickly adapts to the differences in acoustic characteristics of different models of acoustic products, and reduces the cost of cross-product line migration and misjudgment rate.
Smart Images

Figure CN120089162B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of acoustic technologies, and particularly to a method, apparatus, device, and storage medium for detecting abnormal sounds. Background Art
[0002] The acoustic quality of acoustic products is one of the core factors determining user experience and market competitiveness. Acoustic products such as headphones, speakers, and microphones need to undergo a strict abnormal sound detection process before leaving the factory to ensure that there are no defects such as noise, broken sound, abnormal background noise, or structural resonance during their operation. Traditional detection methods highly rely on the subjective judgment of human listeners. By repeatedly playing standardized audio samples and relying on the human ear to identify abnormalities. However, with the exponential growth of the shipment volume of acoustic products and the accelerating speed of product iteration, the disadvantages of manual detection have become increasingly prominent: low detection efficiency, rising labor costs, and the human ear is vulnerable to fatigue, individual hearing differences, and environmental noise interference, resulting in difficulty in ensuring detection consistency. Although some enterprises have introduced automated detection equipment, the existing technologies still have difficulty in achieving high-precision and strong generalization of abnormal sound determination in complex acoustic scenarios and diverse product lines.
[0003] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide a method, apparatus, device, and storage medium for detecting abnormal sounds, aiming to solve the technical problem that traditional acoustic detection methods have low detection accuracy and generalization ability because it is difficult to accurately separate abnormal acoustic features from background interference.
[0005] To achieve the above purpose, this application proposes a method for detecting abnormal sounds, and the method includes:
[0006] In response to an abnormal sound detection instruction, the feature extraction unit in the abnormal sound detection model extracts features from the audio to be detected to obtain the target audio features of the audio to be detected. The feature extraction unit is composed of multiple twin module groups and a semantic feature extraction module, and there is a hybrid attention mechanism in the twin module groups;
[0007] Input the target audio features into the abnormal sound classification unit in the abnormal sound detection model for abnormal sound detection to obtain the category to which the audio of the audio to be detected belongs;
[0008] Obtain the audio detection result of the audio to be detected according to the category to which the audio of the audio to be detected belongs.
[0009] In one embodiment, the step of extracting features from the audio to be detected by the feature extraction unit in the abnormal sound detection model to obtain the target audio features of the audio to be detected includes:
[0010] Perform signal conversion on the audio to be detected to obtain the audio spectrogram of the audio to be detected;
[0011] Perform cutting processing on multiple windows corresponding to the audio spectrogram through a block embedding neural network to obtain multiple block tags corresponding to the audio spectrogram;
[0012] Input the multiple block tags corresponding to the audio spectrogram into the feature extraction unit in the abnormal sound detection model, and perform feature extraction on the multiple block tags through multiple twin module groups in the feature extraction unit to obtain the associated multi-level features corresponding to the audio spectrogram;
[0013] Perform semantic information aggregation on the audio spectrogram according to the associated multi-level features and the semantic feature extraction module in the feature extraction unit to obtain the target audio features corresponding to the audio spectrogram.
[0014] In one embodiment, the step of performing feature extraction on multiple block tags through multiple twin module groups in the feature extraction unit to obtain the associated multi-level features corresponding to the audio spectrogram includes:
[0015] Input multiple block tags into multiple twin module groups in the feature extraction unit, and perform feature extraction on the multiple block tags through the first twin module group in the multiple twin module groups to obtain the first-level features corresponding to each block tag;
[0016] Merge multiple block tags according to the first-level features corresponding to each block tag to obtain multiple merged blocks and the second-level features corresponding to each merged block;
[0017] Perform feature extraction on the audio spectrogram according to the second-level features corresponding to each merged block and the second twin module group in the multiple twin module groups to obtain the third-level features corresponding to the audio spectrogram;
[0018] Perform feature processing on the third-level features corresponding to the audio spectrogram according to a convolutional layer with a preset size to obtain the associated multi-level features corresponding to the audio spectrogram.
[0019] In one embodiment, the step of performing feature extraction on multiple block tags through the first twin module group in the multiple twin module groups to obtain the first-level features corresponding to each block tag includes:
[0020] Input multiple block tags into the first twin module group in the multiple twin module groups, and perform local feature extraction on the multiple block tags through the local attention mechanism in the first twin module group to obtain the local feature representations of each block tag;
[0021] Global feature extraction is performed on multiple block tokens through the global attention mechanism in the first twin module group according to the local feature representations of each block token, and global feature representations of each block token are obtained;
[0022] The first-level features corresponding to each block token are determined according to the global feature representations of each block token and the hierarchical structure of the first twin module group.
[0023] In one embodiment, the step of inputting the target audio feature into the abnormal sound classification unit in the abnormal sound detection model to perform abnormal sound detection and obtaining the audio category to which the audio to be detected belongs includes:
[0024] The target audio feature is input into the abnormal sound classification unit in the abnormal sound detection model, feature mapping is performed on the target audio feature through the fully connected layer in the abnormal sound classification unit, and a category probability map of the audio to be detected is obtained according to the mapping result;
[0025] A global average pooling operation is performed on the category probability map through the abnormal sound classification unit to obtain a category probability vector of the audio to be detected;
[0026] Category determination is performed according to the category probability vector of the audio to be detected to determine the audio category to which the audio to be detected belongs.
[0027] In one embodiment, before the step of performing feature extraction on the audio to be detected through the feature extraction unit in the abnormal sound detection model to obtain the target audio feature of the audio to be detected, the following steps are further included:
[0028] Sample generation is performed according to multiple sample audios and the abnormal sound classification labels of each sample audio to obtain multiple model training audios and the abnormal sound classification labels of each model training audio;
[0029] The multiple model training audios are divided to obtain a model training set and a model validation set;
[0030] The initial detection model is trained according to the model training set, the abnormal sound classification labels of each audio in the model training set, the learning rate warm-up strategy, and the target optimizer to obtain a first detection model, and the initial detection model is composed of multiple initial twin module groups, an initial semantic feature extraction module, and an initial abnormal sound classification unit;
[0031] The performance of the first detection model is evaluated according to the model validation set, and an abnormal sound detection model is determined according to the performance evaluation result.
[0032] In one embodiment, the step of performing sample generation according to multiple sample audios and the abnormal sound classification labels of each sample audio to obtain multiple model training audios and the abnormal sound classification labels of each model training audio includes:
[0033] Randomly sample multiple sample audio to obtain multiple pairs of sample audio;
[0034] Perform linear interpolation on multiple pairs of sample audio, and generate the first generated audio for each pair of sample audio and the abnormal sound classification label for each first generated audio according to the linear interpolation result;
[0035] Randomly mask each first generated audio and each sample audio, and generate multiple second generated audio and the abnormal sound classification label for each second generated audio according to the masking result;
[0036] According to each sample audio and the abnormal sound classification label of each sample audio, each first generated audio and the abnormal sound classification label of each first generated audio, each second generated audio and the abnormal sound classification label of each second generated audio, obtain multiple model training audio and the abnormal sound classification label of each model training audio.
[0037] In addition, to achieve the above object, the present application also proposes an abnormal sound detection device, the abnormal sound detection device includes: an extraction module, configured to respond to an abnormal sound detection instruction, and extract features of a to-be-detected audio through a feature extraction unit in an abnormal sound detection model to obtain target audio features of the to-be-detected audio, the feature extraction unit is composed of a plurality of twin module groups and a semantic feature extraction module, and there is a hybrid attention mechanism in the twin module groups;
[0038] A detection module, configured to input the target audio features into an abnormal sound classification unit in the abnormal sound detection model for abnormal sound detection to obtain the audio category to which the to-be-detected audio belongs;
[0039] A processing module, configured to obtain an audio detection result of the to-be-detected audio according to the audio category to which the to-be-detected audio belongs.
[0040] In addition, to achieve the above object, the present application also proposes an abnormal sound detection device, the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program is configured to implement the steps of the abnormal sound detection method as described above.
[0041] In addition, to achieve the above object, the present application also proposes a storage medium, the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the abnormal sound detection method as described above.
[0042] In addition, to achieve the above object, the present application also provides a computer program product, the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps of the abnormal sound detection method as described above.
[0043] The present application provides a method for detecting abnormal sounds. In response to an abnormal sound detection instruction, the present application extracts features of the audio to be detected through a feature extraction unit in an abnormal sound detection model, and obtains target audio features of the audio to be detected. The feature extraction unit is composed of multiple siamese module groups and a semantic feature extraction module, and there is a hybrid attention mechanism in the siamese module groups; the target audio features are input into an abnormal sound classification unit in the abnormal sound detection model for abnormal sound detection, and the audio category to which the audio to be detected belongs is obtained; an audio detection result of the audio to be detected is obtained according to the audio category to which the audio to be detected belongs. By the above method, by fusing time-frequency domain depth representation and self-attention mechanism, the global correlation across time steps in the acoustic signal is dynamically captured, multi-scale abnormal pattern features are extracted in parallel by using multi-head attention, and efficient feature decoupling in a complex noise environment is realized by combining a lightweight encoder design, which significantly improves the sensitivity of the model to weak abnormal sounds and realizes high-robustness abnormal sound detection. At the same time, it can quickly adapt to the acoustic characteristic differences of different models of acoustic products through a parameter fine-tuning strategy, providing an innovative technical path for solving the problems of high cross-product line migration cost and large fluctuation of misjudgment rate of traditional models. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0046] Figure 1 It is a schematic flow chart provided for Embodiment 1 of the abnormal sound detection method of the present application;
[0047] Figure 2 It is a schematic network architecture diagram provided for Embodiment 1 of the present application;
[0048] Figure 3 It is a schematic flow chart provided for Embodiment 2 of the abnormal sound detection method of the present application;
[0049] Figure 4 It is a schematic diagram of windowed cutting provided for Embodiment 2 of the present application;
[0050] Figure 5 It is a schematic diagram of the flow of the attention mechanism provided for Embodiment 2 of the present application
[0051] Figure 6Schematic flowchart provided for the third embodiment of the abnormal sound detection method of this application;
[0052] Figure 7 Brief schematic flowchart of the abnormal sound detection method provided for the third embodiment of this application;
[0053] Figure 8 Schematic module structure diagram of the abnormal sound detection device for the embodiment of this application;
[0054] Figure 9 Schematic device structure diagram of the hardware operating environment involved in the abnormal sound detection method for the embodiment of this application.
[0055] The implementation, functional features and advantages of the purpose of this application will be further described with reference to the embodiments and the accompanying drawings. Specific implementation manners
[0056] It should be understood that the specific embodiments described herein are only used to explain the technical solution of this application and are not used to limit this application.
[0057] In order to better understand the technical solution of this application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.
[0058] The main solution of the embodiment of this application is: in response to an abnormal sound detection instruction, a feature extraction unit in an abnormal sound detection model is used to extract features from the audio to be detected, obtaining the target audio features of the audio to be detected, the feature extraction unit is composed of a plurality of twin module groups and a semantic feature extraction module, and there is a hybrid attention mechanism in the twin module groups; inputting the target audio features into an abnormal sound classification unit in the abnormal sound detection model for abnormal sound detection, obtaining the audio category to which the audio to be detected belongs; and obtaining the audio detection result of the audio to be detected according to the audio category to which the audio to be detected belongs.
[0059] Existing acoustic product abnormal sound detection technologies mainly focus on signal analysis and traditional machine learning. At the signal analysis level, common methods include frequency response curve comparison, total harmonic distortion (THD) test, signal-to-noise ratio (SNR) calculation, etc. Whether there are abnormalities is determined by setting thresholds. Although such methods can capture some obvious defects, they lack sensitivity to subtle abnormal sounds (such as high-frequency whistling and intermittent background noise) that are sensitive to the human ear and are easily affected by the environmental noise of the production line or product individual differences. At the machine learning level, classification models constructed based on algorithms such as support vector machine (SVM) and random forest perform abnormal classification by extracting time-domain features (such as zero-crossing rate, energy entropy) or frequency-domain features (such as Mel spectrum, cepstral coefficients) of acoustic signals. However, the abnormal sound patterns of acoustic products are highly non-linear and random, and it is difficult for traditional feature engineering to cover such complex scenarios. In recent years, deep learning solutions based on convolutional neural network (CNN) have been attempted for end-to-end detection, but their modeling ability is limited, and their adaptability to product design changes is poor. Frequent retraining of the model restricts the feasibility of large-scale deployment.
[0060] Due to the characteristics of strong randomness in the time-frequency distribution and low proportion of abnormal energy of the abnormal sounds of acoustic products, it is difficult for traditional acoustic detection methods to accurately separate abnormal acoustic features from background interference under the influence of noise interference and product structure differences, resulting in high sensitivity of the detection results to working conditions and weak cross-model generalization ability. This makes the high-precision and low false positive rate abnormal sound detection of acoustic products in industrial scenarios an urgent technical bottleneck to be broken through.
[0061] This application aims at the defects of insufficient modeling of long-term sequence dependencies and great limitations in artificial feature design of traditional methods, and proposes an abnormal sound detection framework based on Transformer: by fusing time-frequency domain deep representations and self-attention mechanisms, dynamically capturing the global correlation across time steps in acoustic signals, using multi-head attention to parallelly extract multi-scale abnormal pattern features, and combining lightweight encoder design to achieve efficient feature decoupling in complex noise environments. This solution can significantly improve the sensitivity of the model to weak abnormal sounds, and at the same time quickly adapt to the acoustic characteristic differences of different models of acoustic products through parameter fine-tuning strategies, providing an innovative technical path to solve the problems of high cross-product line migration cost and large fluctuations in false positive rates of traditional models.
[0062] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a heterophony detection device, etc. that can implement the above functions. Hereinafter, the heterophony detection device will be taken as an example to illustrate this embodiment and the following embodiments.
[0063] Based on this, an embodiment of the present application provides a heterophony detection method, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the heterophony detection method of the present application.
[0064] In this embodiment, the heterophony detection method includes steps S10 to S30:
[0065] Step S10, in response to a heterophony detection instruction, the feature extraction unit in the heterophony detection model extracts features from the audio to be detected, and obtains the target audio features of the audio to be detected. The feature extraction unit is composed of multiple twin module groups and a semantic feature extraction module, and there is a hybrid attention mechanism in the twin module groups.
[0066] It should be noted that the heterophony detection instruction is used to trigger the heterophony detection process of the acoustic product to determine whether there is a heterophony in the audio emitted by the acoustic product. The heterophony detection instruction can be actively initiated by the heterophony detection device when detecting the acoustic product, or can be initiated by the user to the heterophony detection device. This embodiment does not limit this. Heterophony refers to an unexpected and non-design-permitted abnormal sound signal generated in the normal working state of the acoustic product, and its manifestation is significantly different from the normal acoustic output. Heterophony includes, but is not limited to, types such as noise, broken sound, abnormal background noise, and structural resonance.
[0067] It can be understood that the heterophony detection model uses Twins-SVT as the backbone network, and is composed of multiple twin module groups, a semantic feature extraction module, and a heterophony classification unit. The Twins-SVT backbone network adopts the Spatially Separable Self-Attention (SVT) mechanism, which is a model architecture including multiple modules, and combines local and global spatial self-attention mechanisms to effectively process data. The Twins-SVT architecture optimizes the computational complexity problem encountered by the traditional Transformer in processing data in a structured manner. The heterophony detection model can gradually reduce the resolution and increase the number of channels, enabling it to efficiently process features of different scales.
[0068] It should be understood that the rise of the Transformer model has provided a new technological breakthrough for the detection of abnormal sounds in acoustic products. The global context modeling ability initially demonstrated by this model in natural language processing exactly matches the long-time sequence dependence characteristics of acoustic signals. Therefore, this embodiment proposes an algorithm that can fuse time-frequency domain depth representations and utilize the global attention mechanism of the Transformer to achieve highly robust abnormal sound detection, which is of breakthrough significance for improving the cross-model adaptability, complex noise suppression ability, and weak abnormal signal capture accuracy of industrial acoustic product abnormal sound detection.
[0069] In specific implementation, the abnormal sound detection model is specifically set as 4 twin module groups. The 4 twin module groups respectively contain 3, 3, 9, and 3 layers to adapt to time sequence-frequency domain feature interaction. Each layer's corresponding twin module (i.e., Twins module) has a hybrid attention mechanism, namely the local attention mechanism (Local Spatial Attention, LSA) and the global attention mechanism (Global Spatial Attention, GSA). In this embodiment, the local attention mechanism is used to calculate self-attention within the window to capture fine-grained features and reduce the computational complexity. The global attention mechanism is used to supplement global information (such as long-time event associations in audio) through cross-window sparse sampling or shifted windows. It should be noted that the semantic feature extraction module refers to the labeled semantic CNN layer, which is used to integrate information in all frequency intervals. After the last twin module group, a convolutional layer with a preset size is also added. The specific connection method of the feature extraction unit is as Figure 2 shown. Each twin-spatial vision transformer in the figure represents a twin module group. Each twin module in each twin module group has a hybrid attention mechanism (i.e., global attention and local attention). After the last twin module group, a new convolutional layer is added, and then the output of the convolutional layer is subjected to a dimension reorganization operation. Finally, the result of the reshaping operation is sent to the semantic feature extraction module (i.e., the labeled semantic convolutional neural network shown in the figure).
[0070] It can be understood that the abnormal sound detection model is obtained by training the initial detection model with a large number of audio with abnormal sound classification labels. The initial detection model consists of multiple untrained twin module groups, a semantic feature extraction module, and an abnormal sound classification unit.
[0071] After receiving the abnormal sound detection instruction, obtain the audio to be detected generated by the acoustic product that needs to be detected for abnormal sound. Extract the features of multiple patch tokens corresponding to the audio to be detected through the feature extraction unit in the abnormal sound detection module, capture the global relevance across the time part of the audio to be detected, and use multi-head attention to parallelly extract multi-scale abnormal pattern features, so as to obtain the target audio features of the audio to be detected. The target audio features include time-frequency local details and global semantics.
[0072] Step S20, input the target audio features into the abnormal sound classification unit in the abnormal sound detection model for abnormal sound detection to obtain the category to which the audio to be detected belongs.
[0073] It should be noted that the abnormal sound classification unit is the decision-making layer in the abnormal sound detection model, responsible for mapping the target audio features to the category mapping space and completing the abnormal sound detection task. The core components of the abnormal sound classification unit include, but are not limited to, fully connected layers, global average pooling, and activation functions.
[0074] It can be understood that input the target audio features into the abnormal sound classification unit in the abnormal sound detection model, and the abnormal sound classification unit maps the target audio features to the category mapping space to obtain the category to which the audio to be detected belongs. In this embodiment, the category to which the audio belongs refers to the category label to which the audio to be detected is classified, including but not limited to normal audio, noise, broken sound, etc.
[0075] In a feasible implementation manner, step S20 may include steps A11 to A13:
[0076] Step A11, input the target audio features into the abnormal sound classification unit in the abnormal sound detection model, perform feature mapping on the target audio features through the fully connected layer in the abnormal sound classification unit, and obtain the category probability map of the audio to be detected according to the mapping result.
[0077] It should be noted that when performing audio classification through the abnormal sound classification unit, first input the target audio features into the fully connected layer in the abnormal sound classification unit. The fully connected layer maps the channel dimension to the number of categories, generates the original category scores at each time point, and finally outputs the category probability map of the audio to be detected according to the original category scores at the time points generated after mapping. For example, if the size of the target audio features is (T / (8P), 8D), then the fully connected layer maps the feature channels from 8D to the number of categories C, and the output category probability map is a (T / (8P), C) matrix.
[0078] Step A12, perform a global average pooling operation on the category probability map through the abnormal sound classification unit to obtain the category probability vector of the audio to be detected.
[0079] It should be noted that by taking the average of the class probability map along the time dimension through the heterophony classification unit, the time resolution is compressed into a single vector, thereby obtaining a global class probability vector, which is used to represent the comprehensive probability that the audio to be detected belongs to each category. For example, if the class probability map is a (T / (8P), C) matrix, then the class probability vector is a (1, C).
[0080] Step A13, perform class determination according to the class probability vector of the audio to be detected, and determine the audio category to which the audio to be detected belongs.
[0081] It should be noted that by using an activation function (such as the Softmax function), a normalized probability distribution is generated based on the class probability vector, and the class with the highest probability is taken as the prediction result, thereby obtaining the audio category to which the audio to be detected belongs.
[0082] Step S30, obtain the audio detection result of the audio to be detected according to the audio category to which the audio to be detected belongs.
[0083] It should be noted that after identifying the audio category to which the audio to be detected belongs through the heterophony detection model, it is possible to clarify whether there is heterophony in the audio to be detected, and when there is heterophony, the specific category of the heterophony, thereby obtaining the final audio detection result.
[0084] This embodiment provides a heterophony detection method. In this embodiment, in response to a heterophony detection instruction, the feature extraction unit in the heterophony detection model extracts features from the audio to be detected to obtain the target audio features of the audio to be detected. The feature extraction unit is composed of multiple twin module groups and a semantic feature extraction module, and there is a hybrid attention mechanism in the twin module groups; input the target audio features into the heterophony classification unit in the heterophony detection model for heterophony detection to obtain the audio category to which the audio to be detected belongs; obtain the audio detection result of the audio to be detected according to the audio category to which the audio to be detected belongs. By the above method, by fusing time-frequency domain depth representation and self-attention mechanism, dynamically capturing the global correlation across time steps in the acoustic signal, using multi-head attention to parallelly extract multi-scale abnormal pattern features, and combining lightweight encoder design to achieve efficient feature decoupling in complex noise environments, the sensitivity of the model to weak heterophony is significantly improved, and high-robustness heterophony detection is realized. At the same time, it can quickly adapt to the acoustic characteristic differences of different models of acoustic products through the parameter fine-tuning strategy, providing an innovative technical path for solving the problems of high cross-product line migration cost and large fluctuation of misjudgment rate in traditional models.
[0085] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3, step S10, the abnormal sound detection method further includes steps S11 to S14:
[0086] Step S11, in response to an abnormal sound detection instruction, perform signal conversion on the audio to be detected to obtain the audio spectrogram of the audio to be detected.
[0087] It should be noted that when receiving an abnormal sound detection instruction, the abnormal sound detection device will obtain the audio to be detected of the corresponding acoustic product. Since the audio to be detected is a wav file (Waveform Audio File Format), the short-time Fourier transform needs to be performed on the audio to be detected, and the Mel filter bank is applied, and then logarithmic compression and Mel spectrogram generation are performed to obtain the audio spectrogram of the audio to be detected. The width of the audio spectrogram represents the time dimension (duration), the height represents the frequency range, and the length of the time axis is usually much larger than the number of frequency ranges.
[0088] Step S12, cut the multiple windows corresponding to the audio spectrogram through a block embedding neural network to obtain multiple block tokens corresponding to the audio spectrogram.
[0089] It should be noted that in order to effectively capture the relationship between different frequencies at the same time, the spectrogram is windowed and segmented. First, the audio spectrogram is divided into multiple windows w1, w2,..., wn, and then the block embedding neural network is used to further divide each window into patches (blocks).
[0090] It can be understood that the multiple windows corresponding to the audio spectrogram are cut into different block tokens (i.e., patch tokens) by using a block embedding neural network. As Figure 4 shown, taking the block embedding neural network with a kernel size of P x P as an example, the width of the audio spectrogram represents the time series T, the height represents the frequency F, w1, w2,..., wn are multiple windows corresponding to the audio spectrogram, and the multiple windows corresponding to the audio spectrogram are segmented through the block convolutional neural network to obtain a series of block tokens patch tokens.
[0091] Step S13, input the multiple block tokens corresponding to the audio spectrogram into the feature extraction unit in the abnormal sound detection model, and perform feature extraction on the multiple block tokens through multiple twin module groups in the feature extraction unit to obtain the associated multi-level features corresponding to the audio spectrogram.
[0092] It should be noted that multiple patch tokens of the audio spectrogram are sequentially input into the feature extraction unit in the abnormal sound detection model. Specifically, they are input into the twin module group in the feature extraction unit according to the priority of the arrangement time, frequency domain, and window of each block token. Multiple twin module groups gradually extract features from multiple block tokens, and different levels of features are generated through a merging operation after each twin module group. The features output by the last twin module of the last twin module group are processed by a convolutional layer with a preset size, and the processed features are the associated multi-level features corresponding to the audio spectrogram.
[0093] Step S14: Aggregate semantic information of the audio spectrogram according to the associated multi-level features and the semantic feature extraction module in the feature extraction unit to obtain the target audio features corresponding to the audio spectrogram.
[0094] It should be noted that the associated multi-level features output by the last twin module group are input into the semantic feature extraction module (i.e., Token-Semantic CNN). The semantic feature extraction module integrates all frequency interval information along the frequency axis to generate time-dimensional features, and the output result of the semantic feature extraction module is the target audio features corresponding to the audio spectrogram.
[0095] It can be understood that in this embodiment, the Token-Semantic CNN layer is used, with enhanced semantic mapping, supporting multiple tasks simultaneously, achieving cross-band semantic integration through convolution in the frequency band dimension, and generating a global classification result using average pooling in the subsequent process, while retaining the time-domain localization ability.
[0096] In a feasible implementation manner, in step S13: Extracting features from multiple block tokens through multiple twin module groups in the feature extraction unit to obtain the associated multi-level features corresponding to the audio spectrogram may include steps B11 to B14:
[0097] Step B11: Input multiple block tokens into multiple twin module groups in the feature extraction unit, and extract features from multiple block tokens through the first twin module group in the multiple twin module groups to obtain first-level features corresponding to each block token.
[0098] It should be noted that the first twin module group refers to the twin module group located first and used to receive the patch tokens output by the Patch-Embed CNN. Features of multiple block tokens are extracted through the twin modules corresponding to multiple levels in the twin module group. The features of each block token output by the twin module of the last level in the first twin module group are the first-level features corresponding to each block token.
[0099] In this embodiment, Patch-Embed CNN is used to implement local feature extraction. By dividing the image into local blocks of a fixed size, the model can focus more on learning local detailed features, reduce the computational complexity, and improve the parallel processing ability.
[0100] Step B12: Merge the multiple block tags according to the first-level features corresponding to each block tag to obtain multiple merged blocks and the second-level features corresponding to each merged block.
[0101] It should be noted that after each twin module group ends, a merging operation is performed on the multiple block tags based on the first-level features corresponding to each block tag to reduce the resolution and expand the channels, thereby obtaining multiple merged blocks and the second-level features corresponding to each merged block. For example, if the window corresponding to the merging operation is 2×2 and the output of the first twin module group is (T / P × F / P, D), the output after merging is (T / 2P × F / 2P, 2D).
[0102] Step B13: Extract features from the audio spectrogram according to the second-level features corresponding to each merged block and the second twin module group in the multiple twin module groups to obtain the third-level features corresponding to the audio spectrogram.
[0103] It should be noted that the second twin module group refers to all the other twin module groups except the first twin module group in the multiple twin module groups. The second-level features corresponding to each merged block are input into the twin module groups located in the front in the second twin module group, and then feature extraction is gradually performed through the multiple twin module groups. The result output by the last layer of the last twin module group in the second twin module group is the third-level feature corresponding to the audio spectrogram.
[0104] Step B14: Perform feature processing on the third-level features corresponding to the audio spectrogram according to a convolutional layer with a preset size to obtain the associated multi-level features corresponding to the audio spectrogram.
[0105] It should be noted that in this embodiment, the convolutional layer with a preset size is set to a convolutional layer with a kernel size of (3, F / 8P) and a padding size of (1, 0), and it can also be adjusted according to the actual situation. This embodiment does not limit this.
[0106] It can be understood that by introducing a convolutional layer with a preset size, the third-level features are further processed to generate a new feature map. Then, a reshape operation is performed on the feature map generated by the convolutional layer to obtain the associated multi-level features that can meet the operation requirements of the subsequent semantic feature extraction module.
[0107] It can be understood that after each twin module group, a global attention anchor is used to guide the merging position, and adjacent patches with semantic relevance are preferentially merged. After being processed by 4 network groups, the size of the patch tokens is reduced by 8 times from (T / P × F / P, D) to (T / 8P× F / 8P, 8D). In this embodiment, the semantic association region is identified through the attention weight, and patches with high correlation are preferentially merged, enabling the features to gradually transition from local details to overall semantics, and avoiding semantic breakage caused by mechanical merging. It can guide cross-region semantic aggregation, enabling logically merging of patches that are not adjacent in physical position but are semantically relevant, breaking through the limitation of the local receptive field of traditional CNNs.
[0108] In a specific implementation, for example, there are four twin module groups 1 to 4. A merging operation is performed based on the output of group 1, and then the result of the merging operation is used as the input of group 2. A merging operation is performed based on the output of group 2, and then the result of the merging operation is used as the input of group 3. A merging operation is performed based on the output of group 3, and the result of the merging operation is used as the input of group 4. The output of group 4 is used as the input of a convolutional layer with a preset size. After further processing the output of group 4 through the convolutional layer, a reshape operation is performed on the feature map generated by the convolutional layer, and the result of the reshape operation (i.e., the associated multi-level features) is used as the input of the Token-Semantic CNN layer.
[0109] In a feasible implementation manner, in step B11: Extracting features of multiple block tokens through the first twin module group in multiple twin module groups to obtain first-level features corresponding to each block token may include steps C11 to C13:
[0110] Step C11, inputting multiple block tokens into the first twin module group in multiple twin module groups, and performing local feature extraction on the multiple block tokens through the local attention mechanism in the first twin module group to obtain local feature representations of each block token.
[0111] Step C12, performing global feature extraction on the multiple block tokens through the global attention mechanism in the first twin module group according to the local feature representations of each block token to obtain global feature representations of each block token.
[0112] Step C13, determining the first-level features corresponding to each block token according to the global feature representations of each block token and the hierarchical structure of the first twin module group.
[0113] It should be noted that since each twin module group includes multiple levels, there is a twin module in each level, and each twin module includes a local attention mechanism and a global attention mechanism. The patch tokens are input into the first twin module group in sequence. The twin modules at each level in the first twin module group first use the local attention mechanism LSA to extract features and capture local acoustic features, so as to obtain the local feature representations of each block token.
[0114] It can be understood that the twin modules at each level use the global attention mechanism GSA, combine the local feature representations of each block token generated by the local attention mechanism at this level, and perform global feature extraction on each block token to obtain the global feature representations of each block token.
[0115] In specific implementation, after the first level in the first twin module group outputs the global feature representations of each block token, the global feature representations of each block token will be input into the second level in the first twin module group. The second level uses the local attention mechanism and the global attention mechanism, combines the global feature representations of each block token output by the first level, performs feature extraction, and inputs the result output by the second level into the third level, repeating the above steps until all levels in the first twin module group have completed feature extraction. In this embodiment, the first-level feature refers to the result generated after the twin module at the last level in the first twin module uses the global attention mechanism GSA for feature extraction.
[0116] It should be noted that the levels of each twin module group are different, but each level executes the above steps, that is, the output of the previous layer - the LSA of the current layer - the GSA of the current layer - the LSA of the next layer - until the end. In this embodiment, there are 4 twin module groups, corresponding to 3, 3, 9, and 3 levels respectively.
[0117] It can be understood that in this embodiment, a hybrid attention mechanism of LSA-GSA is adopted to reduce the computational amount. For the audio patch tokens arranged in the time-frequency-window order, each attention module will calculate the correlation relationship within a specific continuous frequency band and time range. As the network deepens, the global attention anchor points will merge adjacent windows, enabling the attention calculation to be performed in a larger spatial domain. The alternating use of LSA - GSA is as Figure 5 shown in the figure. The figure shows the downsampling process from high resolution to low resolution. The global attention acts on the original feature map on the right. The size of the original feature map is H×W. The global attention captures long-range dependencies by calculating the correlation weights between all positions (H×W); the square with size m×n in the middle on the right represents the process from global localization to local refinement; the local attention divides the global region into m×n sub-regions, and the size of each sub-region is H / m × W / n, as Figure 5For the bottommost feature map in the right part of the middle, local attention only focuses on short-range dependencies within specific sub-regions, significantly reducing the computational complexity. The local attention mechanism LSA enhances the local pattern sensitivity and focuses on capturing local acoustic features. The global attention mechanism GSA addresses the temporal continuity requirements of audio signals. By alternately using LSA and GSA, the problem of semantic fragmentation caused by local windows is avoided.
[0118] This embodiment provides a method for detecting abnormal sounds. In this embodiment, the audio spectrogram of the audio to be detected is obtained by performing signal conversion on the audio to be detected; multiple windows corresponding to the audio spectrogram are cut through a block embedding neural network to obtain multiple block tokens corresponding to the audio spectrogram; the multiple block tokens corresponding to the audio spectrogram are input into the feature extraction unit in the abnormal sound detection model, and multiple block tokens are subjected to feature extraction through multiple twin module groups in the feature extraction unit to obtain the associated multi-level features corresponding to the audio spectrogram; according to the associated multi-level features and the semantic feature extraction module in the feature extraction unit, semantic information aggregation is performed on the audio spectrogram to obtain the target audio features corresponding to the audio spectrogram. Through the above method, the global correlation across time steps in the acoustic signal can be dynamically captured, multi-scale abnormal pattern features can be extracted in parallel using multi-head attention, and efficient feature decoupling in a complex noise environment is achieved by combining a lightweight encoder design.
[0119] Based on the first embodiment and / or the second embodiment of the present application, in the third embodiment of the present application, the same or similar content as in the above-mentioned first and second embodiments can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 6 ..., before step S10 of the abnormal sound detection method, steps S01 to S04 are further included:
[0120] Step S01, sample generation is performed according to multiple sample audios and the abnormal sound classification labels of each sample audio to obtain multiple model training audios and the abnormal sound classification labels of each model training audio.
[0121] It should be noted that the sample audio is the original audio file used for training, including various types of abnormal sounds and normal audio; the abnormal sound classification label is the category identifier of each sample audio, which is used to indicate whether the audio has an abnormal sound and the specific abnormal sound type when there is an abnormal sound.
[0122] It can be understood that in order to ensure the balance and diversity of samples during model training, the sample audio needs to be processed through data augmentation techniques (such as Mix-up, spectral masking, etc.) to generate more training audios, and each generated training audio has its corresponding abnormal sound classification label. In this embodiment, the model training audio is composed of a large number of sample audios and generated training audios.
[0123] Step S02: Divide multiple model training audios to obtain a model training set and a model validation set.
[0124] It should be noted that, in a random selection manner, among multiple model training audios, a preset proportion (such as 80%) of the model training audios are selected as the model training set, and the remaining model training audios are used as the model validation set.
[0125] It can be understood that, in this embodiment, all model training audios are converted into a single-channel format with a sampling rate of 32 kHz. The short-time Fourier transform (STFT) and Mel spectrogram are calculated using a window length of 1024, a hop size of 320, and 64 Mel bands. The shape of the Mel spectrogram is finally (1024, 64). For each 1000-frame (10 seconds) sample, 24 zero frames are padded (total number of frames T = 1024), and the patch size is set to 4×4, with a patch window length of 256 frames.
[0126] Step S03: Train an initial detection model according to the model training set, the abnormal sound classification labels of each audio in the model training set, a learning rate warm-up strategy, and a target optimizer to obtain a first detection model. The initial detection model is composed of multiple initial twin module groups, an initial semantic feature extraction module, and an initial abnormal sound classification unit.
[0127] It should be noted that the model training audios in the model training set are windowed, and after windowing, they are cut into Patch-tokens by Patch-Embed CNN for iterative training input into the network. The learning rate warm-up strategy in this embodiment means that the learning rates in the first three rounds are set to 0.05, 0.1, and 0.2 in sequence, and then the learning rate is halved every ten rounds until it returns to 0.05. The specific value of the learning rate can be adjusted according to the actual situation, and this embodiment does not limit this.
[0128] It can be understood that the target optimizer is set as the AdamW optimizer in this embodiment, where β1 = 0.9, β2 = 0.999, ε = 1e-8, and the weight decay rate is 0.05. It can also be set as other optimizers, and the parameters can also be adjusted according to the actual situation. This embodiment does not limit this. The multiple initial twin module groups are multiple untrained twin module groups; the initial semantic feature extraction module is an untrained Token-Semantic CNN layer; the initial abnormal sound classification unit is an untrained abnormal sound classification unit, and its core components include a fully connected layer, global average pooling, and an activation function.
[0129] In a specific implementation, the initial detection model is trained using the audio in the model training set and its corresponding abnormal sound classification labels, in combination with a weighted average strategy, a learning rate warm-up strategy, an objective optimizer, and pre-trained weights, so as to obtain a preliminarily trained first detection model. During the training process, the binary cross-entropy loss is calculated based on the predicted vectors of the obtained audio and the abnormal sound classification labels of the audio, thereby guiding the model to optimize the parameters and improve the classification accuracy. In this embodiment, the number of training iterations can be set to 200, or it can be adjusted according to requirements, and this embodiment does not limit this.
[0130] It should be noted that during the training process, in order to solve the problem of class imbalance, the probability of different label audios being extracted is adjusted through a balanced sampler, so that the samples of each category can participate in the training process in a relatively balanced manner. Specifically, the weighted sampling method can be used to assign higher weights to the samples of the minority classes to ensure that their probability of being selected in each iteration is higher.
[0131] Step S04, perform performance evaluation on the first detection model according to the model validation set, and determine the abnormal sound detection model according to the performance evaluation result.
[0132] It should be noted that the first detection model is tested using the audio in the model validation set and its abnormal sound classification labels, and various performance indicators are calculated, including but not limited to accuracy, recall rate, and F1 score, etc. If the performance indicators of the first detection model meet the expected standards, it can be used as the final abnormal sound detection model; otherwise, the hyperparameters or structure of the first detection model are adjusted and retrained to obtain the abnormal sound detection model.
[0133] It can be understood that in this embodiment, an abnormal sound detection model based on transformers and time-frequency information is used to improve the calculation efficiency, reduce the training parameters, and has better compatibility than traditional CNNs.
[0134] In a feasible implementation manner, step S01 may further include steps D11 to D14:
[0135] Step D11, randomly extract multiple sample audios to obtain multiple pairs of sample audios.
[0136] Step D12, perform linear interpolation on the multiple pairs of sample audios, and generate the first generated audio of each pair of sample audios and the abnormal sound classification label of each first generated audio according to the linear interpolation result.
[0137] It should be noted that, by using the Mix-up augmentation technique, multiple sample pairs are randomly selected from multiple sample audios for linear interpolation, so as to generate new audios corresponding to each sample audio pair, and use them as the first generated audios. Combining the abnormal sound classification labels of each sample audio, the abnormal sound classification labels of the first generated audios are determined. In this embodiment, the parameter for controlling the mixing ratio is set to 0.5, so that the new samples tend to average-mix the two sample audios.
[0138] Step D13: Randomly mask each first generated audio and each sample audio, and generate multiple second generated audios and the abnormal sound classification labels of each second generated audio according to the masking results.
[0139] It should be noted that, by using the spectral masking technique, for multiple first generated audios and multiple sample audios, multiple audios are randomly selected, and the corresponding spectrograms are spectrally masked. The new samples generated after the masking are the second generated audios. Combining the abnormal sound classification labels of each sample audio, the abnormal sound classification labels of the second generated audios are determined.
[0140] It can be understood that the spectral masking technique includes time masking and frequency masking. The specific operation of time masking is: randomly select a continuous time period on the time axis of the spectrogram, and set all frequency components within this time period to zero, thus simulating the information loss situation in some time periods of the audio signal.
[0141] The specific operation of frequency masking is to randomly select a frequency band with a certain width on the frequency axis of the spectrogram, and cover the information of this frequency band over the entire time period, thus simulating the loss situation of some frequency components in the audio signal. In this embodiment, the time masking length is set to 128 frames, and the frequency masking width is set to 16 frequency bands, which can also be adjusted according to the actual situation, and this embodiment does not limit this.
[0142] Step D14: Obtain multiple model training audios and the abnormal sound classification labels of each model training audio according to each sample audio and the abnormal sound classification labels of each sample audio, each first generated audio and the abnormal sound classification labels of each first generated audio, and each second generated audio and the abnormal sound classification labels of each second generated audio.
[0143] It should be noted that all the new generated audios and the original sample audios are summarized to obtain multiple model training audios; and based on the abnormal sound classification labels of each new generated audio and the abnormal sound classification labels of the sample audios, the abnormal sound classification labels of each model training audio are obtained.
[0144] This embodiment provides a method for detecting abnormal sounds. In this embodiment, multiple model training audio and their abnormal sound classification labels are obtained by generating samples according to multiple sample audios and their abnormal sound classification labels; the multiple model training audio are divided to obtain a model training set and a model validation set; the initial detection model is trained according to the model training set, the abnormal sound classification labels of the audios in the model training set, the learning rate warm-up strategy, and the target optimizer to obtain a first detection model, where the initial detection model is composed of multiple initial twin module groups, an initial semantic feature extraction module, and an initial abnormal sound classification unit; the performance of the first detection model is evaluated according to the model validation set, and an abnormal sound detection model is determined according to the performance evaluation result. In this way, the high robustness of the model is ensured, the generalization ability of the model is improved, the training efficiency and stability are improved, and the convergence process of the model is accelerated.
[0145] Exemplarily, to help understand the implementation process of the abnormal sound detection method obtained by combining this embodiment with the above Embodiment 1 and Embodiment 2, please refer to Figure 7 , Figure 7 A schematic diagram of the brief process of an abnormal sound detection method is provided. Specifically: S1, a sweep signal containing multiple test frequencies is sent to the target speaker, and the audio signal emitted by the target speaker is collected. S2, the Fourier transform and Mel spectrogram of the audio signal are calculated to obtain the spectral signal of the audio varying with frequency. 80% of the samples in the dataset are randomly selected as the training set, and the remaining are used as the validation set. After windowing, it is cut into Patch-tokens by Patch-Embed CNN for input into the network. S3, the detection model gradually reduces the resolution and increases the number of channels, enabling it to efficiently process features of different scales; Twins-SVT adopts the Spatially Separable Self-Attention (SVT) mechanism, which divides the calculation into two steps: Local Self-Attention (LSA): similar to the local window attention of CNN convolution, reducing the computational complexity; Global Self-Attention (GSA): a cross-window attention mechanism to supplement global information. S4, set the training parameters of the model: Select Twins-SVT as the backbone network, use pre-trained weights, set the number of training epochs to 200, use the AdamW optimizer (β1 = 0.9, β2 = 0.999, ε = 1e-8, weight decay rate 0.05), and adopt the learning rate warm-up strategy: the learning rates in the first three rounds are set to 0.05, 0.1, and 0.2 in sequence, and then the learning rate is halved every ten rounds until it returns to 0.05.
[0146] S5. Training the model: Input the Patch-tokens corresponding to the spectrograms obtained in S2 into the improved network in S3, and train the network according to the parameters in S4 to obtain an abnormal sound detection model. S6. Model inference: During inference, first convert the wav file of the detected audio into a spectrogram, and input the patch-tokens corresponding to the spectrogram images to be classified into the classification model trained in S5 for inference to obtain the classification result.
[0147] Through the abnormal sound detection algorithm of the transformer with a multi-level structure proposed in this embodiment, it focuses on solving the problem of insufficient detection accuracy in the abnormal sound detection of acoustic products due to the lack of correlation of long-time acoustic signals and the difficulty in capturing weak abnormal patterns. It can achieve the following beneficial effects including but not limited to: 1) Using the abnormal sound detection algorithm based on the transformer and time-frequency information to improve the calculation efficiency, reduce the training parameters, and have better compatibility than the traditional CNN. (2) Using Patch-Embed CNN to achieve local feature extraction. By dividing the image into local blocks of a fixed size, the model can focus more on learning local detail features, reduce the calculation complexity, and improve the parallel processing ability. (3) Identify the semantic association area through the attention weight, and preferentially merge the patches with high correlation, so that the features gradually transition from local details to overall semantics, avoiding semantic breaks caused by mechanical merging. It can guide the semantic aggregation across regions, enabling the logical merging of patches that are not adjacent in physical location but are semantically related, breaking through the limitation of the local receptive field of the traditional CNN. (4) Use the mechanism of alternating local-global attention (LSA-GSA). LSA enhances the sensitivity to local patterns and focuses on capturing local acoustic features, while GSA solves the requirement for the temporal continuity of audio signals. By alternately using LSA and GSA, the problem of semantic fragmentation caused by local windows is avoided. (5) Use the labeled semantic CNN layer after the final transformer module (i.e., the twin module group). The semantic mapping is enhanced, and multi-task is supported at the same time. Cross-band semantic integration is achieved through convolution in the frequency band dimension, and finally global classification results are generated by average pooling, retaining the time-domain localization ability.
[0148] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the abnormal sound detection method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.
[0149] This application also provides an abnormal sound detection device. Please refer to Figure 8 , the abnormal sound detection device includes:
[0150] An extraction module 10, configured to, in response to a foreign sound detection instruction, extract features of an audio to be detected through a feature extraction unit in a foreign sound detection model, so as to obtain target audio features of the audio to be detected, where the feature extraction unit is composed of a plurality of siamese module groups and a semantic feature extraction module, and there is a hybrid attention mechanism in the siamese module groups.
[0151] A detection module 20, configured to input the target audio features into a foreign sound classification unit in the foreign sound detection model for foreign sound detection, so as to obtain the audio category to which the audio to be detected belongs.
[0152] A processing module 30, configured to obtain an audio detection result of the audio to be detected according to the audio category to which the audio to be detected belongs.
[0153] Optionally, the extraction module 10 is further configured to:
[0154] Perform signal conversion on the audio to be detected to obtain an audio spectrogram of the audio to be detected; perform cutting processing on a plurality of windows corresponding to the audio spectrogram through a block embedding neural network to obtain a plurality of block tokens corresponding to the audio spectrogram; input the plurality of block tokens corresponding to the audio spectrogram into the feature extraction unit in the foreign sound detection model, and extract features of the plurality of block tokens through the plurality of siamese module groups in the feature extraction unit to obtain associated multi-level features corresponding to the audio spectrogram; perform semantic information aggregation on the audio spectrogram according to the associated multi-level features and the semantic feature extraction module in the feature extraction unit to obtain target audio features corresponding to the audio spectrogram.
[0155] Optionally, the extraction module 10 is further configured to:
[0156] Input a plurality of block tokens into the plurality of siamese module groups in the feature extraction unit, and extract features of the plurality of block tokens through a first siamese module group in the plurality of siamese module groups to obtain first-level features corresponding to each block token; merge the plurality of block tokens according to the first-level features corresponding to each block token to obtain a plurality of merged blocks and second-level features corresponding to each merged block; extract features of the audio spectrogram according to the second-level features corresponding to each merged block and a second siamese module group in the plurality of siamese module groups to obtain third-level features corresponding to the audio spectrogram; perform feature processing on the third-level features corresponding to the audio spectrogram through a convolutional layer with a preset size to obtain associated multi-level features corresponding to the audio spectrogram.
[0157] Optionally, the extraction module 10 is further configured to:
[0158] Input multiple block tokens into the first twin module group among multiple twin module groups, perform local feature extraction on the multiple block tokens through the local attention mechanism in the first twin module group to obtain local feature representations of each block token; perform global feature extraction on the multiple block tokens through the global attention mechanism in the first twin module group according to the local feature representations of each block token to obtain global feature representations of each block token; determine first-level features corresponding to each block token according to the global feature representations of each block token and the hierarchical structure of the first twin module group.
[0159] Optionally, the detection module 20 is further configured to:
[0160] Input the target audio feature into the abnormal sound classification unit in the abnormal sound detection model, perform feature mapping on the target audio feature through the fully connected layer in the abnormal sound classification unit, and obtain the category probability map of the audio to be detected according to the mapping result; perform global average pooling operation on the category probability map through the abnormal sound classification unit to obtain the category probability vector of the audio to be detected; perform category determination according to the category probability vector of the audio to be detected to determine the audio category to which the audio to be detected belongs.
[0161] Optionally, the extraction module 10 is further configured to:
[0162] Generate samples according to multiple sample audios and the abnormal sound classification labels of each sample audio to obtain multiple model training audios and the abnormal sound classification labels of each model training audio; divide the multiple model training audios to obtain a model training set and a model validation set; train an initial detection model according to the model training set, the abnormal sound classification labels of each audio in the model training set, the learning rate warm-up strategy, and the target optimizer to obtain a first detection model, where the initial detection model is composed of multiple initial twin module groups, an initial semantic feature extraction module, and an initial abnormal sound classification unit; perform performance evaluation on the first detection model according to the model validation set, and determine the abnormal sound detection model according to the performance evaluation result.
[0163] Optionally, the extraction module 10 is further configured to:
[0164] Randomly extract multiple sample audio to obtain multiple pairs of sample audio; perform linear interpolation on multiple pairs of sample audio, and generate the first generated audio of each pair of sample audio and the abnormal sound classification label of each first generated audio according to the linear interpolation result; randomly mask each first generated audio and each sample audio, and generate multiple second generated audio and the abnormal sound classification label of each second generated audio according to the masking result; obtain multiple model training audio and the abnormal sound classification label of each model training audio according to each sample audio and the abnormal sound classification label of each sample audio, each first generated audio and the abnormal sound classification label of each first generated audio, and each second generated audio and the abnormal sound classification label of each second generated audio.
[0165] The abnormal sound detection device provided by the present application adopts the abnormal sound detection method in the above-mentioned embodiment, and can solve the technical problem that the traditional acoustic detection method has low detection accuracy and generalization ability because it is difficult to accurately separate abnormal acoustic features from background interference. Compared with the prior art, the beneficial effects of the abnormal sound detection device provided by the present application are the same as those of the abnormal sound detection method provided by the above-mentioned embodiment, and other technical features in the abnormal sound detection device are the same as those disclosed in the above-mentioned embodiment method, and will not be elaborated here.
[0166] The present application provides an abnormal sound detection device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the abnormal sound detection method in the first embodiment above.
[0167] Refer to the following Figure 9 , which shows a schematic structural diagram of an abnormal sound detection device suitable for implementing the embodiment of the present application. The abnormal sound detection device in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description), PMPs (Portable Media Player), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The abnormal sound detection device shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present application.
[0168] As Figure 9As shown, the abnormal sound detection device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the abnormal sound detection device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the abnormal sound detection device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an abnormal sound detection device having various systems, it should be understood that it is not required to implement or have all the systems shown. Instead, more or fewer systems may be implemented or had.
[0169] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.
[0170] The abnormal sound detection device provided by the present application adopts the abnormal sound detection method in the above embodiment, and can solve the technical problem that the traditional acoustic detection method has low detection accuracy and generalization ability due to the difficulty in accurately separating abnormal acoustic features from background interference. Compared with the prior art, the beneficial effects of the abnormal sound detection device provided by the present application are the same as those of the abnormal sound detection method provided by the above embodiment, and other technical features in the abnormal sound detection device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.
[0171] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0172] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0173] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the abnormal sound detection method in the above embodiments.
[0174] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems, or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM: Random Access Memory), read-only memory (ROM: Read Only Memory), erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0175] The above computer-readable storage medium can be included in the abnormal sound detection device; it can also exist separately without being assembled into the abnormal sound detection device.
[0176] The above computer-readable storage medium carries one or more programs, which, when executed by the abnormal sound detection device, cause the abnormal sound detection device to: in response to an abnormal sound detection instruction, extract features of the audio to be detected through a feature extraction unit in the abnormal sound detection model to obtain target audio features of the audio to be detected, where the feature extraction unit is composed of a plurality of twin module groups and a semantic feature extraction module, and there is a hybrid attention mechanism in the twin module groups; input the target audio features into an abnormal sound classification unit in the abnormal sound detection model to perform abnormal sound detection and obtain the audio category to which the audio to be detected belongs; and obtain an audio detection result of the audio to be detected according to the audio category to which the audio to be detected belongs.
[0177] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0178] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0179] The modules involved in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.
[0180] The readable storage medium provided in this application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above abnormal sound detection method, which can solve the technical problem that the traditional acoustic detection method has low detection accuracy and generalization ability due to the difficulty in accurately separating abnormal acoustic features from background interference. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the abnormal sound detection method provided in the above embodiments, and will not be elaborated here.
[0181] This application also provides a computer program product, including a computer program, and the steps of the above abnormal sound detection method are implemented when the computer program is executed by a processor.
[0182] The computer program product provided in this application can solve the technical problem that the traditional acoustic detection method has low detection accuracy and generalization ability due to the difficulty in accurately separating abnormal acoustic features from background interference. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the abnormal sound detection method provided in the above embodiments, and will not be elaborated here.
[0183] The above are only some embodiments of this application, and do not limit the patent scope of this application. Any equivalent structural transformation made using the content of the specification and drawings of this application under the technical concept of this application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of this application.
Claims
1. A method for detecting abnormal noise, characterized in that, The method includes: In response to a heterophony detection instruction, a feature extraction unit in a heterophony detection model extracts features from the audio to be detected to obtain target audio features of the audio to be detected. The feature extraction unit consists of a semantic feature extraction module and multiple twin module groups. There is a hybrid attention mechanism in the twin module groups. The twin module groups include twin spatial vision transformers. There is a convolutional layer after the last twin module group. The output of the convolutional layer is subjected to a dimension reorganization operation, and the result of the dimension reorganization operation is input into the semantic feature extraction module; Input the target audio features into a heterophony classification unit in the heterophony detection model for heterophony detection to obtain the audio category to which the audio to be detected belongs; Obtain the audio detection result of the audio to be detected according to the audio category to which the audio to be detected belongs.
2. The method according to claim 1, wherein The step of extracting features from the audio to be detected by a feature extraction unit in a heterophony detection model to obtain target audio features of the audio to be detected includes: Perform signal conversion on the audio to be detected to obtain an audio spectrogram of the audio to be detected; Cut multiple windows corresponding to the audio spectrogram through a block embedding neural network to obtain multiple block tokens corresponding to the audio spectrogram; Input the multiple block tokens corresponding to the audio spectrogram into a feature extraction unit in a heterophony detection model, and extract features from the multiple block tokens through multiple twin module groups in the feature extraction unit to obtain associated multi-level features corresponding to the audio spectrogram; Aggregate semantic information of the audio spectrogram according to the associated multi-level features and the semantic feature extraction module in the feature extraction unit to obtain target audio features corresponding to the audio spectrogram.
3. The method according to claim 2, wherein The step of extracting features from the multiple block tokens through multiple twin module groups in the feature extraction unit to obtain associated multi-level features corresponding to the audio spectrogram includes: Input the multiple block tokens into multiple twin module groups in the feature extraction unit, and extract features from the multiple block tokens through the first twin module group in the multiple twin module groups to obtain first-level features corresponding to each block token; Merge the multiple block tokens according to the first-level features corresponding to each block token to obtain multiple merged blocks and second-level features corresponding to each merged block; Extract features from the audio spectrogram according to the second-level features corresponding to each merged block and the second twin module group in the multiple twin module groups to obtain third-level features corresponding to the audio spectrogram; Perform feature processing on the third-level features corresponding to the audio spectrogram through a convolutional layer with a preset size to obtain associated multi-level features corresponding to the audio spectrogram.
4. The method according to claim 3, wherein The step of extracting features from the multiple block tokens through the first twin module group in the multiple twin module groups to obtain first-level features corresponding to each block token includes: Input the multiple block tokens into the first twin module group in the multiple twin module groups, and perform local feature extraction on the multiple block tokens through the local attention mechanism in the first twin module group to obtain local feature representations of each block token; Global feature extraction is performed on multiple block tokens through the global attention mechanism in the first twin module group according to the local feature representations of each block token, and the global feature representations of each block token are obtained; The first-level features corresponding to each block token are determined according to the global feature representations of each block token and the hierarchical structure of the first twin module group.
5. The method according to claim 1, wherein The step of inputting the target audio feature into the abnormal sound classification unit in the abnormal sound detection model to perform abnormal sound detection and obtaining the audio category to which the audio to be detected belongs includes: Input the target audio feature into the abnormal sound classification unit in the abnormal sound detection model, perform feature mapping on the target audio feature through the fully connected layer in the abnormal sound classification unit, and obtain the category probability map of the audio to be detected according to the mapping result; Perform global average pooling operation on the category probability map through the abnormal sound classification unit to obtain the category probability vector of the audio to be detected; Perform category determination according to the category probability vector of the audio to be detected to determine the audio category to which the audio to be detected belongs.
6. The method according to any one of claims 1 to 5, characterized in that, Before the step of obtaining the target audio feature of the audio to be detected by performing feature extraction on the audio to be detected through the feature extraction unit in the abnormal sound detection model, it further includes: Sample generation is performed according to multiple sample audios and the abnormal sound classification labels of each sample audio to obtain multiple model training audios and the abnormal sound classification labels of each model training audio; Divide multiple model training audios to obtain a model training set and a model validation set; Train the initial detection model according to the model training set, the abnormal sound classification labels of each audio in the model training set, the learning rate warm-up strategy, and the target optimizer to obtain the first detection model. The initial detection model is composed of multiple initial twin module groups, an initial semantic feature extraction module, and an initial abnormal sound classification unit; Perform performance evaluation on the first detection model according to the model validation set, and determine the abnormal sound detection model according to the performance evaluation result.
7. The method according to claim 6, wherein The step of performing sample generation according to multiple sample audios and the abnormal sound classification labels of each sample audio to obtain multiple model training audios and the abnormal sound classification labels of each model training audio includes: Randomly extract multiple sample audios to obtain multiple sample audio pairs; Perform linear interpolation on multiple sample audio pairs, and generate the first generated audio of each sample audio pair and the abnormal sound classification label of each first generated audio according to the linear interpolation result; Randomly mask each first generated audio and each sample audio, and generate multiple second generated audios and the abnormal sound classification labels of each second generated audio according to the masking result; According to each sample audio and the abnormal sound classification label of each sample audio, the first generated audio of each first generated audio and the abnormal sound classification label of each first generated audio, and the second generated audio of each second generated audio and the abnormal sound classification label of each second generated audio, obtain multiple model training audios and the abnormal sound classification labels of each model training audio.
8. A abnormal sound detection device, characterized in that, The abnormal sound detection device includes: An extraction module, configured to, in response to a foreign sound detection instruction, extract features of an audio to be detected through a feature extraction unit in a foreign sound detection model, so as to obtain target audio features of the audio to be detected, where the feature extraction unit is composed of a semantic feature extraction module and a plurality of twin module groups, a hybrid attention mechanism exists in the twin module groups, a twin spatial vision transformer is included in the twin module groups, a convolutional layer exists after the last twin module group, a dimension reorganization operation is performed on the output of the convolutional layer, and the result of the dimension reorganization operation is input into the semantic feature extraction module; A detection module, configured to input the target audio features into a foreign sound classification unit in the foreign sound detection model to perform foreign sound detection, so as to obtain the category to which the audio of the audio to be detected belongs; A processing module, configured to obtain an audio detection result of the audio to be detected according to the category to which the audio of the audio to be detected belongs.
9. A abnormal sound detection device, characterized in that, The device includes: a memory, a processor, and a foreign sound detection program stored on the memory and executable on the processor, where the foreign sound detection program is configured to implement the steps of the foreign sound detection method according to any one of claims 1 to 7.
10. A storage medium, characterized in that, A foreign sound detection program is stored on the storage medium, and when the foreign sound detection program is executed by a processor, the steps of the foreign sound detection method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Video classification method, method, device and equipment for constructing classification model
CN114882399A
Wind turbine generator operation and maintenance decision-making system based on artificial intelligence
CN118690313A