Abnormal sound detection method and device, equipment and storage medium
By using the combination of spectral diagrams and efficient classification models in industrial application scenarios, the problem of low accuracy of odd sound detection in the prior art is solved, and more efficient and accurate odd sound detection is achieved, improving product quality and production efficiency.
Patent Information
- Application Number
- CN202510560163.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The accuracy of the previous technology is low in industrial environments, and is affected by factors such as complex noise, machine jitter and product differences, resulting in inconsistent detection results.
A strange sound detection method is adopted to obtain the audio signal to be detected in industrial application scenarios, determine its time-frequency resolution and signal characteristic information, generate the spectrum diagram to be detected, and input it into an efficient classification model based on pre-trained feature extraction backbone network, cross-stage coordinated attention mechanism and progressive feature pyramid training for inference to obtain strange sound detection results.
It improves the accuracy of abnormal sound detection, enhances the sensitivity to low signal-to-noise ratio abnormal sound, reduces information faults caused by spectrogram resolution jumps, and improves product production efficiency and quality.
Smart Images

Figure CN120089163A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of industrial detection technologies, and particularly to a method, apparatus, device, and storage medium for abnormal sound detection. Background Art
[0002] In the industrial field, abnormal sound detection is a key link to ensure the quality of acoustic products. Currently, the common methods for abnormal sound detection rely on manual auditory judgment or rough recognition by traditional Convolutional Neural Network (CNN) models. However, in an industrial environment, the noise is complex and variable, and the above methods will be affected by factors such as slight machine jitter and product differences, resulting in inconsistent abnormal sound detection results, and the traditional CNN model often makes misrecognition. Based on this, the accuracy of detecting abnormal sounds by the above methods is relatively low.
[0003] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide a method, apparatus, device, and storage medium for abnormal sound detection, aiming to solve the technical problem of relatively low accuracy of detecting abnormal sounds in the prior art.
[0005] To achieve the above purpose, this application proposes an abnormal sound detection method, and the method includes: Obtain an audio signal to be detected in an industrial application scenario, and determine the time-frequency resolution and signal characteristic information of the audio signal to be detected; Generate a spectrogram to be detected according to the time-frequency resolution and the signal characteristic information; Input the spectrogram to be detected into a target efficient classification model, and obtain the current classification information output by the target efficient classification model after inference; wherein, the target efficient classification model is trained according to a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid. Determine the abnormal sound detection result of the audio signal to be detected according to the current classification information.
[0006] In one embodiment, the step of generating a spectrogram to be detected according to the time-frequency resolution and the signal characteristic information includes: Perform dimension detection on the audio signal to be detected; When the dimension detection result meets the preset requirements, create a spectrogram converter according to the time-frequency resolution and the signal characteristic information; Configure multi-dimensional window parameters for the spectrogram converter; The spectrogram converter based on the configuration parameters converts the to-be-detected audio signal to obtain a to-be-detected spectrogram.
[0007] In one embodiment, the step of inputting the to-be-detected spectrogram into the target efficient classification model includes: Obtain a historical audio signal set, and convert the historical audio signal set to obtain a historical spectrogram set; Randomly select the historical spectrogram set according to a target ratio, and determine a sample training set and a sample validation set according to the selection result; Train the current efficient classification model according to the sample training set, the pre-trained feature extraction backbone network, the cross-stage co-attention mechanism, and the progressive feature pyramid; Calculate the loss value between the predicted value and the true value of the current efficient classification model based on the target loss function with weights; Perform validation training on the current efficient classification model according to the sample validation set until the loss value converges to a preset value, and determine the target efficient classification model.
[0008] In one embodiment, the step of training the current efficient classification model according to the sample training set, the pre-trained feature extraction backbone network, the cross-stage co-attention mechanism, and the progressive feature pyramid includes: Obtain the channels of the sample training spectrograms in the sample training set; Unify the channels of the sample training spectrograms through a first-specification convolution, and perform global average pooling on the sample training spectrograms according to the unified channels; Generate a channel weight matrix according to the sample training spectrograms after global average pooling; Generate a spatial weight matrix according to the sample training spectrograms; Based on the cross-stage co-attention mechanism, generate an attention mask according to the spatial weight matrix and the channel weight matrix, and apply the attention mask to the sample training set; Train the current efficient classification model according to the sample training set after applying the attention mask, the pre-trained feature extraction backbone network, and the progressive feature pyramid.
[0009] In one embodiment, the step of generating a spatial weight matrix according to the sample training spectrograms includes: Perform max pooling on the sample training spectrograms, and perform average pooling on the sample training spectrograms after max pooling; Concatenate the sample training spectrograms after average pooling along the unified channels; Generate a spatial weight map according to the second-rule convolution and the concatenated sample training spectrograms; Normalize the spatial weight map to obtain a spatial weight matrix.
[0010] In one embodiment, the step of training the current efficient classification model according to the sample training set after applying the attention mask, the pre-trained feature extraction backbone network, and the progressive feature pyramid includes: Determine multi-level features with enhanced key information according to the sample training set after applying the attention mask and the pre-trained feature extraction backbone network; wherein, the multi-level features include a first enhanced feature, a second enhanced feature, and a third enhanced feature. Adjust the number of channels of the first enhanced feature according to the progressive feature pyramid. Perform double upsampling on the first enhanced feature after adjusting the channel data, and element-wise add the first collected feature to the feature after convolving the second enhanced feature to obtain a first added feature. Perform double upsampling on the first added feature, and add the second collected feature to the feature after convolving the first enhanced feature to obtain a second added feature. Process the second added feature based on a third specification convolution, and element-wise add the processed second added feature to the first added feature to obtain a third added feature. Perform downsampling on the third added feature, and add the third collected feature to the first enhanced feature after adjusting the channel data to obtain a fourth added feature. Perform global average pooling on the second added feature, the third added feature, and the fourth added feature respectively. Train the current efficient classification model according to the pre-trained feature extraction backbone network and the added features after global average pooling.
[0011] In one embodiment, after the step of validating and training the current efficient classification model according to the sample validation set until the loss value converges to a preset value, the method further includes: Determine a sample test set according to the selection result. Test the current efficient classification model after validation training according to the sample test set. Obtain the recall rate and precision rate of the current efficient classification model after validation training according to the model test result. When both the recall rate and the precision rate meet the preset conditions, determine the target efficient classification model.
[0012] In addition, to achieve the above object, the present application also proposes a heterophony detection device, and the heterophony detection device includes: A determination module, configured to obtain an audio signal to be detected in an industrial application scenario, and determine the time-frequency resolution and signal characteristic information of the audio signal to be detected; A generation module, configured to generate a spectrogram to be detected according to the time-frequency resolution and the signal characteristic information; An inference module, configured to input the spectrogram to be detected into a target efficient classification model, and obtain the current classification information output after inference by the target efficient classification model; wherein, the target efficient classification model is trained according to a pre-trained feature extraction backbone network, a cross-stage co-attention mechanism, and a progressive feature pyramid; The determination module is further configured to determine the abnormal sound detection result of the audio signal to be detected according to the current classification information.
[0013] In addition, to achieve the above object, the present application also provides an abnormal sound detection device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the abnormal sound detection method as described above.
[0014] In addition, to achieve the above object, the present application also provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the abnormal sound detection method as described above.
[0015] One or more technical solutions proposed by the present application have at least the following technical effects: by obtaining an audio signal to be detected in an industrial application scenario, and determining the time-frequency resolution and signal characteristic information of the audio signal to be detected; generating a spectrogram to be detected according to the time-frequency resolution and the signal characteristic information; inputting the spectrogram to be detected into a target efficient classification model, and obtaining the current classification information output after inference by the target efficient classification model; wherein, the target efficient classification model is trained according to a pre-trained feature extraction backbone network, a cross-stage co-attention mechanism, and a progressive feature pyramid; determining the abnormal sound detection result of the audio signal to be detected according to the current classification information. In the above manner, by converting the audio signal to be detected from a sound form to a spectrogram to be detected in an image form, rich input information is provided for the target efficient classification model, that is, the cross-stage co-attention mechanism is used to improve the sensitivity of the model to abnormal sounds with low signal-to-noise ratio, and the progressive feature pyramid is used to reduce the information discontinuity caused by the jump of the spectrogram resolution, and then the abnormal sound detection result is determined in combination with the current classification information, so as to effectively improve the accuracy of abnormal sound detection, and further improve the production efficiency and quality of products. Description of the Drawings
[0016] The accompanying drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with this application and, together with the description, are used to explain the principles of this application.
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic flowchart provided for the first embodiment of the abnormal sound detection method of this application; Figure 2 It is a schematic flowchart of the audio signal conversion process provided for the first embodiment of the abnormal sound detection method of this application; Figure 3 It is a schematic flowchart of the processing process of the progressive feature pyramid provided for the first embodiment of the abnormal sound detection method of this application; Figure 4 It is a schematic flowchart provided for the specific implementation of step S20 in the first embodiment of the abnormal sound detection method of this application; Figure 5 It is a schematic diagram of the module structure of the abnormal sound detection device in the embodiment of this application; Figure 6 It is a schematic diagram of the device structure of the hardware operating environment involved in the abnormal sound detection method in the embodiment of this application.
[0019] The realization of the purpose, functional features, and advantages of this application will be further described in conjunction with the embodiments with reference to the accompanying drawings. Specific Embodiments
[0020] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, an abnormal sound detection device, etc. that can implement the above functions. Hereinafter, taking the abnormal sound detection device as an example, this embodiment and the following embodiments will be described.
[0021] Based on this, the embodiments of this application provide an abnormal sound detection method, referring to Figure 1 , Figure 1 It is a schematic flowchart of the first embodiment of the abnormal sound detection method of this application.
[0022] In this embodiment, the abnormal sound detection method includes steps S10 to S40: Step S10, obtain the audio signal to be detected in the industrial application scenario, and determine the time-frequency resolution and signal characteristic information of the audio signal to be detected.
[0023] It should be noted that the industrial application scenario refers to an application scenario with complex and variable noise in the field of industrial detection. In this industrial application scenario, the abnormal sound intensity is low or similar to the background noise. At this time, the accuracy of abnormal sound detection by traditional CNN models is low. Therefore, this embodiment proposes a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a target efficient classification model obtained by progressive feature pyramid training. The cross-stage cooperative attention mechanism is used to improve the sensitivity of the model to abnormal sounds with low signal-to-noise ratio, and the progressive feature pyramid is used to reduce the information discontinuity caused by the jump of spectrogram resolution, so as to effectively improve the accuracy of abnormal sound detection.
[0024] It should be understood that the audio signal to be detected refers to the audio signal for which abnormal sound detection is required in the industrial application scenario. Since the chip sweep signal is a segmented frequency modulation signal design optimized for capturing mechanical abnormal acoustic features, the chip sweep signal can be used to actively stimulate the acoustic response of the speaker to be tested, and the audio signal to be detected can be collected in real time. To ensure the consistency of the collected data, the sampling rate can be 44.1k and the duration can be 5 seconds.
[0025] Step S20: Generate a spectrogram to be detected according to the time-frequency resolution and the signal characteristic information.
[0026] It can be understood that the time-frequency resolution characterizes the ability of the audio signal to be detected to depict details in the time and frequency dimensions. The time dimension reflects the pattern of the signal changing over time, such as abnormal sound bursts and periodic noise. The frequency dimension characterizes the energy distribution in different frequency bands, such as high-frequency pulse noise and low-frequency mechanical resonance. Pixel values are used to correspond to the energy intensity at specific time points and frequency intervals. The signal characteristic information characterizes the characteristic information of the audio signal to be detected. To provide rich input information for the target efficient classification model, the audio signal to be detected is converted from a sound form to an image form of the spectrogram to be detected according to the time-frequency resolution and the signal characteristic information.
[0027] It should be noted that, in addition to the above method for generating the spectrogram to be detected, another method for generating the spectrogram to be detected is proposed in this embodiment, specifically: after obtaining the audio signal to be detected in the industrial application scenario, scan the sequence of the audio signal to be detected, and based on the scanning results, all local maximum points and minimum points are determined, and the maximum points and minimum points are connected by multiple spline interpolations to form the audio signal fluctuation boundary. At this time, the envelope removal process can be performed on the audio signal to be detected according to the audio signal fluctuation boundary, and the intrinsic mode signal components are determined according to the processing results. The intrinsic mode signal components are input into the Hilbert spectrum generation module by calling the parameter input interface, and the spectrogram to be detected generated and output by the Hilbert spectrum generation module is obtained. Among them, the horizontal axis of the spectrogram to be detected represents time, the vertical axis represents frequency, and the color intensity represents the energy or amplitude corresponding to the time point and frequency.
[0028] Step S30: Input the spectrogram to be detected into the target efficient classification model, and obtain the current classification information output by the target efficient classification model after inference; wherein, the target efficient classification model is trained based on a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid.
[0029] It should be understood that the current classification information includes different types of spectrogram image information, for example, background noise type spectrogram image information, normal machine operation type spectrogram image information, and abnormal machine operation type spectrogram image information, etc.
[0030] It can be understood that the target efficient classification model refers to a classification model that infers and outputs the current classification information based on the spectrogram to be detected. The target efficient classification model can be an EfficientNet model. The target efficient classification model can be obtained by training a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid. Among them, the pre-trained feature extraction backbone network can be an EfficientNet-B4 network, which can effectively utilize shallow and deep features, enhance the reuse and transmission of features, and is suitable for processing high-dimensional spectrograms. The MBConv module in the EfficientNet-B4 network integrates an inverted residual structure and an SE (Squeeze-Excitation) channel attention mechanism. Through a depth / width / resolution scaling strategy with a composite coefficient φ = 1.5, the model capacity and computational efficiency are balanced, and the time-frequency multi-scale features of the spectrogram are captured. On this basis, a cross-stage cooperative attention mechanism is also introduced to achieve cross-level feature interaction using the cross-stage cooperative attention mechanism, enhance the response ability to weak abnormal sounds, and improve the sensitivity of the model to abnormal sounds with low signal-to-noise ratio. Using a progressive feature pyramid, through bottom-up lateral connections, the high temporal resolution features of the shallow layer of EfficientNet are retained, and deformable convolution-based dynamic feature fusion is used to accurately align features at different levels in the time-frequency domain, reducing information breaks caused by jumps in spectrogram resolution. Combining a channel reweighting mechanism, the contribution degrees of features at different levels are automatically adjusted to adapt to the scale changes of time-frequency patterns in the spectrogram.
[0031] It should also be noted that in addition to the above method of training the target efficient classification model, this embodiment also proposes another method of training the target efficient classification model. Specifically, after obtaining the historical spectrogram set through conversion, data augmentation is performed on the historical spectrogram set in the manner of a causal augmentation strategy to achieve the purpose of improving the sensitivity of key features. On the other hand, in terms of the network architecture, this embodiment can also select a network architecture that combines a sparse activation network and a fractal neural network, in which the fully connected layers are recursively nested and the inter-class joint nodes are automatically pruned. During training, only 10% - 30% of the neurons in each layer are activated through a gating mechanism, and the model volume is reduced by repeating nested convolutional blocks. The training process is as follows: After the data-augmented historical spectrogram set, the data-augmented historical spectrogram set is input into the network architecture that combines a sparse activation network and a fractal neural network. At this time, the network architecture will capture the multi-scale features of the data-augmented historical spectrogram set bidirectionally to effectively improve the efficiency of feature capture. After training is completed, a sample validation set is also used to verify the trained efficient classification model and a sample test set is used for testing. When the training requirements are met, the target efficient classification model is obtained. Compared with other model training methods, the above method can effectively improve the efficiency and accuracy of training the model.
[0032] Further, before the step of inputting the spectrogram to be detected into the target efficient classification model, the following steps are also included: obtaining a historical audio signal set, and performing conversion on the historical audio signal set to obtain a historical spectrogram set; randomly selecting from the historical spectrogram set according to a target ratio, and determining a sample training set and a sample validation set according to the selection result; training the current efficient classification model according to the sample training set, a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid; calculating the loss value between the predicted value and the true value of the current efficient classification model based on the target loss function with weights; performing validation training on the current efficient classification model according to the sample validation set until the loss value converges to a preset value, and determining the target efficient classification model.
[0033] It can be understood that the historical audio signal set refers to a set composed of each historical audio signal. Refer to Figure 2 , Figure 2 As shown in the schematic diagram of the audio signal conversion process, specifically: on the left is the historical audio signal set, and on the right is the historical spectrogram set. The conversion strategy used can be the Short-Time Fourier Transform (STFT) strategy. The advantage of this is that by dividing the signal through a time window and converting the one-dimensional audio into a three-dimensional spectrogram containing time-frequency-energy according to the STFT conversion strategy, it can simultaneously capture the instantaneous frequency changes and continuous frequency characteristics of abnormal sounds.
[0034] It should be understood that after obtaining the historical spectrogram set, randomly select from the historical spectrogram set according to the target ratio. The target ratio can be 80%. After randomly selecting the sample training set, randomly and evenly divide the remaining 20% to obtain a sample validation set and a sample test set. That is, after determining the pre-trained feature extraction backbone network, add a cross-stage cooperative attention mechanism and a progressive feature pyramid, and train the current efficient classification model according to the sample training set. The target loss function can be a binary cross-entropy function with weights. At this time, calculate the loss value between the predicted value and the true value of the current efficient classification model based on the target loss function with weights, and perform validation training on the current efficient classification model according to the sample validation set until the loss value converges to a preset value. At this time, it can be determined that the current efficient classification model after validation training is the target efficient classification model.
[0035] Further, the steps of training the current efficient classification model according to the sample training set, the pre-trained feature extraction backbone network, the cross-stage cooperative attention mechanism, and the progressive feature pyramid include: obtaining the channels of the sample training spectrograms in the sample training set; unifying the channels of the sample training spectrograms through a first-specification convolution, and performing global average pooling on the sample training spectrograms according to the unified channels; generating a channel weight matrix based on the sample training spectrograms after global average pooling; generating a spatial weight matrix based on the sample training spectrograms; based on the cross-stage cooperative attention mechanism, generating an attention mask according to the spatial weight matrix and the channel weight matrix, and applying the attention mask to the sample training set; training the current efficient classification model according to the sample training set after applying the attention mask, the pre-trained feature extraction backbone network, and the progressive feature pyramid.
[0036] It should be understood that, in order to effectively improve the accuracy of training the current efficient classification model, a cross-stage cooperative attention mechanism is introduced in this embodiment. The cross-stage cooperative attention mechanism is a dual-attention mechanism that dynamically enhances the key information in the feature map from different dimensions through channel weight allocation and spatial region focusing. The cross-stage cooperative attention mechanism can be inserted before the downsampling layer to enable high-order features to reverse-correct the expression of low-order features. For the spectrogram input to a certain stage of the pre-trained feature extraction backbone network, the channels of the sample training spectrogram are unified through a first-specification convolution to reduce the computational amount. The first specification can be 1×1, and global average pooling is performed on the sample training spectrogram according to the unified channels, compressing each channel into 1 scalar. Through two fully connected layers, the first layer reduces the dimension, and the second layer restores the original number of channels. Then, the ReLU activation function is introduced, and the Sigmoid function is used to generate a channel weight matrix based on the sample training spectrogram after global average pooling, with a range of [0,1]. Specifically: 。
[0037] Wherein, represents the channel weight matrix, represents global average pooling.
[0038] It can be understood that after generating the spatial weight matrix and the channel weight matrix, the attention mask is generated in an element-wise multiplication manner. Specifically: 。
[0039] Wherein, represents the attention mask, represents the channel weight matrix, represents the spatial weight matrix.
[0040] It should be understood that after generating the attention mask, the generated attention mask is applied to the original sample training set to achieve the purpose of retaining the residual connection and avoiding the vanishing gradient. Specifically: 。
[0041] Among them, represents the sample training set after applying the attention mask, represents the attention mask, represents the sample training set.
[0042] Furthermore, the step of generating the spatial weight matrix according to the sample training spectrogram includes: performing max pooling on the sample training spectrogram, and performing average pooling on the sample training spectrogram after max pooling; concatenating the sample training spectrogram after average pooling along the unified channels; generating a spatial weight map according to the second rule convolution and the concatenated sample training spectrogram; and normalizing the spatial weight map to obtain the spatial weight matrix.
[0043] It can be understood that in order to effectively improve the accuracy of generating the spatial weight matrix, after obtaining the sample training spectrogram, max pooling and average pooling are respectively performed, and then the sample training spectrogram after average pooling is concatenated along the unified channels. The second specification can be 7×7, and after generating the spatial weight map, the spatial weight map is normalized through Sigmoid. Specifically: 。
[0044] Among them, represents the spatial weight matrix, represents average pooling, represents max pooling.
[0045] Further, the step of training the current efficient classification model according to the sample training set after applying the attention mask, the pre-trained feature extraction backbone network, and the progressive feature pyramid includes: determining multi-level features with enhanced key information according to the sample training set after applying the attention mask and the pre-trained feature extraction backbone network; wherein, the multi-level features include a first enhanced feature, a second enhanced feature, and a third enhanced feature; adjusting the number of channels of the first enhanced feature according to the progressive feature pyramid; doubling the upsampling of the first enhanced feature after adjusting the channel data, and element-wise adding the first collected feature and the feature after convolving the second enhanced feature to obtain a first added feature; doubling the upsampling of the first added feature, and adding the second collected feature and the feature after convolving the first enhanced feature to obtain a second added feature; processing the second added feature based on a third specification convolution, and element-wise adding the processed second added feature and the first added feature to obtain a third added feature; downsampling the third added feature, and adding the third collected feature and the first enhanced feature after adjusting the channel data to obtain a fourth added feature; respectively performing global average pooling on the second added feature, the third added feature, and the fourth added feature; training the current efficient classification model according to the pre-trained feature extraction backbone network and the added features after global average pooling.
[0046] It should be understood that the progressive feature pyramid in this embodiment introduces bidirectional progressive fusion on the basis of the traditional pyramid, enhances the semantic consistency of cross-scale features through multi-stage interaction, and particularly improves the expression ability of small targets and detailed features. After the key information enhancement of the cross-stage collaborative attention mechanism, the pre-trained feature extraction backbone network outputs multi-level features. For example, a first enhanced feature C3', a second enhanced feature C4', and a third enhanced feature C5'. The resolution of these multi-level features decreases, and the semantic level increases.
[0047] It can be understood that referring to Figure 3 , Figure 3It is a schematic diagram of the processing flow of the progressive feature pyramid. Specifically, after obtaining the first enhanced feature, the second enhanced feature, and the third enhanced feature, the number of channels of the first enhanced feature is adjusted according to the progressive feature pyramid and the first specification convolution, and the first acquisition feature is element-wise added to the feature after convolving the second enhanced feature. The upsampling aliasing is eliminated through the third specification convolution, and the third specification can be 3×3. Then, the first added feature is double upsampled, the second acquisition feature is added to the feature after convolving the first enhanced feature, the processed second added feature is element-wise added to the first added feature, and the third added feature is downsampled. At this time, multi-scale features [P3, N4, N5] can be obtained, that is, the second added feature P3, the third added feature N4, and the fourth added feature N5. At this time, global average pooling is also performed on the second added feature, the third added feature, and the fourth added feature respectively, and after splicing, it is input into the classification head, which can utilize high-frequency details and global context simultaneously to improve the classification robustness.
[0048] It should be noted that after obtaining the added feature after global average pooling, set the training parameters of the model. According to the pre-trained feature extraction backbone network and weights, the image size can be set to 380×380, the training epochs are set to 200, the Adam optimizer is used, the initial learning rate is 0.001, warm up for the first 5 epochs, dynamically adjust the learning rate through the cosine annealing algorithm, and train the current efficient classification model.
[0049] Further, after the step of validating and training the current efficient classification model according to the sample validation set until the loss value converges to a preset value, it further includes: determining the sample test set according to the selection result; testing the current efficient classification model after validation training according to the sample test set; obtaining the recall rate and precision rate of the current efficient classification model after validation training according to the model test result; when both the recall rate and the precision rate meet the preset conditions, determining the target efficient classification model.
[0050] It can be understood that after determining the sample test set according to the selection result, the current efficient classification model after validation training is tested according to the sample test set. At this time, the metrics used for measurement can be the recall rate and the precision rate. When it is determined that both the recall rate and the precision rate meet the preset conditions, directly determine the current efficient classification model after validation training as the target efficient classification model. On the contrary, when it is determined that the recall rate and / or the precision rate do not meet the preset conditions, adjust the weights in real time. For example, when the recall rate is low, automatically increase the weight of the current category.
[0051] Step S40, determine the abnormal sound detection result of the audio signal to be detected according to the current classification information.
[0052] It should be understood that after obtaining the current classification information output by the target efficient classification model, the current classification information is subjected to category recognition to achieve the purpose of determining the abnormal sound detection result of the audio signal to be detected from the image dimension, that is, to distinguish abnormal sounds from normal sounds. Among them, background noise and the sounds generated by the normal operation of the machine are normal sounds, and the sounds of abnormal operation of the machine belong to abnormal sounds.
[0053] In this embodiment, the audio signal to be detected in the industrial application scenario is obtained, and the time-frequency resolution and signal characteristic information of the audio signal to be detected are determined; a spectrogram to be detected is generated according to the time-frequency resolution and the signal characteristic information; the spectrogram to be detected is input into the target efficient classification model, and the current classification information output after the target efficient classification model performs inference is obtained; wherein, the target efficient classification model is trained according to a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid; the abnormal sound detection result of the audio signal to be detected is determined according to the current classification information. In the above manner, by adopting the method of converting the audio signal to be detected from the sound form into the spectrogram to be detected in the image form, rich input information is provided for the target efficient classification model, that is, the cross-stage cooperative attention mechanism is used to improve the sensitivity of the model to abnormal sounds with low signal-to-noise ratio, and the progressive feature pyramid is used to reduce the information discontinuity caused by the resolution jump of the spectrogram. Then, the abnormal sound detection result is determined by combining the current classification information, so as to effectively improve the accuracy of detecting abnormal sounds, and further improve the production efficiency and quality of products.
[0054] Based on the first embodiment of the present application, a specific implementation manner for further limiting step S20 in the first embodiment is proposed. For the same or similar content in this specific implementation manner and the first embodiment, reference may be made to the above introduction and will not be repeated hereinafter. Please refer to Figure 4 , and this specific implementation manner includes steps S201 to S204: Step S201, perform dimension detection on the audio signal to be detected.
[0055] Step S202, when the dimension detection result meets the preset requirements, create a spectrogram converter according to the time-frequency resolution and the signal characteristic information.
[0056] It can be understood that the spectrogram converter refers to a device used to convert the audio signal to be detected into a spectrogram to be detected. After performing dimension detection on the audio signal to be detected, it is necessary to determine whether the dimension detection result meets the preset requirements. If so, it indicates that the audio signal to be detected is a one-dimensional time series, and at this time, a spectrogram converter is created according to the time-frequency resolution and the signal characteristic information.
[0057] Step S203, configure multi-dimensional window parameters for the spectrogram converter.
[0058] It should be understood that after creating the spectrogram converter, multi-dimensional window parameters can also be configured for the spectrogram converter, and the multi-dimensional window parameters include but are not limited to the fast Fourier transform window size, the step size of the sliding window, and the window length, etc.
[0059] Step S204, based on the spectrogram converter after configuring the parameters, convert the to-be-detected audio signal to obtain a to-be-detected spectrogram.
[0060] It can be understood that after obtaining the to-be-detected audio signal, input the to-be-detected audio signal into the spectrogram converter after configuring the parameters, and the spectrogram converter after configuring the parameters converts and outputs the to-be-detected spectrogram. The dimension of the to-be-detected spectrogram can be two-dimensional. In addition, for the convenience of the inference of the target efficient classification model, the to-be-detected spectrogram can also be converted to the decibel scale. By adjusting the maximum decibel parameter, the dynamic range of the decibel scale can be flexibly controlled to adapt to different application requirements.
[0061] In this embodiment, dimension detection is performed on the to-be-detected audio signal; when the dimension detection result meets the preset requirements, a spectrogram converter is created according to the time-frequency resolution and the signal characteristic information; multi-dimensional window parameters are configured for the spectrogram converter; based on the spectrogram converter after configuring the parameters, the to-be-detected audio signal is converted to obtain a to-be-detected spectrogram. By the above method, when it is determined that the dimension detection result of the to-be-detected audio signal meets the preset requirements, a spectrogram converter is created in combination with the time-frequency resolution and the signal characteristic information, and multi-dimensional window parameters are configured for the spectrogram converter, and then the to-be-detected audio signal is input into the spectrogram converter after configuring the parameters, and the spectrogram converter after configuring the parameters outputs the to-be-detected spectrogram, so as to effectively improve the accuracy of obtaining the to-be-detected spectrogram.
[0062] This application also provides a foreign sound detection device, please refer to Figure 5 , the foreign sound detection device includes: A determination module 10, configured to obtain a to-be-detected audio signal in an industrial application scenario, and determine the time-frequency resolution and signal characteristic information of the to-be-detected audio signal.
[0063] A generation module 20, configured to generate a to-be-detected spectrogram according to the time-frequency resolution and the signal characteristic information.
[0064] An inference module 30, configured to input the to-be-detected spectrogram into a target efficient classification model, and obtain the current classification information output after inference by the target efficient classification model; wherein, the target efficient classification model is trained according to a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid.
[0065] The determining module 10 is further configured to determine the abnormal sound detection result of the audio signal to be detected according to the current classification information.
[0066] In this embodiment, by acquiring the audio signal to be detected in an industrial application scenario, and determining the time-frequency resolution and signal characteristic information of the audio signal to be detected; generating a spectrogram to be detected according to the time-frequency resolution and the signal characteristic information; inputting the spectrogram to be detected into a target efficient classification model, and obtaining the current classification information output after inference by the target efficient classification model; wherein, the target efficient classification model is trained according to a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid; determining the abnormal sound detection result of the audio signal to be detected according to the current classification information. In the above manner, by adopting the method of converting the audio signal to be detected from a sound form to a spectrogram to be detected in an image form, rich input information is provided for the target efficient classification model, that is, the cross-stage cooperative attention mechanism is used to improve the sensitivity of the model to abnormal sounds with low signal-to-noise ratio, and the progressive feature pyramid is used to reduce the information discontinuity caused by the resolution jump of the spectrogram, and then the abnormal sound detection result is determined in combination with the current classification information, so as to effectively improve the accuracy of detecting abnormal sounds, and further improve the production efficiency and quality of products.
[0067] The abnormal sound detection device provided in this application adopts the abnormal sound detection method in the above embodiment, and can solve the technical problem that the accuracy of detecting abnormal sounds in the prior art is relatively low. Compared with the prior art, the beneficial effects of the abnormal sound detection device provided in this application are the same as those of the abnormal sound detection method provided in the above embodiment, and other technical features in the abnormal sound detection device are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.
[0068] In one embodiment, the generating module 20 is further configured to perform dimension detection on the audio signal to be detected; when the dimension detection result meets a preset requirement, create a spectrogram converter according to the time-frequency resolution and the signal characteristic information; configure multi-dimensional window parameters for the spectrogram converter; and convert the audio signal to be detected based on the spectrogram converter after configuring the parameters to obtain a spectrogram to be detected.
[0069] In one embodiment, the inference module 30 is further configured to obtain a set of historical audio signals, convert the set of historical audio signals to obtain a set of historical spectrograms; randomly select from the set of historical spectrograms according to a target ratio, and determine a sample training set and a sample validation set according to the selection result; train a current high-efficiency classification model according to the sample training set, a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid; calculate a loss value between a predicted value and a true value of the current high-efficiency classification model based on a target loss function with weights; perform validation training on the current high-efficiency classification model according to the sample validation set until the loss value converges to a preset value, and determine a target high-efficiency classification model.
[0070] In one embodiment, the inference module 30 is further configured to obtain channels of sample training spectrograms in the sample training set; unify the channels of the sample training spectrograms through first-specification convolution, and perform global average pooling on the sample training spectrograms according to the unified channels; generate a channel weight matrix according to the sample training spectrograms after global average pooling; generate a spatial weight matrix according to the sample training spectrograms; generate an attention mask based on the cross-stage cooperative attention mechanism according to the spatial weight matrix and the channel weight matrix, and apply the attention mask to the sample training set; train a current high-efficiency classification model according to the sample training set after applying the attention mask, a pre-trained feature extraction backbone network, and a progressive feature pyramid.
[0071] In one embodiment, the inference module 30 is further configured to perform max pooling on the sample training spectrograms, and perform average pooling on the sample training spectrograms after max pooling; splice the sample training spectrograms after average pooling along the unified channels; generate a spatial weight map according to second-rule convolution and the spliced sample training spectrograms; perform normalization processing on the spatial weight map to obtain a spatial weight matrix.
[0072] In one embodiment, the inference module 30 is further configured to determine multi-level features with enhanced key information according to the sample training set after applying the attention mask and the pre-trained feature extraction backbone network; wherein, the multi-level features include a first enhanced feature, a second enhanced feature, and a third enhanced feature; adjust the number of channels of the first enhanced feature according to the progressive feature pyramid; perform double upsampling on the first enhanced feature after adjusting the channel data, and element-wise add the first acquired feature to the feature obtained by convolving the second enhanced feature to obtain a first added feature; perform double upsampling on the first added feature, and add the second acquired feature to the feature obtained by convolving the first enhanced feature to obtain a second added feature; process the second added feature based on a third specification convolution, and element-wise add the processed second added feature to the first added feature to obtain a third added feature; perform downsampling on the third added feature, and add the third acquired feature to the first enhanced feature after adjusting the channel data to obtain a fourth added feature; perform global average pooling on the second added feature, the third added feature, and the fourth added feature respectively; train the current efficient classification model according to the pre-trained feature extraction backbone network and the added features after global average pooling.
[0073] In one embodiment, the inference module 30 is further configured to determine a sample test set according to the selection result; test the currently efficient classification model after verification training according to the sample test set; obtain the recall rate and precision rate of the currently efficient classification model after verification training according to the model test result; and determine the target efficient classification model when both the recall rate and the precision rate meet the preset conditions.
[0074] This application provides a heterophonic sound detection device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the heterophonic sound detection method in the first embodiment above.
[0075] Next, refer to Figure 6, which shows a schematic structural diagram of a heterophonic sound detection device suitable for implementing the embodiments of the present application. The heterophonic sound detection device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description), PMPs (Portable Media Player), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The shown heterophonic sound detection device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0076] As Figure 6 shown, the heterophonic sound detection device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to the program stored in the ROM (Read Only Memory) 1002 or the program loaded from the storage device 1003 into the RAM (Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the heterophonic sound detection device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the heterophonic sound detection device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a heterophonic sound detection device having various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.
[0077] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. The computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the disclosed embodiments of the present application are executed.
[0078] The abnormal sound detection device provided by the present application adopts the abnormal sound detection method in the above-mentioned embodiment, and can solve the technical problem of low accuracy in detecting abnormal sounds in the prior art. Compared with the prior art, the beneficial effects of the abnormal sound detection device provided by the present application are the same as those of the abnormal sound detection method provided by the above-mentioned embodiment, and other technical features in the abnormal sound detection device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.
[0079] It should be understood that each part disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0080] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0081] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the abnormal sound detection method in the above-mentioned embodiment.
[0082] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0083] The above computer-readable storage medium can be included in the abnormal sound detection device; or it can exist independently without being assembled into the abnormal sound detection device.
[0084] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++; and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0085] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems and methods according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0086] The modules described in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.
[0087] The readable storage medium provided by the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned abnormal sound detection method, which can solve the technical problem of low accuracy of detecting abnormal sounds in the prior art. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the abnormal sound detection method provided by the above embodiments, and will not be elaborated here.
[0088] The above are only some embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A method for detecting abnormal sound, characterized in that: The method comprises: Acquire an audio signal to be detected in an industrial application scenario, and determine the time-frequency resolution and signal characteristic information of the audio signal to be detected; Generate a spectrogram to be detected according to the time-frequency resolution and the signal characteristic information; Inputting the to-be-detected spectrogram into a target efficient classification model, and obtaining the current classification information output by the target efficient classification model after inference; wherein the target efficient classification model is obtained based on a pre-trained feature extraction backbone network, a cross-stage collaborative attention mechanism, and a progressive feature pyramid training; The abnormal sound detection result of the audio signal to be detected is determined according to the current classification information.
2. The method according to claim 1, characterized in that The step of generating a spectrogram to be detected according to the time-frequency resolution and the signal characteristic information comprises: Performing dimension detection on the audio signal to be detected; When the dimension detection result meets the preset requirements, creating a spectrogram converter according to the time-frequency resolution and the signal characteristic information; Configuring multi-dimensional window parameters for the spectrogram converter; The spectrogram converter based on the configured parameters converts the audio signal to be detected to obtain the spectrogram to be detected.
3. The method according to claim 1, characterized in that Before the step of inputting the to-be-detected spectrogram into the target efficient classification model, the method further includes: Acquire a historical audio signal set, and convert the historical audio signal set to obtain a historical spectrogram set; Randomly selecting the historical spectrogram set according to the target ratio, and determining a sample training set and a sample verification set according to the selection result; Training a current efficient classification model based on the sample training set, the pre-trained feature extraction backbone network, the cross-stage collaborative attention mechanism, and the progressive feature pyramid; Calculate the loss value between the predicted value and the true value of the current efficient classification model based on the target loss function carrying the weight; The current efficient classification model is verified and trained according to the sample verification set until the loss value converges to a preset value, and a target efficient classification model is determined.
4. The method according to claim 3, characterized in that The step of training the current efficient classification model according to the sample training set, the pre-trained feature extraction backbone network, the cross-stage collaborative attention mechanism and the progressive feature pyramid includes: Obtaining a channel of a sample training spectrogram in the sample training set; Unifying the channels of the sample training spectrogram by a first specification convolution, and performing global average pooling on the sample training spectrogram according to the unified channels; Generate a channel weight matrix based on the sample training spectrogram after global average pooling; Generating a spatial weight matrix according to the sample training spectrogram; Based on the cross-stage collaborative attention mechanism, an attention mask is generated according to the spatial weight matrix and the channel weight matrix, and the attention mask is applied to the sample training set; The current efficient classification model is trained based on the sample training set after the attention mask, the pre-trained feature extraction backbone network and the progressive feature pyramid.
5. The method according to claim 4, characterized in that The step of generating a spatial weight matrix according to the sample training spectrogram comprises: Performing maximum pooling on the sample training spectrogram, and performing average pooling on the sample training spectrogram after maximum pooling; Concatenate the average pooled sample training spectrograms along the unified channels; Generate a spatial weight map based on the sample training spectrogram after convolution and concatenation according to the second rule; The spatial weight map is normalized to obtain a spatial weight matrix.
6. The method according to claim 4, characterized in that The step of training the current efficient classification model according to the sample training set after applying the attention mask, the pre-trained feature extraction backbone network and the progressive feature pyramid includes: Determine a multi-level feature enhanced with key information according to the sample training set after the attention mask is applied and the pre-trained feature extraction backbone network; wherein the multi-level feature includes a first enhanced feature, a second enhanced feature and a third enhanced feature; Adjusting the number of channels of the first enhanced feature according to the progressive feature pyramid; Double upsampling the first enhanced feature after adjusting the channel data, and adding the first acquisition feature and the feature after convolution of the second enhanced feature element by element to obtain a first added feature; Double upsampling the first added feature, and adding the second collected feature to the feature after convolution of the first enhanced feature to obtain a second added feature; Processing the second added feature based on a third specification convolution, and adding the processed second added feature to the first added feature element by element to obtain a third added feature; Downsampling the third added feature, and adding the third acquisition feature to the first enhanced feature after adjusting the channel data to obtain a fourth added feature; Performing global average pooling on the second added feature, the third added feature, and the fourth added feature respectively; The current efficient classification model is trained based on the pre-trained feature extraction backbone network and the added features after global average pooling.
7. The method according to claim 3, characterized in that After the step of performing verification training on the current efficient classification model according to the sample verification set until the loss value converges to a preset value, the method further includes: Determine the sample test set according to the selection results; Testing the current efficient classification model after verification training according to the sample test set; Obtaining the recall rate and precision rate of the current efficient classification model after the verification training according to the model test results; When the recall rate and the precision rate both meet the preset conditions, the target efficient classification model is determined.
8. A device for detecting abnormal sound, characterized in that: The device comprises: A determination module, used to obtain an audio signal to be detected in an industrial application scenario, and determine the time-frequency resolution and signal characteristic information of the audio signal to be detected; A generating module, used for generating a spectrogram to be detected according to the time-frequency resolution and the signal characteristic information; An inference module, used for inputting the to-be-detected spectrogram into a target efficient classification model, and obtaining the current classification information output by the target efficient classification model after inference; wherein the target efficient classification model is obtained based on a pre-trained feature extraction backbone network, a cross-stage collaborative attention mechanism, and a progressive feature pyramid training; The determination module is further used to determine the abnormal sound detection result of the audio signal to be detected according to the current classification information.
9. An abnormal sound detection device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the abnormal sound detection method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the abnormal sound detection method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Air pipe leakage signal detection method based on non-uniform frequency spectrogram
CN116577037A
Hot rolled steel strip surface defect detection method based on improved YOLOv5s network
CN117132827A
Industrial part surface defect detection method and device based on YOLOv8
CN118333940A
Multi-feature time-frequency domain partial discharge insulation defect detection and classification system and method
CN118465465A
Dam crack detection method and device, electronic equipment and storage medium
CN118747838A
Cited By
Audio function detection method and device, equipment and storage medium
CN120727032A