Abnormal Sound Detection Method, Device, Equipment and Storage Medium

By using pre-trained feature extraction backbone network and progressive feature pyramid training in industrial detection, the problem of low accuracy of odd sound detection is solved, more efficient odd sound recognition is achieved, and product quality is improved.

CN120089163BActive Publication Date: 2025-07-22GOERTEK INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510560163.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-22
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The prior art has low accuracy in the industrial field of abnormal sound detection, especially in complex and variable environments. Traditional methods are affected by machine jitter and product differences, resulting in inconsistent detection results and frequent identification errors.

Method used

The pre-trained feature extraction backbone network, a cross-stage collaborative attention mechanism, and a progressive feature pyramid training target efficient classification model, generate a spectral map through time-frequency resolution and signal characteristic information, and use the cross-stage collaborative attention mechanism to improve the model's sensitivity to low signal-to-noise ratio antonyms, reduce information faults, and determine the antonyms detection results based on the current classification information.

Benefits of technology

Improve the accuracy of abnormal sound detection, thereby improving the production efficiency and quality of the product.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089163B_ABST
    Figure CN120089163B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, equipment and storage medium for detecting abnormal sounds, which relates to the field of industrial detection technology, including: determining the time-frequency resolution and signal characteristic information of the audio signal to be detected; generating a spectrogram to be detected according to the time-frequency resolution and signal characteristic information; obtaining the current classification information output by the target efficient classification model; and determining the abnormal sound detection result of the audio signal to be detected according to the current classification information. Through the above-mentioned method, a method of converting the audio signal to be detected from the sound form to the spectrogram to be detected in the image form is adopted to provide rich input information for the target efficient classification model, that is, using the cross-stage collaborative attention mechanism to improve the sensitivity of the model to low signal-to-noise ratio abnormal sounds, using the progressive feature pyramid to reduce the information fault caused by the jump of the spectrogram resolution, and then combining the current classification information to determine the abnormal sound detection result, so as to effectively improve the accuracy of detecting abnormal sounds, thereby improving the production efficiency and quality of the product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of industrial detection technologies, and particularly to a method, device, equipment, and storage medium for abnormal sound detection. Background Art

[0002] In the industrial field, abnormal sound detection is a key link to ensure the quality of acoustic products. Currently, the common methods for abnormal sound detection rely on manual auditory judgment or rough recognition by traditional Convolutional Neural Network (CNN) models. However, the noise in the industrial environment is complex and variable, and the above methods will be affected by factors such as slight machine jitter and product differences, resulting in inconsistent abnormal sound detection results, and the traditional CNN model often makes recognition errors. Based on this, the accuracy of detecting abnormal sounds by the above methods is relatively low.

[0003] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of this application is to provide a method, device, equipment, and storage medium for abnormal sound detection, aiming to solve the technical problem of relatively low accuracy in detecting abnormal sounds in the prior art.

[0005] To achieve the above purpose, this application proposes an abnormal sound detection method, and the method includes:

[0006] Obtain the audio signal to be detected in the industrial application scenario, and determine the time-frequency resolution and signal characteristic information of the audio signal to be detected;

[0007] Generate a spectrogram to be detected according to the time-frequency resolution and the signal characteristic information;

[0008] Input the spectrogram to be detected into the target efficient classification model, and obtain the current classification information output after the target efficient classification model performs inference; wherein, the target efficient classification model is trained according to a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid.

[0009] Determine the abnormal sound detection result of the audio signal to be detected according to the current classification information.

[0010] In one embodiment, the step of generating a spectrogram to be detected according to the time-frequency resolution and the signal characteristic information includes:

[0011] Perform dimension detection on the audio signal to be detected;

[0012] When the dimension detection result meets the preset requirements, a spectrogram converter is created according to the time-frequency resolution and the signal characteristic information;

[0013] Configure multi-dimensional window parameters for the spectrogram converter;

[0014] Based on the spectrogram converter after configuring the parameters, the audio signal to be detected is converted to obtain a spectrogram to be detected.

[0015] In one embodiment, the step of inputting the spectrogram to be detected into the target efficient classification model includes:

[0016] Obtain a historical audio signal set, and convert the historical audio signal set to obtain a historical spectrogram set;

[0017] Randomly select the historical spectrogram set according to a target ratio, and determine a sample training set and a sample validation set according to the selection result;

[0018] Train the current efficient classification model according to the sample training set, the pre-trained feature extraction backbone network, the cross-stage cooperative attention mechanism, and the progressive feature pyramid;

[0019] Calculate the loss value between the predicted value and the true value of the current efficient classification model based on the target loss function with weights;

[0020] Verify and train the current efficient classification model according to the sample validation set until the loss value converges to a preset value, and determine the target efficient classification model.

[0021] In one embodiment, the step of training the current efficient classification model according to the sample training set, the pre-trained feature extraction backbone network, the cross-stage cooperative attention mechanism, and the progressive feature pyramid includes:

[0022] Obtain the channels of the sample training spectrograms in the sample training set;

[0023] Unify the channels of the sample training spectrograms through a first-specification convolution, and perform global average pooling on the sample training spectrograms according to the unified channels;

[0024] Generate a channel weight matrix according to the sample training spectrograms after global average pooling;

[0025] Generate a spatial weight matrix according to the sample training spectrograms;

[0026] Based on the cross-stage cooperative attention mechanism, generate an attention mask according to the spatial weight matrix and the channel weight matrix, and apply the attention mask to the sample training set;

[0027] Train the current efficient classification model according to the sample training set after applying the attention mask, the pre-trained feature extraction backbone network, and the progressive feature pyramid.

[0028] In one embodiment, the step of generating the spatial weight matrix according to the sample training spectrogram includes:

[0029] Perform max pooling on the sample training spectrogram, and perform average pooling on the sample training spectrogram after max pooling;

[0030] Concatenate the sample training spectrogram after average pooling along the unified channels;

[0031] Generate a spatial weight map according to the second rule convolution and the concatenated sample training spectrogram;

[0032] Normalize the spatial weight map to obtain the spatial weight matrix.

[0033] In one embodiment, the step of training the current efficient classification model according to the sample training set after applying the attention mask, the pre-trained feature extraction backbone network, and the progressive feature pyramid includes:

[0034] Determine the multi-level features enhanced with key information according to the sample training set after applying the attention mask and the pre-trained feature extraction backbone network; wherein, the multi-level features include the first enhanced feature, the second enhanced feature, and the third enhanced feature;

[0035] Adjust the number of channels of the first enhanced feature according to the progressive feature pyramid;

[0036] Double-upsample the first enhanced feature after adjusting the channel data, and element-wise add the first acquired feature to the feature after convolving the second enhanced feature to obtain the first added feature;

[0037] Double-upsample the first added feature, and add the second acquired feature to the feature after convolving the first enhanced feature to obtain the second added feature;

[0038] Process the second added feature based on the third specification convolution, and element-wise add the processed second added feature to the first added feature to obtain the third added feature;

[0039] Downsample the third added feature, and add the third acquired feature to the first enhanced feature after adjusting the channel data to obtain the fourth added feature;

[0040] Perform global average pooling on the second added feature, the third added feature, and the fourth added feature respectively;

[0041] Train the current efficient classification model based on the pre-trained feature extraction backbone network and the added features after global average pooling.

[0042] In one embodiment, after the step of validating and training the current efficient classification model according to the sample validation set until the loss value converges to a preset value, the method further includes:

[0043] Determine a sample test set according to the selection result;

[0044] Test the current efficient classification model after validation training according to the sample test set;

[0045] Obtain the recall rate and precision rate of the current efficient classification model after validation training according to the model test result;

[0046] When both the recall rate and the precision rate meet the preset conditions, determine the target efficient classification model.

[0047] In addition, to achieve the above object, the present application further provides a heterophony detection device, which includes:

[0048] A determination module, configured to obtain a to-be-detected audio signal in an industrial application scenario, and determine the time-frequency resolution and signal characteristic information of the to-be-detected audio signal;

[0049] A generation module, configured to generate a to-be-detected spectrogram according to the time-frequency resolution and the signal characteristic information;

[0050] An inference module, configured to input the to-be-detected spectrogram into the target efficient classification model, and obtain the current classification information output after inference by the target efficient classification model; wherein, the target efficient classification model is trained based on a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid;

[0051] The determination module is further configured to determine the heterophony detection result of the to-be-detected audio signal according to the current classification information.

[0052] In addition, to achieve the above object, the present application further provides a heterophony detection device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the computer program is configured to implement the steps of the heterophony detection method as described above.

[0053] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the heterophony detection method as described above are implemented.

[0054] One or more technical solutions proposed in this application have at least the following technical effects: by acquiring the audio signal to be detected in an industrial application scenario, and determining the time-frequency resolution and signal characteristic information of the audio signal to be detected; generating a spectrogram to be detected according to the time-frequency resolution and the signal characteristic information; inputting the spectrogram to be detected into a target efficient classification model, and obtaining the current classification information output by the target efficient classification model after inference; wherein, the target efficient classification model is trained according to a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid; determining the abnormal sound detection result of the audio signal to be detected according to the current classification information. By the above method, by converting the audio signal to be detected from a sound form to a spectrogram to be detected in the form of an image, rich input information is provided for the target efficient classification model, that is, the cross-stage cooperative attention mechanism is used to improve the sensitivity of the model to abnormal sounds with low signal-to-noise ratio, and the progressive feature pyramid is used to reduce the information discontinuity caused by the jump of the spectrogram resolution, and then the abnormal sound detection result is determined in combination with the current classification information, so as to effectively improve the accuracy of detecting abnormal sounds, and further improve the production efficiency and quality of products. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.

[0056] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0057] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the abnormal sound detection method of this application;

[0058] Figure 2 It is a schematic flowchart of the audio signal conversion process provided for Embodiment 1 of the abnormal sound detection method of this application;

[0059] Figure 3 It is a schematic flowchart of the processing process of the progressive feature pyramid provided for Embodiment 1 of the abnormal sound detection method of this application;

[0060] Figure 4 It is a schematic flowchart provided for the specific implementation manner of step S20 in Embodiment 1 of the abnormal sound detection method of this application;

[0061] Figure 5 It is a schematic diagram of the module structure of the abnormal sound detection device in the embodiment of this application;

[0062] Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the abnormal sound detection method in the embodiments of this application.

[0063] The realization of the purpose, functional characteristics and advantages of this application will be further described in conjunction with the embodiments with reference to the accompanying drawings. Specific embodiments

[0064] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, an abnormal sound detection device, etc. that can realize the above functions. The following takes the abnormal sound detection device as an example to illustrate this embodiment and the following embodiments.

[0065] Based on this, the embodiments of this application provide an abnormal sound detection method, with reference to Figure 1 , Figure 1 This is a schematic flowchart of the first embodiment of the abnormal sound detection method of this application.

[0066] In this embodiment, the abnormal sound detection method includes steps S10 to S40:

[0067] Step S10, obtain the audio signal to be detected in the industrial application scenario, and determine the time-frequency resolution and signal characteristic information of the audio signal to be detected.

[0068] It should be noted that the industrial application scenario refers to an application scenario with complex and variable noise in the industrial detection field. The abnormal sound intensity in this industrial application scenario is low or similar to the background noise. At this time, the accuracy of abnormal sound detection by the traditional CNN model is low. For this reason, this embodiment proposes a pre-trained feature extraction backbone network, a cross-stage collaborative attention mechanism, and a target efficient classification model obtained by progressive feature pyramid training. The cross-stage collaborative attention mechanism is used to improve the sensitivity of the model to abnormal sounds with low signal-to-noise ratio, and the progressive feature pyramid is used to reduce the information break caused by the jump of the spectrogram resolution, so as to effectively improve the accuracy of abnormal sound detection.

[0069] It should be understood that the audio signal to be detected refers to the audio signal that needs to be detected for abnormal sounds in the industrial application scenario. Since the chip sweep signal is a segmented frequency modulation signal design optimized for capturing mechanical abnormal acoustic characteristics, the chip sweep signal can be used to actively stimulate the acoustic response of the speaker to be tested and collect the audio signal to be detected in real time. To ensure the consistency of the collected data, the sampling can be 44.1k and the duration can be 5 seconds.

[0070] Step S20, generate a spectrogram to be detected according to the time-frequency resolution and the signal characteristic information.

[0071] It is understandable that the time-frequency resolution characterizes the ability to depict the details of the audio signal to be detected in the time and frequency dimensions. The time dimension reflects the pattern of the signal changing over time. For example, abnormal sound bursts and periodic noises. The frequency dimension characterizes the energy distribution in different frequency bands. For example, high-frequency pulse noises and low-frequency mechanical resonances. Pixel values correspond to the energy intensity at specific time points and frequency intervals. The signal characteristic information characterizes the characteristic information of the audio signal to be detected. It provides rich input information for the target efficient classification model, and converts the audio signal to be detected from the sound form to the spectrogram to be detected in the image form according to the time-frequency resolution and the signal characteristic information.

[0072] It should be noted that, in addition to the above method of generating the spectrogram to be detected, this embodiment also proposes another method of generating the spectrogram to be detected. Specifically, after obtaining the audio signal to be detected in the industrial application scenario, scan the sequence of the audio signal to be detected, and based on the scanning results, find all local maximum points and minimum points, and connect the maximum points and minimum points through multiple spline interpolations to form the audio signal fluctuation boundary. At this time, the envelope removal process can be performed on the audio signal to be detected according to the audio signal fluctuation boundary, and the intrinsic mode signal component can be determined according to the processing result. The intrinsic mode signal component is input into the Hilbert spectrum generation module by calling the parameter input interface, and the spectrogram to be detected generated and output by the Hilbert spectrum generation module is obtained. Among them, the horizontal axis of the spectrogram to be detected represents time, the vertical axis represents frequency, and the color intensity represents the energy or amplitude corresponding to the time point and frequency.

[0073] Step S30, input the spectrogram to be detected into the target efficient classification model, and obtain the current classification information inferred and output by the target efficient classification model; among them, the target efficient classification model is trained according to a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid.

[0074] It should be understood that the current classification information includes different types of spectrogram image information. For example, background noise type spectrogram image information, machine normal operation type spectrogram image information, and machine abnormal operation type spectrogram image information, etc.

[0075] It can be understood that the target efficient classification model refers to a classification model that infers and outputs the current classification information based on the spectrogram to be detected. The target efficient classification model can be an EfficientNet model. The target efficient classification model can be obtained by training a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid. Among them, the pre-trained feature extraction backbone network can be an EfficientNet-B4 network, which can effectively utilize shallow and deep features, enhance the reuse and transmission of features, and is suitable for processing high-dimensional spectrograms. The MBConv module in the EfficientNet-B4 network integrates an inverted residual structure and an SE (Squeeze-Excitation) channel attention mechanism. Through a depth / width / resolution scaling strategy with a composite coefficient φ = 1.5, the model capacity and computational efficiency are balanced to capture the time-frequency multi-scale features of the spectrogram. On this basis, a cross-stage cooperative attention mechanism is also introduced to achieve cross-level feature interaction using the cross-stage cooperative attention mechanism, enhance the response ability to weak abnormal sounds, and improve the sensitivity of the model to abnormal sounds with low signal-to-noise ratio. Using a progressive feature pyramid, through bottom-up lateral connections, the high temporal resolution features of the shallow layer of EfficientNet are retained, and deformable convolution is used for dynamic feature fusion to accurately align features at different levels in the time-frequency domain, reduce information breaks caused by spectrogram resolution jumps, and combine a channel reweighting mechanism to automatically adjust the contribution degrees of features at different levels to adapt to the scale changes of time-frequency patterns in the spectrogram.

[0076] It should also be noted that in addition to the above method of training the target efficient classification model, this embodiment also proposes another method of training the target efficient classification model. Specifically, after obtaining the historical spectrogram set through conversion, data augmentation is performed on the historical spectrogram set in the manner of a causal augmentation strategy to achieve the purpose of improving the sensitivity of key features. On the other hand, in terms of the network architecture, this embodiment can also select a network architecture that combines a sparse activation network and a fractal neural network, in which the fully connected layers are recursively nested and the joints between classes are automatically pruned. During training, only 10% - 30% of the neurons in each layer are activated through a gating mechanism, and the model volume is reduced by repeating nested convolutional blocks. The training process is as follows: After the data-augmented historical spectrogram set, the data-augmented historical spectrogram set is input into the network architecture that combines a sparse activation network and a fractal neural network. At this time, the network architecture will capture the multi-scale features of the data-augmented historical spectrogram set bidirectionally to effectively improve the efficiency of feature capture. After training is completed, a sample validation set is also used to verify the trained efficient classification model and a sample test set is used for testing. When the training requirements are met, the target efficient classification model is obtained. Compared with other model training methods, the above method can effectively improve the efficiency and accuracy of training the model.

[0077] Further, before the step of inputting the spectrogram to be detected into the target efficient classification model, the following steps are further included: obtaining a historical audio signal set, and converting the historical audio signal set to obtain a historical spectrogram set; randomly selecting from the historical spectrogram set according to a target ratio, and determining a sample training set and a sample validation set according to the selection result; training the current efficient classification model according to the sample training set, a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid; calculating the loss value between the predicted value and the true value of the current efficient classification model based on a target loss function with weights; validating and training the current efficient classification model according to the sample validation set until the loss value converges to a preset value, and determining the target efficient classification model.

[0078] It can be understood that the historical audio signal set refers to a set composed of various historical audio signals. Refer to Figure 2 , Figure 2 FIG. [FIGURE NUMBER] is a schematic diagram of the audio signal conversion process. Specifically, on the left is the historical audio signal set, and on the right is the historical spectrogram set. The conversion strategy used can be the Short-Time Fourier Transform (STFT) strategy. The advantage of this is that by dividing the signal through a time window and converting the one-dimensional audio into a three-dimensional spectrogram containing time-frequency-energy according to the STFT conversion strategy, it can simultaneously capture the instantaneous frequency changes and continuous frequency characteristics of different sounds.

[0079] It should be understood that after obtaining the historical spectrogram set, randomly select from the historical spectrogram set according to the target ratio. The target ratio can be 80%. After randomly selecting the sample training set, randomly and evenly divide the remaining 20% to obtain a sample validation set and a sample test set. That is, after determining the pre-trained feature extraction backbone network, add a cross-stage cooperative attention mechanism and a progressive feature pyramid, and train the current efficient classification model according to the sample training set. The target loss function can be a binary cross-entropy function with weights. At this time, calculate the loss value between the predicted value and the true value of the current efficient classification model based on the target loss function with weights, and perform validation training on the current efficient classification model according to the sample validation set until the loss value converges to a preset value. At this time, the current efficient classification model after validation training can be determined as the target efficient classification model.

[0080] Further, the steps of training the current efficient classification model according to the sample training set, the pre-trained feature extraction backbone network, the cross-stage cooperative attention mechanism, and the progressive feature pyramid include: obtaining the channels of the sample training spectrograms in the sample training set; unifying the channels of the sample training spectrograms through a first-specification convolution, and performing global average pooling on the sample training spectrograms according to the unified channels; generating a channel weight matrix based on the sample training spectrograms after global average pooling; generating a spatial weight matrix based on the sample training spectrograms; based on the cross-stage cooperative attention mechanism, generating an attention mask according to the spatial weight matrix and the channel weight matrix, and applying the attention mask to the sample training set; training the current efficient classification model according to the sample training set after applying the attention mask, the pre-trained feature extraction backbone network, and the progressive feature pyramid.

[0081] It should be understood that, in order to effectively improve the accuracy of training the current efficient classification model, a cross-stage cooperative attention mechanism is introduced in this embodiment. The cross-stage cooperative attention mechanism is a dual attention mechanism. By channel weight assignment and spatial region focusing, it dynamically enhances the key information in the feature map from different dimensions. The cross-stage cooperative attention mechanism can be inserted before the downsampling layer to enable high-order features to reverse-correct the expression of low-order features. For the spectrogram input to a certain stage of the pre-trained feature extraction backbone network, the channels of the sample training spectrogram are unified through a first-specification convolution to reduce the computational amount. The first specification can be 1×1, and global average pooling is performed on the sample training spectrogram according to the unified channels, compressing each channel into 1 scalar. Through two fully connected layers, the first layer reduces the dimension, and the second layer restores the original number of channels. Then, the ReLU activation function is introduced, and the Sigmoid function is used to generate a channel weight matrix based on the sample training spectrogram after global average pooling, with a range of [0,1]. Specifically:

[0082] 。

[0083] Wherein, represents the channel weight matrix, represents global average pooling.

[0084] It can be understood that after generating the spatial weight matrix and the channel weight matrix, an attention mask is generated in an element-wise multiplication manner. Specifically:

[0085] 。

[0086] Wherein, represents the attention mask, represents the channel weight matrix, represents the spatial weight matrix.

[0087] It should be understood that after generating the attention mask, the generated attention mask is applied to the original sample training set to achieve the purpose of retaining the residual connection and avoiding gradient disappearance. Specifically:

[0088] .

[0089] Among them, represents the sample training set after applying the attention mask, represents the attention mask, represents the sample training set.

[0090] Furthermore, the step of generating the spatial weight matrix according to the sample training spectrogram includes: performing max pooling on the sample training spectrogram, and performing average pooling on the sample training spectrogram after max pooling; concatenating the sample training spectrogram after average pooling along the unified channels; generating a spatial weight map according to the second rule convolution and the concatenated sample training spectrogram; and performing normalization processing on the spatial weight map to obtain the spatial weight matrix.

[0091] It can be understood that in order to effectively improve the accuracy of generating the spatial weight matrix, after obtaining the sample training spectrogram, max pooling and average pooling are respectively performed, and then the sample training spectrogram after average pooling is concatenated along the unified channels. The second specification can be 7×7, and after generating the spatial weight map, the spatial weight map is normalized through Sigmoid. Specifically:

[0092] .

[0093] Among them, represents the spatial weight matrix, represents average pooling, represents max pooling.

[0094] Further, the step of training the current efficient classification model according to the sample training set after applying the attention mask, the pre-trained feature extraction backbone network, and the progressive feature pyramid includes: determining multi-level features with enhanced key information according to the sample training set after applying the attention mask and the pre-trained feature extraction backbone network; wherein, the multi-level features include a first enhanced feature, a second enhanced feature, and a third enhanced feature; adjusting the number of channels of the first enhanced feature according to the progressive feature pyramid; doubling the upsampling of the first enhanced feature after adjusting the channel data, and element-wise adding the first collected feature to the feature after convolving the second enhanced feature to obtain a first added feature; doubling the upsampling of the first added feature, and adding the second collected feature to the feature after convolving the first enhanced feature to obtain a second added feature; processing the second added feature based on a third specification convolution, and element-wise adding the processed second added feature to the first added feature to obtain a third added feature; downsampling the third added feature, and adding the third collected feature to the first enhanced feature after adjusting the channel data to obtain a fourth added feature; performing global average pooling on the second added feature, the third added feature, and the fourth added feature respectively; training the current efficient classification model according to the pre-trained feature extraction backbone network and the added features after global average pooling.

[0095] It should be understood that the progressive feature pyramid in this embodiment introduces bidirectional progressive fusion on the basis of the traditional pyramid, enhances the semantic consistency of cross-scale features through multi-stage interaction, and particularly improves the expression ability of small targets and detailed features. After the key information enhancement by the cross-stage collaborative attention mechanism, the pre-trained feature extraction backbone network outputs multi-level features. For example, the first enhanced feature C3', the second enhanced feature C4', and the third enhanced feature C5'. The resolution of these multi-level features decreases and the semantic level increases.

[0096] It can be understood that referring to Figure 3 , Figure 3It is a schematic diagram of the processing flow of the progressive feature pyramid. Specifically, after obtaining the first enhanced feature, the second enhanced feature, and the third enhanced feature, the number of channels of the first enhanced feature is adjusted according to the progressive feature pyramid and the first specification convolution, and the first acquired feature is element-wise added to the feature obtained by convolving the second enhanced feature. The upsampling aliasing is eliminated through the third specification convolution, and the third specification can be 3×3. Then, the first added feature is double upsampled, the second acquired feature is added to the feature obtained by convolving the first enhanced feature, the processed second added feature is element-wise added to the first added feature, and the third added feature is downsampled. At this time, multi-scale features [P3, N4, N5] can be obtained, that is, the second added feature P3, the third added feature N4, and the fourth added feature N5. At this time, global average pooling is also performed on the second added feature, the third added feature, and the fourth added feature respectively, and after splicing, it is input into the classification head, which can utilize both high-frequency details and global context to improve the classification robustness.

[0097] It should be noted that after obtaining the added feature after global average pooling, the training parameters of the model are set. According to the pre-trained feature extraction backbone network and weights, the image size can be set to 380×380, the training epochs are set to 200, the Adam optimizer is used, the initial learning rate is 0.001, the first 5 epochs are for warm-up, and the learning rate is dynamically adjusted through the cosine annealing algorithm, and the current efficient classification model is trained.

[0098] Further, after the step of validating and training the current efficient classification model according to the sample validation set until the loss value converges to a preset value, it further includes: determining a sample test set according to the selection result; testing the current efficient classification model after validation training according to the sample test set; obtaining the recall rate and precision rate of the current efficient classification model after validation training according to the model test result; and determining the target efficient classification model when both the recall rate and the precision rate meet the preset conditions.

[0099] It can be understood that after determining the sample test set according to the selection result, the current efficient classification model after validation training is tested according to the sample test set. At this time, the metrics used for measurement can be the recall rate and the precision rate. When it is determined that both the recall rate and the precision rate meet the preset conditions, the current efficient classification model after validation training is directly determined as the target efficient classification model. On the contrary, when it is determined that the recall rate and / or the precision rate do not meet the preset conditions, the weights are adjusted in real time. For example, when the recall rate is low, the weight of the current category is automatically increased.

[0100] Step S40, determining the abnormal sound detection result of the audio signal to be detected according to the current classification information.

[0101] It should be understood that after obtaining the current classification information output by the target efficient classification model, the current classification information is subjected to category recognition to achieve the purpose of determining the abnormal sound detection result of the audio signal to be detected from the image dimension, that is, to distinguish abnormal sounds from normal sounds. Among them, background noise and the sounds generated by the normal operation of the machine are normal sounds, and the sounds of abnormal operation of the machine belong to abnormal sounds.

[0102] In this embodiment, the audio signal to be detected in the industrial application scenario is obtained, and the time-frequency resolution and signal characteristic information of the audio signal to be detected are determined; a spectrogram to be detected is generated according to the time-frequency resolution and the signal characteristic information; the spectrogram to be detected is input into the target efficient classification model, and the current classification information output after the target efficient classification model is inferred is obtained; wherein, the target efficient classification model is trained according to a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid; the abnormal sound detection result of the audio signal to be detected is determined according to the current classification information. In the above manner, by adopting the spectrogram to be detected that converts the audio signal to be detected from the sound form to the image form, rich input information is provided for the target efficient classification model, that is, the cross-stage cooperative attention mechanism is used to improve the sensitivity of the model to abnormal sounds with low signal-to-noise ratio, and the progressive feature pyramid is used to reduce the information discontinuity caused by the resolution jump of the spectrogram, and then the abnormal sound detection result is determined in combination with the current classification information, so as to effectively improve the accuracy of detecting abnormal sounds, and further improve the production efficiency and quality of products.

[0103] Based on the first embodiment of the present application, a specific implementation manner for further limiting step S20 in the first embodiment is proposed. For the same or similar content in this specific implementation manner and the first embodiment, reference can be made to the above introduction and will not be repeated hereinafter. Please refer to Figure 4 , this specific implementation manner includes steps S201 to S204:

[0104] Step S201, perform dimension detection on the audio signal to be detected.

[0105] Step S202, when the dimension detection result meets the preset requirements, create a spectrogram converter according to the time-frequency resolution and the signal characteristic information.

[0106] It can be understood that the spectrogram converter refers to a device for converting the audio signal to be detected into a spectrogram to be detected. After performing dimension detection on the audio signal to be detected, it is necessary to determine whether the dimension detection result meets the preset requirements. If so, it indicates that the audio signal to be detected is a one-dimensional time series. At this time, a spectrogram converter is created according to the time-frequency resolution and the signal characteristic information.

[0107] Step S203, configure multi-dimensional window parameters for the spectrogram converter.

[0108] It should be understood that after creating the spectrogram converter, multi-dimensional window parameters can also be configured for the spectrogram converter, and the multi-dimensional window parameters include but are not limited to the fast Fourier transform window size, the step size of the sliding window, and the window length, etc.

[0109] Step S204, based on the spectrogram converter after configuring the parameters, convert the to-be-detected audio signal to obtain a to-be-detected spectrogram.

[0110] It can be understood that after obtaining the to-be-detected audio signal, input the to-be-detected audio signal into the spectrogram converter after configuring the parameters, and the spectrogram converter after configuring the parameters converts and outputs the to-be-detected spectrogram. The dimension of the to-be-detected spectrogram can be two-dimensional. In addition, for the convenience of the inference of the target efficient classification model, the to-be-detected spectrogram can also be converted to the decibel scale. By adjusting the maximum decibel parameter, the dynamic range of the decibel scale can be flexibly controlled to adapt to different application requirements.

[0111] In this embodiment, dimension detection is performed on the to-be-detected audio signal; when the dimension detection result meets the preset requirements, a spectrogram converter is created according to the time-frequency resolution and the signal characteristic information; multi-dimensional window parameters are configured for the spectrogram converter; based on the spectrogram converter after configuring the parameters, the to-be-detected audio signal is converted to obtain a to-be-detected spectrogram. In the above manner, when it is determined that the dimension detection result of the to-be-detected audio signal meets the preset requirements, a spectrogram converter is created in combination with the time-frequency resolution and the signal characteristic information, and multi-dimensional window parameters are configured for the spectrogram converter, and then the to-be-detected audio signal is input into the spectrogram converter after configuring the parameters, and the spectrogram converter after configuring the parameters outputs the to-be-detected spectrogram, so as to effectively improve the accuracy of obtaining the to-be-detected spectrogram.

[0112] This application also provides a foreign sound detection device, please refer to Figure 5 , the foreign sound detection device includes:

[0113] A determination module 10, configured to obtain a to-be-detected audio signal in an industrial application scenario, and determine the time-frequency resolution and signal characteristic information of the to-be-detected audio signal.

[0114] A generation module 20, configured to generate a to-be-detected spectrogram according to the time-frequency resolution and the signal characteristic information.

[0115] An inference module 30, configured to input the to-be-detected spectrogram into a target efficient classification model, and obtain the current classification information output after inference by the target efficient classification model; wherein, the target efficient classification model is trained according to a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid.

[0116] The determining module 10 is further configured to determine the abnormal sound detection result of the audio signal to be detected according to the current classification information.

[0117] In this embodiment, the audio signal to be detected in the industrial application scenario is obtained, and the time-frequency resolution and signal characteristic information of the audio signal to be detected are determined; a spectrogram to be detected is generated according to the time-frequency resolution and the signal characteristic information; the spectrogram to be detected is input into the target efficient classification model, and the current classification information output after the target efficient classification model is inferred is obtained; wherein, the target efficient classification model is trained according to a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid; the abnormal sound detection result of the audio signal to be detected is determined according to the current classification information. By the above method, the spectrogram to be detected, which converts the audio signal to be detected from the sound form to the image form, is used to provide rich input information for the target efficient classification model, that is, the cross-stage cooperative attention mechanism is used to improve the sensitivity of the model to abnormal sounds with low signal-to-noise ratio, and the progressive feature pyramid is used to reduce the information discontinuity caused by the resolution jump of the spectrogram, and then the abnormal sound detection result is determined in combination with the current classification information, so as to effectively improve the accuracy of abnormal sound detection, and further improve the production efficiency and quality of products.

[0118] The abnormal sound detection device provided by this application adopts the abnormal sound detection method in the above embodiment, and can solve the technical problem of low accuracy of detecting abnormal sounds in the prior art. Compared with the prior art, the beneficial effects of the abnormal sound detection device provided by this application are the same as those of the abnormal sound detection method provided by the above embodiment, and other technical features in the abnormal sound detection device are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.

[0119] In one embodiment, the generating module 20 is further configured to perform dimension detection on the audio signal to be detected; when the dimension detection result meets the preset requirements, create a spectrogram converter according to the time-frequency resolution and the signal characteristic information; configure multi-dimensional window parameters for the spectrogram converter; and convert the audio signal to be detected based on the spectrogram converter after configuring the parameters to obtain a spectrogram to be detected.

[0120] In one embodiment, the inference module 30 is further configured to obtain a historical audio signal set, convert the historical audio signal set to obtain a historical spectrogram set; randomly select from the historical spectrogram set according to a target ratio, and determine a sample training set and a sample validation set according to the selection result; train a current efficient classification model according to the sample training set, a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid; calculate a loss value between a predicted value and a true value of the current efficient classification model based on a target loss function with weights; perform validation training on the current efficient classification model according to the sample validation set until the loss value converges to a preset value, and determine a target efficient classification model.

[0121] In one embodiment, the inference module 30 is further configured to obtain channels of sample training spectrograms in the sample training set; unify the channels of the sample training spectrograms through first-specification convolution, and perform global average pooling on the sample training spectrograms according to the unified channels; generate a channel weight matrix according to the sample training spectrograms after global average pooling; generate a spatial weight matrix according to the sample training spectrograms; generate an attention mask based on the cross-stage cooperative attention mechanism according to the spatial weight matrix and the channel weight matrix, and apply the attention mask to the sample training set; train a current efficient classification model according to the sample training set after applying the attention mask, a pre-trained feature extraction backbone network, and a progressive feature pyramid.

[0122] In one embodiment, the inference module 30 is further configured to perform max pooling on the sample training spectrograms, and perform average pooling on the sample training spectrograms after max pooling; splice the sample training spectrograms after average pooling along the unified channels; generate a spatial weight map according to second-rule convolution and the spliced sample training spectrograms; perform normalization processing on the spatial weight map to obtain a spatial weight matrix.

[0123] In one embodiment, the inference module 30 is further configured to determine multi-level features with enhanced key information according to the sampled training set after applying the attention mask and the pre-trained feature extraction backbone network; wherein the multi-level features include a first enhanced feature, a second enhanced feature, and a third enhanced feature; adjust the number of channels of the first enhanced feature according to the progressive feature pyramid; perform double upsampling on the first enhanced feature after adjusting the channel data, and element-wise add the first acquired feature to the feature obtained by convolving the second enhanced feature to obtain a first added feature; perform double upsampling on the first added feature, and add the second acquired feature to the feature obtained by convolving the first enhanced feature to obtain a second added feature; process the second added feature based on a third specification convolution, and element-wise add the processed second added feature to the first added feature to obtain a third added feature; perform downsampling on the third added feature, and add the third acquired feature to the first enhanced feature after adjusting the channel data to obtain a fourth added feature; perform global average pooling on the second added feature, the third added feature, and the fourth added feature respectively; and train the current high-efficiency classification model according to the pre-trained feature extraction backbone network and the added features after global average pooling.

[0124] In one embodiment, the inference module 30 is further configured to determine a sampled test set according to the selection result; test the current high-efficiency classification model after verification training according to the sampled test set; obtain the recall rate and precision rate of the current high-efficiency classification model after verification training according to the model test result; and determine the target high-efficiency classification model when both the recall rate and the precision rate meet the preset conditions.

[0125] The present application provides a heterophony detection device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the heterophony detection method in the first embodiment above.

[0126] Next, refer to Figure 6, which shows a schematic structural diagram of a heterophony detection device suitable for implementing the embodiments of the present application. The heterophony detection device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description), PMPs (Portable Media Player), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The shown heterophony detection device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0127] As Figure 6 shown, the heterophony detection device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the ROM (Read Only Memory) 1002 or the program loaded from the storage device 1003 into the RAM (Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the heterophony detection device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the heterophony detection device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a heterophony detection device with various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems may be implemented or had.

[0128] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. The computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the disclosed embodiments of the present application are executed.

[0129] The abnormal sound detection device provided by the present application adopts the abnormal sound detection method in the above-mentioned embodiment, and can solve the technical problem of relatively low accuracy of detecting abnormal sounds in the prior art. Compared with the prior art, the beneficial effects of the abnormal sound detection device provided by the present application are the same as those of the abnormal sound detection method provided by the above-mentioned embodiment, and other technical features in the abnormal sound detection device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.

[0130] It should be understood that each part disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0131] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0132] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the abnormal sound detection method in the above-mentioned embodiment.

[0133] The computer-readable storage medium provided by this application can, for example, be a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0134] The above computer-readable storage medium can be included in the abnormal sound detection device; or it can exist independently without being assembled into the abnormal sound detection device.

[0135] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0136] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems and methods according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0137] The modules described in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0138] The readable storage medium provided by the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned abnormal sound detection method, which can solve the technical problem of low accuracy in detecting abnormal sounds in the prior art. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the abnormal sound detection method provided by the above embodiments, and will not be elaborated here.

[0139] The above are only some embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A method for detecting abnormal sound, characterized in that, The method includes: Obtaining an audio signal to be detected in an industrial application scenario, and determining the time-frequency resolution and signal characteristic information of the audio signal to be detected; Generating a spectrogram to be detected according to the time-frequency resolution and the signal characteristic information; Inputting the spectrogram to be detected into a target efficient classification model, and obtaining the current classification information output after inference by the target efficient classification model; wherein, the target efficient classification model is trained according to a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid; Determining the abnormal sound detection result of the audio signal to be detected according to the current classification information; Before the step of inputting the spectrogram to be detected into the target efficient classification model, it further includes: Obtaining the channels of the sample training spectrograms in the sample training set; Unifying the channels of the sample training spectrograms through first-specification convolution, and performing global average pooling on the sample training spectrograms according to the unified channels; Generating a channel weight matrix according to the sample training spectrograms after global average pooling; Generating a spatial weight matrix according to the sample training spectrograms; Based on the cross-stage cooperative attention mechanism, generating an attention mask according to the spatial weight matrix and the channel weight matrix, and applying the attention mask to the sample training set; Training the current efficient classification model according to the sample training set after applying the attention mask, the pre-trained feature extraction backbone network, and the progressive feature pyramid; Determining the target efficient classification model based on the target loss function with weights according to the current efficient classification model and the sample validation set.

2. The method according to claim 1, wherein The step of generating a spectrogram to be detected according to the time-frequency resolution and the signal characteristic information includes: Performing dimension detection on the audio signal to be detected; When the dimension detection result meets the preset requirements, creating a spectrogram converter according to the time-frequency resolution and the signal characteristic information; Configuring multi-dimensional window parameters for the spectrogram converter; Converting the audio signal to be detected based on the spectrogram converter after configuring the parameters to obtain a spectrogram to be detected.

3. The method according to claim 1, wherein The step of determining the target efficient classification model based on the target loss function with weights according to the current efficient classification model and the sample validation set includes: Obtaining a historical audio signal set, and converting the historical audio signal set to obtain a historical spectrogram set; Randomly selecting the historical spectrogram set according to a target ratio, and determining the sample validation set according to the selection result; Calculating the loss value between the predicted value and the true value of the current efficient classification model based on the target loss function with weights; Performing validation training on the current efficient classification model according to the sample validation set until the loss value converges to a preset value, and determining the target efficient classification model.

4. The method according to claim 1, wherein The step of generating a spatial weight matrix according to the sample training spectrograms includes: Performing max pooling on the sample training spectrograms, and performing average pooling on the sample training spectrograms after max pooling; Concatenating the sample training spectrograms after average pooling along the unified channels; Generate a spatial weight map based on the spectrogram generated from the samples after convolution and splicing according to the second rule; Normalize the spatial weight map to obtain a spatial weight matrix.

5. The method according to claim 1, wherein The steps of training the current efficient classification model according to the sample training set with the attention mask applied, the pre-trained feature extraction backbone network, and the progressive feature pyramid include: Determine multi-level features with enhanced key information based on the sample training set with the attention mask applied and the pre-trained feature extraction backbone network; wherein, the multi-level features include a first enhanced feature, a second enhanced feature, and a third enhanced feature; Adjust the number of channels of the first enhanced feature according to the progressive feature pyramid; Perform double upsampling on the first enhanced feature after adjusting the channel data, and element-wise add the first collected feature to the feature obtained by convolving the second enhanced feature to obtain a first added feature; Perform double upsampling on the first added feature, and add the second collected feature to the feature obtained by convolving the first enhanced feature to obtain a second added feature; Process the second added feature based on a third specification convolution, and element-wise add the processed second added feature to the first added feature to obtain a third added feature; Perform downsampling on the third added feature, and add the third collected feature to the first enhanced feature after adjusting the channel data to obtain a fourth added feature; Perform global average pooling on the second added feature, the third added feature, and the fourth added feature respectively; Train the current efficient classification model according to the pre-trained feature extraction backbone network and the added features after global average pooling.

6. The method according to claim 3, characterized in that, After the step of validating and training the current efficient classification model according to the sample validation set until the loss value converges to a preset value, it further includes: Determine a sample test set according to the selection result; Test the current efficient classification model after validation training according to the sample test set; Obtain the recall rate and precision rate of the current efficient classification model after validation training according to the model test result; When both the recall rate and the precision rate meet the preset conditions, determine the target efficient classification model.

7. An abnormal sound detection device, characterized in that, The device includes: A determination module, configured to obtain a to-be-detected audio signal in an industrial application scenario, and determine the time-frequency resolution and signal characteristic information of the to-be-detected audio signal; A generation module, configured to generate a to-be-detected spectrogram according to the time-frequency resolution and the signal characteristic information; An inference module, configured to input the to-be-detected spectrogram into the target efficient classification model, and obtain the current classification information output after inference by the target efficient classification model; wherein, the target efficient classification model is trained according to a pre-trained feature extraction backbone network, a cross-stage cooperative attention mechanism, and a progressive feature pyramid; The determination module is further configured to determine the abnormal sound detection result of the to-be-detected audio signal according to the current classification information; The inference module is further configured to obtain the channels of the sample training spectrogram in the sample training set; unify the channels of the sample training spectrogram through a first specification convolution, and perform global average pooling on the sample training spectrogram according to the unified channels; generate a channel weight matrix based on the sample training spectrogram after global average pooling; generate a spatial weight matrix based on the sample training spectrogram; based on the cross-stage co-attention mechanism, generate an attention mask according to the spatial weight matrix and the channel weight matrix, and apply the attention mask to the sample training set; train the current efficient classification model according to the sample training set after applying the attention mask, the pre-trained feature extraction backbone network, and the progressive feature pyramid; determine the target efficient classification model based on the weighted objective loss function according to the current efficient classification model and the sample validation set.

8. An abnormal sound detection device, characterized in that, The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the abnormal sound detection method according to any one of claims 1 to 6.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the abnormal sound detection method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Air pipe leakage signal detection method based on non-uniform frequency spectrogram

    CN116577037A