Machine abnormal sound detection method and system based on dynamic attention mechanism

By introducing a dynamic attention mechanism and a full-dimensional dynamic convolution module in machine abnormal sound detection, the weight of the convolution operation is adaptively adjusted, and the problem of identifying abnormal sounds in a high-noise environment is solved in the existing technology, and real-time monitoring capabilities with high accuracy and low latency are achieved.

CN120108423AActive Publication Date: 2025-06-06ZHEJIANG GONGSHANG UNIVERSITY

Patent Information

Application Number
CN202510162421.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

Existing machine abnormal sound detection technology is difficult to accurately identify machine abnormal sound in high noise environments, and traditional complex models have large calculation volume and high delay, making them not suitable for real-time monitoring systems.

Method used

Using a machine anomaly sound detection method based on a dynamic attention mechanism, the frequency and time domain characteristics of machine audio are extracted by building a feature extraction network, combined with a full-dimensional dynamic convolution module and an improved CBAM attention mechanism, the weight of the convolution operation is adaptively adjusted to improve the model's abnormal sound recognition ability in complex environments.

Benefits of technology

It improves the accuracy and stability of machine abnormal sound detection, can accurately detect weak abnormal sounds in a noisy environment, supports real-time monitoring, reduces dependence on manual inspection, and optimizes resource configuration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108423A_ABST
    Figure CN120108423A_ABST
Patent Text Reader

Abstract

The invention discloses a machine abnormal sound detection method and system based on a dynamic attention mechanism, and the method comprises the steps: constructing a feature extraction network and a detection network, extracting the frequency domain-time domain features of machine audio in the feature extraction network through a log-Mel spectrogram and an improved WaveNet and TgramNet network, obtaining a feature fusion spectrogram, and carrying out the detection of the abnormal sound of a machine. A mobile FaceNet network and a detection network of an OCBAM dynamic attention mechanism are included, machine audio is obtained, data labeling of a specific machine type and a working state is carried out on the machine audio, a feature extraction network is used for carrying out feature extraction on a machine audio data set, a detection model is trained, and the detection model is used for detecting abnormal sound. According to the invention, the problem of difficulty in machine abnormal sound detection in a complex industrial environment is solved, the safety and efficiency of industrial production can be effectively improved, the production process and quality control are optimized, and the downtime and maintenance cost of equipment are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of neural network audio detection, and specifically relates to a method and system for detecting abnormal machine sounds based on a dynamic attention mechanism. Background Art

[0002] Machine abnormal sound detection technology is an important branch of the field of sound signal analysis and is widely used in scenarios such as fault prediction of industrial equipment and production environment monitoring. With the continuous development of industrial automation, the operating status of equipment has an important impact on production efficiency, product quality and factory safety. Traditional equipment monitoring methods mainly rely on manual inspections or regular maintenance, but these methods are often limited by the timeliness and coverage of manual labor, and are prone to missing subtle abnormal signals, affecting the safety and stability of production.

[0003] Abnormal sounds of machines usually come from problems such as wear, looseness, and friction of mechanical parts. These abnormal sounds often have low frequencies and weak loudness in the early stages, making them difficult to detect in time through traditional methods. The use of sound-based monitoring technology can collect, analyze, and process the sound signals generated during equipment operation in real time, identify abnormal situations in time, and thus avoid the spread of faults or accidents, and ensure the normal operation of the equipment.

[0004] At present, sound-based anomaly detection mainly relies on machine learning and deep learning technologies, especially the application of convolutional neural networks (CNN) and recurrent neural networks (RNN) in anomaly detection. Although traditional models can achieve good detection results in some scenarios, in high-noise environments, the difference between abnormal machine sounds and normal sounds is often very subtle, resulting in the accuracy of existing models being limited. Although complex network models can improve detection accuracy, they have large computational complexity and high latency, making them unsuitable for direct application in real-time monitoring systems. Summary of the invention

[0005] In order to solve the shortcomings of the prior art and achieve the purpose of improving the accuracy and stability of abnormal sound detection of machines, the present invention adopts the following technical solutions:

[0006] The machine abnormal sound detection method based on dynamic attention mechanism includes the following steps:

[0007] Step 1: Build a feature extraction network to extract the frequency domain and time domain features of the machine audio to characterize the acoustic features of the target machine audio, and fuse the frequency domain and time domain features to obtain an acoustic feature fusion spectrogram;

[0008] Step 2: Construct a machine abnormal sound detection network based on the dynamic attention mechanism, replace the convolution module in the spatial attention with the full-dimensional dynamic convolution module, the full-dimensional dynamic convolution module dynamically generates kernel weights, progressively adjusts the weights of the convolution operation, and provides adaptive convolution operations according to the different dimensional characteristics of the features. The adaptive function enables the full-dimensional dynamic convolution module to capture more complex and changeable patterns in the data, so that the model can adaptively extract contextual information and further improve the recognition ability of different types of abnormal sounds; the traditional spatial attention module uses a standard convolution layer to generate a spatial attention map, while the present invention uses ODConv to generate spatial attention weights. Integrating ODConv into CBAM can enhance the module's ability to adapt to different input features, thereby improving the attention mechanism. By dynamically adjusting the convolution kernel according to the input features, the modified CBAM can capture more subtle spatial and channel dependencies;

[0009] Step 3: Obtain machine audio data and annotate it;

[0010] Step 4: Based on the annotated machine audio data, frequency domain and time domain features are extracted through the constructed feature extraction network, and the constructed machine abnormal sound detection network is trained;

[0011] Step 5: The machine audio to be detected is passed through the feature extraction network and the trained machine abnormal sound detection network to obtain the final detection result.

[0012] Furthermore, the step 1 comprises the following steps:

[0013] Step 1.1: Use the spectrogram to extract the frequency domain features of the machine audio features; extract the time domain features with periodic and pattern features; extract the time domain features with a long time span;

[0014] Step 1.2: Use the concatenation strategy to fuse the extracted frequency domain and time domain features to integrate the multi-dimensional information of the machine audio signal and provide a more comprehensive feature representation for subsequent anomaly detection;

[0015] The feature fusion spectrum is expressed as follows:

[0016] F TWM =Concat(F Mel ,F Wave ,F Tgram )

[0017] Among them, F TWM Represents the fused feature spectrum, F Mel represents the frequency domain features extracted from the log-Mel spectrum, F Wave represents the time domain features extracted by the improved WaveNet network, F TgramRepresents the time domain features extracted by the TgramNet network.

[0018] Furthermore, in step 1.1, the original waveform is pre-emphasized to enhance high-frequency information, and then the audio signal is framed, windowed, and short-time Fourier transform (STFT) is performed, and a Mel filter group is applied on this basis to obtain a log-Mel spectrogram frequency domain representation, which can effectively capture the spectral features in the audio signal; the log-Mel spectrogram is obtained as follows:

[0019]

[0020] Among them, F Mel represents the log-mel spectrogram, x represents the input machine audio, represents Mel filter bank and STFT represents short-time Fourier transform.

[0021] Furthermore, in step 1.1, a neural network with a depthwise convolution structure is used to extract time domain features with periodic and pattern characteristics, especially for processing periodic signals in machine audio; the improved WaveNet network not only meets the lightweight requirements, but also can effectively capture the rapid changes and temporal changes of sound signals, and is suitable for high-dimensional time series data; the features derived by the improved WaveNet network can be expressed as:

[0022] F Wave =WN(x)

[0023] Among them, F Wave represents the features derived from the WaveNet network, the function WN(·) represents the improved WaveNet model for generating latent feature representations with causal relationships, and x represents the input audio of the model.

[0024] Furthermore, in the step 1.1, a neural network is constructed, including a one-dimensional large kernel convolution (Large Kernel Convolution) and a group of convolutional neural network blocks. The neural network block includes a layer normalization (LayerNormalization), a LeakyReLU activation function and a one-dimensional convolution. The large kernel convolution is used as the initial processing layer to capture a wide range of temporal context information. At the same time, the robustness of the features is enhanced by layer normalization and LeakyReLU activation function; the features derived by the TgramNet network can be expressed as:

[0025] F Tgram =TN(x)

[0026] Among them, F Tgramrepresents the features derived from the TgramNet network, the function TN(·) represents the TgramNet model, which is used to compensate for the missing abnormal information in the log-Mel spectrogram, and x represents the input audio of the model.

[0027] Furthermore, in step 2, the formula for dynamic attention mechanism processing is as follows:

[0028]

[0029] Among them, x represents the joint input feature map of machine audio frequency domain and time domain features, y represents the output feature map after being processed by the dynamic attention mechanism, and α wi Indicates w i The scalar attention weight of the kernel, α ci ∈R Cin , α fi ∈R Cout , α si ∈R k×k represents the attention weights calculated along the input dimension, output dimension, and spatial dimension, * represents the multiplication operation, and m is the index parameter, which represents the mth weight matrix and its corresponding attention mechanism. Spatial attention si assigns different weights to each spatial position of the convolution kernel to enhance the ability to capture spatial information; input channel attention ci assigns different weights to each input channel to enhance the ability to distinguish input features; output channel attention fi assigns different weights to each output channel to enhance the ability to express output features; convolution kernel attention w i Assign a scalar weight to the entire convolution kernel to adjust the global importance of the convolution kernel.

[0030] Further, the dynamic attention mechanism in step 2 includes channel attention and spatial attention; the channel attention first performs maximum pooling and average pooling operations on the features fused in step 1, and then performs addition operations after passing through the convolution layer, obtains the output features of the channel attention after the activation function, multiplies the output features of the channel attention with the fused features, and uses the features after the multiplication operation as the input of the spatial attention; the spatial attention performs maximum pooling and average pooling operations on the input features, and then passes through the full-dimensional dynamic convolution module, and obtains the output features of the spatial attention after the activation function, multiplies the output features of the spatial attention with the features after the multiplication operation, and obtains the output features of the dynamic attention mechanism for the detection of the machine abnormal sound detection network. The channel attention mechanism assigns different weights to different feature channels in the channel dimension, thereby emphasizing the contribution to the key features; while the spatial attention mechanism performs weighting according to the spatial structure of the feature map, focusing on the feature expression of important areas. Through the layer-by-layer superposition of channel and spatial attention, the model's ability to express key features can be effectively enhanced, especially when processing multidimensional features, the accuracy of feature extraction can be improved.

[0031] Furthermore, the machine abnormal sound detection network in step 2 includes, in sequence, a first convolution, a depthwise separable convolution, batch normalization, a second convolution, the dynamic attention mechanism, a linear global depthwise convolution, and a linear convolution.

[0032] Furthermore, the machine abnormal sound detection network in step 2 adopts the M-Arcface loss function to dynamically adjust the angle margin according to the complexity or uncertainty of the sample, better control the intra-class compactness and inter-class separation, introduce the margin adjustment module, and assign appropriate margin values ​​to different categories or samples through the attention mechanism. The M-Arcface loss function formula is as follows:

[0033] cos(θ+m)=cos(θ)cos(m)-sin(θ)sin(m)

[0034] Among them, m represents the dynamic adjustment of the angle margin, which is used to introduce additional intervals in the angle space to enhance the discriminability of the features, and θ is the angle vector representation.

[0035] A machine abnormal sound detection system based on a dynamic attention mechanism includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the machine abnormal sound detection method based on the dynamic attention mechanism.

[0036] The advantages and beneficial effects of the present invention are:

[0037] The present invention extracts the frequency domain-time domain features of machine audio through log-Mel spectrogram, improved WaveNet and TgramNet networks, obtains feature fusion spectrogram, effectively improves the feature information performance of machine working audio, and provides the subsequent abnormality detection model with richer features of machine audio information; through the construction of machine abnormal sound detection network, a dynamic attention mechanism is introduced to automatically adjust the model's attention to different audio features, so that the model can still accurately detect weak abnormal sounds in complex environments, improve the accuracy of abnormality detection, and solve the problem that subtle abnormal sounds are difficult to identify in the prior art, especially in noisy industrial environments, abnormal sounds can still be effectively extracted and detected; by recording different working states of different types of machines to obtain machine audio, data annotation and data preprocessing are performed on the obtained machine audio, and a machine audio data set is constructed; the model adopts a lightweight design, can efficiently process audio data, supports real-time abnormal sound monitoring, and is particularly suitable for real-time monitoring in industrial production. The present invention can adapt to different types of machines and changing working environments. By adaptively learning the sound characteristics under different machine states, the universality and adaptability of the model are improved. By improving the accuracy and real-time performance of abnormal sound detection, the downtime of equipment failures is reduced, maintenance costs and production losses are reduced, while the dependence on manual inspections is reduced and resource allocation is optimized. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a method flow chart of an embodiment of the present invention.

[0039] Figure 2 Schematic diagram of the structure of the feature extraction network in the embodiment of the present invention.

[0040] Figure 3 It is a structural diagram of the OCBAM dynamic attention mechanism module in an embodiment of the present invention.

[0041] Figure 4 It is a schematic diagram of the structure of OConv full-dimensional dynamic convolution in an embodiment of the present invention.

[0042] Figure 5 Schematic diagram of the structure of the OCBAM-MFN network in an embodiment of the present invention. DETAILED DESCRIPTION

[0043] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.

[0044] In the application fields of industrial automation, intelligent manufacturing, equipment monitoring and intelligent maintenance, the abnormal sound detection technology of machines has received extensive attention. In recent years, with the rapid development of deep learning technology, deep learning models have emerged that use neural networks to extract key features from audio signals to achieve abnormal sound detection. However, the current deep learning models still have certain limitations in dealing with complex noise environments, nonlinear sound features and weak abnormal signals, resulting in a decrease in detection accuracy and robustness, especially in high noise environments and variable working conditions.

[0045] In view of the above problems, the machine abnormal sound detection method based on dynamic attention mechanism proposed in this application optimizes the feature extraction network structure, introduces a multi-scale dynamic attention mechanism, and enhances the model's ability to capture fine-grained features of audio signals, which is conducive to improving the accuracy and stability of abnormal sound detection and supporting the efficient deployment of real-time monitoring tasks, such as Figure 1 As shown, the specific steps include:

[0046] Step 1: Build a feature extraction network to extract the frequency domain and time domain features of the machine audio to characterize the acoustic features of the target machine audio, and fuse the frequency domain and time domain features to obtain an acoustic feature fusion spectrogram, including the following steps:

[0047] Step 1.1: Use the log-Mel spectrogram to extract the frequency domain features of the machine audio features; perform pre-emphasis on the original waveform to enhance the high-frequency information, then divide the audio signal into frames, add windows, and perform short-time Fourier transform (STFT), and apply the Mel filter group on this basis to obtain the log-Mel spectrogram frequency domain representation, which can effectively capture the spectral features in the audio signal. The log-Mel spectrogram can be obtained as follows:

[0048]

[0049] Among them, F Mel represents the log-mel spectrogram, x represents the input machine audio, represents Mel filter bank and STFT represents short-time Fourier transform.

[0050] The improved WaveNet network meets the requirements of lightweight. It uses a depthwise convolution structure and can extract time domain features with periodic and pattern characteristics. It is particularly suitable for processing periodic signals in machine audio. The improved WaveNet network not only meets the requirements of lightweight, but also can effectively capture the rapid changes and time series changes of sound signals. It is suitable for high-dimensional time series data. The features derived by the improved WaveNet network can be expressed as:

[0051] F Wave =WN(x)

[0052] Among them, F Wave represents the features derived from the WaveNet network, the function WN(·) represents the improved WaveNet model for generating latent feature representations with causal relationships, and x represents the input audio of the model.

[0053] The TgramNet network consists of a one-dimensional large kernel convolution and three CNN blocks consisting of a layer normalization, a leaky relu and a one-dimensional convolution. The large kernel convolution is used as the initial processing layer to capture a wider range of temporal context information. At the same time, the robustness of the features is enhanced through layer normalization and leaky relu activation functions. TgramNet is suitable for extracting time domain features with long time spans in sound signals. The features derived by the TgramNet network can be expressed as:

[0054] F Tgram =TN(x)

[0055] Among them, F Tgram represents the features derived from the TgramNet network, the function TN(·) represents the TgramNet model, which is used to compensate for the missing abnormal information in the log-Mel spectrogram, and x represents the input audio of the model.

[0056] Step 1.2: Use the Concat strategy to fuse the frequency domain and time domain features extracted by the above operation, integrate the multi-dimensional information of the machine audio signal, and provide a more comprehensive feature representation for subsequent anomaly detection;

[0057] The feature fusion spectrum is expressed as follows:

[0058] F TWM =Concat(F Mel ,F Wave ,F Tgram )

[0059] Among them, F TWM Represents the fused feature spectrum, F Mel represents the frequency domain features extracted from the log-Mel spectrum, F Wave represents the time domain features extracted by the improved WaveNet network, F Tgram Represents the time domain features extracted by the TgramNet network.

[0060] Step 2: Build a machine abnormal sound detection network based on the improved dynamic attention mechanism; introduce the following on the basis of the existing MobileFaceNet (MFN) architecture: Figure 3 The full-dimensional dynamic convolution OCBAM (Omni-dimensional Dynamic Convolution Convolutional Block Attention Module) dynamic attention mechanism shown in the figure adopts the M-Arcface loss function to construct a new machine abnormal sound detection network OCBAM-MFN.

[0061] CBAM is an attention mechanism that enhances the features of standard convolutional neural networks (CNNs) by focusing on important features and suppressing unnecessary features. The CBAM block contains a hybrid attention mechanism that utilizes channel attention and spatial attention mechanisms. Given a fused feature map F∈R h×w×c , where h, w, c represent height, width and channel respectively, and f x×x Denoting the kernel size of the convolutional layer, the channel attention and spatial attention modules in CBAM can be calculated by the following equations.

[0062]

[0063] Among them, W 0 , W 1 represents the weight of the multi-layer perceptron (MLP), σ represents the Sigmoid activation function, The feature maps corresponding to the average and maximum values ​​are collected separately. The two attention modules will calculate the input feature maps in turn, and then multiply these attention maps according to the input feature maps to transform the features according to the input. Based on the above two formulas, the calculation of the CBAM module is as follows:

[0064]

[0065] in, Represent element-wise multiplication and addition, F ′ represents the result of channel attention, F ″ Represents the result of spatial attention.

[0066] The dynamic convolution module (ODConv) is introduced into the CBAM attention mechanism to replace the ordinary convolution module in the spatial attention. The ODconv structure is as follows: Figure 4 a to Figure 4d. Unlike traditional convolutions with fixed weights, ODConv dynamically generates kernel weights. This adaptive function enables ODConv to capture more complex and changeable patterns in the data. By gradually adjusting the weights of the convolution operation (including dimensions such as position, channel, filter, and kernel), it provides more adaptive convolution operations based on the different dimensional characteristics of the features. This enables the model to adaptively extract contextual information and further improve the ability to recognize different types of abnormal sounds. The calculation formula is as follows:

[0067]

[0068] Among them, x represents the joint input feature map of machine audio frequency domain and time domain features, y represents the output feature map after being processed by the dynamic attention mechanism, and α wi Indicates w i The scalar attention weight of the kernel, α ci ∈R Cin , α fi ∈R Cout , α si ∈R k×k Represents three new attention weights, calculated along the input dimension, output dimension, and spatial dimension. * represents the multiplication operation, and m is the index parameter, which represents the mth weight matrix and its corresponding attention mechanism.

[0069] In ODConv, each convolution kernel w i Dynamic adjustment is performed through four attention mechanisms:

[0070] Spatial attention si assigns different weights to each spatial position of the convolution kernel to enhance the ability to capture spatial information;

[0071] Input channel attention ci assigns different weights to each input channel to enhance the ability to distinguish input features;

[0072] Output channel attention fi assigns different weights to each output channel to enhance the expressiveness of output features;

[0073] Convolution kernel attention w i Assign a scalar weight to the entire convolution kernel to adjust the global importance of the convolution kernel.

[0074] Figure 4 a and Figure 4 In b, in order to more clearly show the structure of the convolution kernel and the branches of dynamic convolution, the superscript m is introduced to represent the mth component or sub-kernel of the i-th convolution kernel. The same is true for n, which is used to show the branching situation.

[0075] The traditional spatial attention module uses standard convolutional layers to generate spatial attention maps, while the present invention uses ODConv to generate spatial attention weights. Integrating ODConv into CBAM can enhance the module's ability to adapt to different input features, thereby improving the attention mechanism. By dynamically adjusting the convolution kernel according to the input features, the modified CBAM can capture more subtle spatial and channel dependencies.

[0076] The network structure of the machine abnormal sound detection network OCBAM-MFN is as follows Figure 5 As shown in the figure, after the 1×1 convolution of MobileFaceNet, the OCBAM dynamic attention mechanism module is introduced before the 7×7 linear global depth convolution. The dynamic attention module dynamically combines the channel attention and spatial attention mechanisms to weight different parts of the input feature map. The channel attention mechanism assigns different weights to different feature channels in the channel dimension, thereby emphasizing the contribution to key features; while the spatial attention mechanism performs weighting according to the spatial structure of the feature map, focusing on the feature expression of important areas. Through the layer-by-layer superposition of channel and spatial attention, the model's ability to express key features can be effectively enhanced, especially when processing multi-dimensional features, it can improve the accuracy of feature extraction.

[0077] The M-Arcface loss function dynamically adjusts the angle margin m according to the complexity or uncertainty of the sample to better control the intra-class compactness and inter-class separation. The margin adjustment module is introduced to assign appropriate margin values ​​to different categories or samples through the attention mechanism. M-Arcface loss formula:

[0078] cos(θ+m)=cos(θ)cos(m)-sin(θ)sin(m)

[0079] Where m represents the dynamic adjustment of the angle margin, which is used to introduce additional intervals in the angle space to enhance the discriminability of the features, and θ represents the angle vector representation.

[0080] Step 3: Obtain machine audio data and annotate it; the collected and annotated machine audio data specifically includes sound recordings of different types of machines in normal and faulty states, and the annotations cover the specific types of machines and the corresponding working states. In order to improve the generalization ability of the model, the annotation needs to cover machine audio under various working conditions and different load states as much as possible. Each audio sample should be labeled as "Normal" or "Anomaly". These data will serve as training data sets to provide effective supervision information for the model.

[0081] In an embodiment of the present invention, the collected and annotated audio data set is divided into a training data set, a verification data set, and a test data set in a ratio of 6:2:2. The verification data set and the test data set are used for verification and testing of a machine abnormal sound detection network.

[0082] Step 4: Training the model. After extracting the frequency domain and time domain features from the training data set through the constructed feature extraction network, the training data set is input into the constructed machine abnormal sound detection network to complete the training.

[0083] Specifically, the labeled training data set is input into the constructed feature extraction network to extract frequency domain and time domain features; the feature fusion spectrogram is obtained by processing the audio samples; these feature maps are used as input and passed into the machine abnormal sound detection network OCBAM-MFN network for training; during the training process, the cross entropy loss function and M-Arcface are used to optimize the model to ensure that it can make accurate classification between normal and abnormal sounds.

[0084] Step 5: Input the machine audio into the trained model to complete abnormal machine sound detection and obtain the final detection result.

[0085] Specifically, the trained machine abnormal sound detection model can be used for actual machine abnormal sound detection tasks; when performing detection, the machine audio to be detected is input into the trained model. After feature extraction and model processing, the model will output an abnormality score to indicate whether there is an abnormal sound.

[0086] By implementing the invention of the machine abnormal sound detection method based on the dynamic attention mechanism, it is possible to efficiently perform machine abnormal sound detection in a resource-constrained environment and provide high detection accuracy and versatility. This method is not only suitable for fault detection of a single device, but can also achieve efficient cross-device abnormal sound detection on multiple devices, and has broad application prospects.

[0087] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some or all of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A machine abnormal sound detection method based on dynamic attention mechanism, characterized by The steps include: Step 1: Build a feature extraction network to extract the frequency domain and time domain features of the machine audio to characterize the acoustic features of the target machine audio, and fuse the frequency domain and time domain features to obtain an acoustic feature fusion spectrogram; Step 2: Build a machine abnormal sound detection network based on the dynamic attention mechanism, and replace the convolution module in the spatial attention with the full-dimensional dynamic convolution module. The full-dimensional dynamic convolution module dynamically generates kernel weights and progressively adjusts the weights of the convolution operation, providing adaptive convolution operations according to the different dimensional characteristics of the features. Step 3: Obtain machine audio data and annotate it; Step 4: Based on the annotated machine audio data, frequency domain and time domain features are extracted through the constructed feature extraction network, and the constructed machine abnormal sound detection network is trained; Step 5: The machine audio to be detected is passed through the feature extraction network and the trained machine abnormal sound detection network to obtain the final detection result.

2. The method for detecting abnormal machine sound based on dynamic attention mechanism according to claim 1, characterized in that: The step 1 comprises the following steps: Step 1.1: Use the spectrogram to extract the frequency domain features of the machine audio features; extract the time domain features with periodicity and pattern features; extract the time domain features with a long time span; Step 1.2: Use the splicing strategy to fuse the extracted frequency domain and time domain features.

3. The method for detecting abnormal machine sound based on dynamic attention mechanism according to claim 2, characterized in that: In step 1.1, the original waveform is pre-emphasized to enhance high-frequency information, and then the audio signal is framed, windowed, and short-time Fourier transformed, and a Mel filter group is applied on this basis to obtain a log-Mel spectrogram frequency domain representation.

4. The method for detecting abnormal machine sound based on dynamic attention mechanism according to claim 2, characterized in that: In the step 1.1, a neural network with a deep separable convolutional structure is used to extract time domain features with periodic and pattern characteristics, especially for processing periodic signals in machine audio.

5. The method for detecting abnormal machine sound based on dynamic attention mechanism according to claim 2, characterized in that: In the step 1.1, a neural network is constructed, including a large kernel convolution and a set of convolutional neural network blocks, and the neural network block includes layer normalization, activation function and convolution.

6. The method for detecting abnormal machine sound based on dynamic attention mechanism according to claim 1, characterized in that: In step 2, the formula for dynamic attention mechanism processing is as follows: Among them, x represents the joint input feature map of machine audio frequency domain and time domain features, y represents the output feature map after being processed by the dynamic attention mechanism, and α wi Indicates w i The scalar attention weight of the kernel, α ci ∈R Cin , α fi ∈R Cout , α si ∈R k×k represents the attention weights calculated along the input dimension, output dimension, and spatial dimension, * represents the multiplication operation, and m is the index parameter, which represents the mth weight matrix and its corresponding attention mechanism.

7. The method for detecting abnormal machine sound based on dynamic attention mechanism according to claim 1, characterized in that: The dynamic attention mechanism in step 2 includes channel attention and spatial attention; the channel attention first performs maximum pooling and average pooling operations on the features fused in step 1, and then performs an addition operation after passing through the convolution layer, and obtains the output features of the channel attention after the activation function, and multiplies the output features of the channel attention with the fused features, and uses the features after the multiplication operation as the input of the spatial attention; the spatial attention performs maximum pooling and average pooling operations on the input features, and then passes through the full-dimensional dynamic convolution module, and obtains the output features of the spatial attention after the activation function, and multiplies the output features of the spatial attention with the features after the multiplication operation to obtain the output features of the dynamic attention mechanism, which are used for the detection of machine abnormal sound detection network.

8. The method for detecting abnormal machine sound based on dynamic attention mechanism according to claim 1, characterized in that: The machine abnormal sound detection network in step 2 includes, in sequence, a first convolution, a depthwise separable convolution, batch normalization, a second convolution, the dynamic attention mechanism, a linear global depthwise convolution, and a linear convolution.

9. The method for detecting abnormal machine sound based on dynamic attention mechanism according to claim 1, characterized in that: The machine abnormal sound detection network in step 2 adopts the M-Arcface loss function to dynamically adjust the angle margin according to the complexity or uncertainty of the sample, better control the intra-class compactness and inter-class separation, introduce the margin adjustment module, and assign appropriate margin values ​​to different categories or samples through the attention mechanism. The M-Arcface loss function formula is as follows: cos(θ+m)=cos(θ)cos(m)-sin(θ)sin(m) Among them, m represents the dynamic adjustment of the angle margin, which is used to introduce additional intervals in the angle space to enhance the discriminability of the features, and θ is the angle vector representation.

10. A machine abnormal sound detection system based on a dynamic attention mechanism, comprising a memory and one or more processors, characterized in that: The memory stores executable codes, and when the one or more processors execute the executable codes, they are used to implement the machine abnormal sound detection method based on the dynamic attention mechanism described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Transform-based slender object target detection method

    CN115546468A

  • Image semantic segmentation method based on dynamic convolution attention

    CN116129115A

  • Multi-modal feature target detection method based on dynamic convolution and attention mechanism

    CN116452937A

  • Audio and video processing method and device for handheld terminal and law enforcement recorder

    CN117456405A

  • Sound event detection method based on cross-model two-stage training

    CN117877516A

Cited By

  • Audio-based fault diagnosis method and apparatus, and electronic device

    CN115910100A

  • Audio-based fault diagnosis method, device and electronic equipment

    CN115910100B