Micromotor abnormal sound classification method and device based on multi-scale feature fusion and attention mechanism

By introducing multi-scale feature fusion and attention mechanism methods in micromotor fault diagnosis, the problems of insufficient feature information extraction and low classification accuracy in the prior art are solved, and higher fault diagnosis accuracy and robustness are achieved.

CN120048285AActive Publication Date: 2025-05-27MINZHUO ELECTRIC CO LTD

Patent Information

Application Number
CN202510067686.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-27
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The existing micromotor fault diagnosis methods based on sound signals have problems such as insufficient feature information extraction, low classification accuracy, and poor robustness to complex noise and nonlinear faults.

Method used

The micromotor differential sound classification method based on multi-scale feature fusion and attention mechanism is adopted to extract multi-level features of micromotor sound signals through multi-scale convolutional neural subnetwork and attention mechanism, and enhance the robustness of the model through channel and spatial attention mechanisms.

Benefits of technology

It significantly improves the accuracy and accuracy of micromotor fault diagnosis, enhances the robustness of complex noise and nonlinear faults, and improves the generalization ability and overfitting suppression effect of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048285A_ABST
    Figure CN120048285A_ABST
Patent Text Reader

Abstract

The invention provides a micromotor abnormal sound classification method and device based on multi-scale feature fusion and an attention mechanism. The micromotor abnormal sound classification method comprises the following steps: 1, acquiring sound signal data; 2, sound signal preprocessing; step 3, sound data feature extraction; 4, performing multi-scale feature fusion: performing multi-level feature information extraction on the comprehensive feature vector by adopting convolutional neural sub-networks of three different-scale convolution kernels to obtain a feature map; a channel attention mechanism and a space attention mechanism are integrated in each convolutional neural sub-network, the importance of each channel feature vector is weighted through the channel attention mechanism, and the region of interest of the feature vector is dynamically adjusted in the spatial dimension through the space attention mechanism; and 5, performing feature classification and judging the running state or the fault type of the current micromotor. According to the method, fusion of multi-scale convolution and an attention mechanism is introduced, the robustness of the model in a complex environment is enhanced, and the accuracy of fault classification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of motor manufacturing, and more specifically, to a method and device for classifying abnormal sounds of micro-motors based on multi-scale feature fusion and attention mechanism. Background Art

[0002] Micro-motors are widely used in many industrial and household electrical appliances, such as fans, power tools, pump equipment, etc. Due to the fact that micro-motors are prone to faults such as rotor imbalance, bearing wear, gear friction, and foreign object interference during long-term operation, these faults not only affect the operating efficiency of micro-motors, but may also cause equipment shutdown or more serious damage. Therefore, real-time monitoring of the operating state of micro-motors, timely diagnosis and early warning of faults are the keys to improving equipment reliability and reducing maintenance costs.

[0003] Currently, the fault diagnosis methods of micro-motors mainly include monitoring based on vibration signals, temperature changes, and sound signals. As a non-intrusive monitoring method, sound signals have significant advantages in the fault diagnosis of micro-motors. By collecting the sound signals of micro-motors during operation and combining signal processing and feature analysis, the fault types can be accurately identified. However, the existing fault diagnosis methods based on sound signals mostly rely on traditional signal processing techniques, such as time-domain analysis, frequency-domain analysis, feature extraction, etc., and there are problems such as insufficient extraction of feature information, low classification accuracy, and poor robustness to complex noise and non-linear faults. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a method and device for classifying abnormal sounds of micro-motors based on multi-scale feature fusion and attention mechanism. The method of the present invention introduces the fusion of multi-scale convolution and attention mechanism, enhances the robustness of the model in complex environments, and improves the accuracy and precision of fault classification at the same time.

[0005] To achieve the above purpose, the present invention is realized through the following technical solutions: A method for classifying abnormal sounds of micro-motors based on multi-scale feature fusion and attention mechanism, characterized in that it includes the following steps:

[0006] The first step, sound signal data acquisition: Collect the sound signals generated by the micro-motor under different working states; the collected sound signals are: normal operating sound signals, rotor imbalance sound signals, bearing wear sound signals, gear friction sound signals, and sound signals mixed with foreign objects;

[0007] The second step, sound signal preprocessing: Perform noise reduction processing on the collected sound signals to obtain sound data, so as to reduce the influence of environmental noise and interference;

[0008] Step 3, sound data feature extraction: Extract features from the sound data after preprocessing the sound signal; the extracted features include spectrogram, Mel-frequency cepstral coefficients MFCC, and zero-crossing rate ZCR, and they are fused through feature concatenation to form a comprehensive feature vector;

[0009] Step 4, multi-scale feature fusion: Use convolutional neural sub-networks with three different scales of convolutional kernels to extract multi-level feature information from the comprehensive feature vector to obtain feature maps; Channel attention and spatial attention mechanisms are integrated in each convolutional neural sub-network. The importance of each channel feature vector is weighted through the channel attention mechanism, and the attention area of the feature vector is dynamically adjusted in the spatial dimension through the spatial attention mechanism;

[0010] Step 5, feature classification and determination of the operating state or fault type of the current micro-motor: Flatten and regularize the feature maps to obtain a global feature vector; Use multiple layers of neurons to extract the global feature vector, and map the features to the operating mode or fault state for classification; Generate the probability distribution of each category through the Softmax activation function, and determine the operating state or fault type of the current micro-motor by selecting the maximum probability value, realizing the classification of abnormal sounds of the micro-motor; Among them, the operating mode is the normal state, and the fault states are rotor imbalance state, bearing wear state, gear friction state, and state with foreign objects mixed in.

[0011] In Step 2, the VSNLMS algorithm is used to denoise the collected sound signal:

[0012] (1) Gradient term calculation: At each sampling point k, calculate the gradient term h i (k), which represents the product of the current error signal and the input signal and is used for subsequent step size adjustment;

[0013] h i (k) = d(k)·y(k - i)

[0014] where h i (k) is the gradient term at the current sampling point, representing the product of the error and the input signal, and is used for the dynamic update of the step size; d(k) represents the current error signal, and y(k - i) is the delayed sample of the input signal at the kth moment;

[0015] (2) Update of the step size factor: The step size factor u(k) is dynamically adjusted at each sampling point to adapt to the current error change, and the formula is as follows:

[0016] u(k + 1) = u(k) + α·h i (k)·h i (k - 1)

[0017] Among them, u(k) is the step size factor that controls the amplitude of weight update and is dynamically adjusted within each sampling period; α is the adjustment rate factor used to control the rate of step size update.

[0018] (3) Calculation of the upper limit of the step size factor: To prevent the filter from becoming unstable due to an overly large step size factor, VSNLMS sets an upper limit q for the step size. max , and the calculation formula is:

[0019]

[0020] Among them, y T (k)·y(k) represents the power of the input signal at the k-th moment.

[0021] (4) Range limitation of the step size factor: The value of the step size factor u(k) is restricted within the range of [q min , q max to ensure stability; the specific constraint conditions are as follows:

[0022]

[0023] q max and q min are the upper and lower limits of the step size factor, ensuring the stability of the step size under different states;

[0024] (5) Weight update: Finally, the adjusted step size factor u(k) is used to update the filter weights to reduce the error and make the filter adapt to the changes in the input signal:

[0025] z i (k + 1) = z i (k) + 2·u(k)·h i (k)

[0026] z i (k) is the weight coefficient of the filter, which is updated according to the step size in each sampling period to minimize the error; z i (k + 1) represents the updated value of the i-th weight coefficient at the k + 1 moment.

[0027] The present invention uses the VSNLMS algorithm to denoise the original sound signal, reduce the influence of environmental noise and interference, and ensure the signal quality input to the feature extraction module. The VSNLMS algorithm is a variable step-size adaptive filtering method used to solve the contradiction between the steady-state error and the convergence speed of the LMS and NLMS algorithms. VSNLMS dynamically adjusts the step-size factor so that the filter can balance fast convergence and low steady-state error in different states. In the VSNLMS algorithm, the step-size factor u(k) is dynamically adjusted according to the power of the current input signal and the error signal, and the maximum value q of the step-size is calculated in each sampling period max , to prevent instability caused by an overly large step-size.

[0028] In the third step, the short-time Fourier transform STFT is used to generate spectrogram features to capture the frequency distribution of the sound data for detecting low-frequency and high-frequency components:

[0029]

[0030] where S(t,f) represents the spectral amplitude of the signal at time t and frequency f, representing the value of the spectrogram; x[n] is the input discrete-time signal; w[n-t] is the window function centered at time t, and the window function is used to intercept the local segment of the signal and limit the analysis range at the current moment; e -j2πfn is the kernel function of the Fourier transform that converts the time-domain signal into the frequency-domain signal, where j is the imaginary unit, 2πf is the angular frequency.

[0031] In the third step, the Mel-frequency cepstral coefficients MFCC are extracted to capture the timbre information of the sound data, which helps to distinguish different fault types:

[0032] First, the Fourier transform is performed on the sound data to obtain the spectrum; secondly, the power spectrum is converted to the Mel scale through the Mel filter bank to better conform to the auditory characteristics of the human ear; where the conversion formula for the Mel scale is:

[0033]

[0034] The Mel filter bank consists of a group of triangular filters. Each Mel filter extracts the energy in a different frequency range, and then the logarithm is taken for the output energy E of each filter m , and the discrete cosine transform DCT is performed on the result to obtain the Mel cepstral coefficients MFCC:

[0035]

[0036] where MFCC c is the c-th MFCC coefficient; M is the number of Mel filters; E mis the energy of the m-th Mel filter; C represents the number of MFCC coefficients extracted.

[0037] In the third step, the zero-crossing rate ZCR is extracted to capture the roughness or frequency characteristics of the sound data. The calculation formula of the zero-crossing rate ZCR for each frame is as follows:

[0038]

[0039] where ZCR is the zero-crossing rate of this frame, representing the number of times the signal crosses the zero point within this frame; N is the number of sampling points within this frame, that is, the frame length; x[n] is the value of the discrete-time signal at the n-th sampling point; sgn(x[n - 1]) is the sign function, representing the positive or negative nature of the signal.

[0040] In the fourth step, the convolutional neural sub-networks with three different scale convolutional kernels are respectively the convolutional neural sub-network one, the convolutional neural sub-network two, and the convolutional neural sub-network three;

[0041] Among them, the convolutional neural sub-network one uses a larger convolutional kernel to capture low-frequency components and includes three convolutional layers; the comprehensive feature vector first passes through 24 5×5 convolutional kernels of the first convolutional layer, then performs 4×2 max pooling and uses the ReLU activation function; then it passes through 48 5×5 convolutional kernels of the second convolutional layer, undergoes max pooling and ReLU activation; finally, it passes through 48 5×5 convolutional kernels of the third convolutional layer, and after pooling and ReLU activation, it outputs the feature map;

[0042] The convolutional neural sub-network two uses medium-sized convolutional kernels to focus on intermediate-frequency features and consists of two convolutional layers in total; the comprehensive feature vector first passes through 32 3×3 convolutional kernels of the first convolutional layer, uses L2 regularization to reduce the risk of overfitting, and undergoes 4×2 max pooling and ReLU activation; then it passes through 48 3×3 convolutional kernels of the second convolutional layer, and after ReLU activation and max pooling operations, it outputs the feature map;

[0043] The convolutional neural sub-network three uses smaller convolutional kernels to capture high-frequency components and contains eight convolutional layers; the comprehensive feature vector passes through 32 3×3 convolutional kernels of the first two convolutional layers, uses ReLU activation and 2×2 pooling; then it passes through 64 3×3 convolutional kernels of the next four convolutional layers and continues to use ReLU activation; finally, it passes through 128 3×3 convolutional kernels of the last two convolutional layers, and after ReLU activation, it outputs the feature map;

[0044] The convolutional neural sub-networks with three different scale convolutional kernels are used to extract multi-level feature information from the comprehensive feature vector.

[0045] In each sub-network of the present invention, convolutional kernels of different scales are used to capture multi-level feature information in the abnormal sound signals of the micro-motor, thereby improving the classification accuracy. Additionally, a channel attention mechanism and a spatial attention mechanism are integrated into each sub-network. By dynamically adjusting the weights of the feature maps, the model can focus on more important feature regions, reduce the interference of irrelevant noise, and thus improve the classification accuracy. Among them, the channel attention mechanism weights the importance of each channel, and the spatial attention mechanism dynamically adjusts the attention region of the feature map in the spatial dimension.

[0046] In the fifth step, obtaining the global feature vector by flattening and regularizing the feature map means that the feature maps output by Convolutional Neural Sub-network One, Convolutional Neural Sub-network Two, and Convolutional Neural Sub-network Three are flattened, converted into feature vectors, and input into the fully connected layer. The fully connected layer uses the Dropout technique for regularization to reduce the risk of overfitting and improve the generalization ability of the model; a global feature vector is obtained through the aggregation of the fully connected layer.

[0047] In the fifth step, using multiple layers of neurons to extract the global feature vector and mapping the features to the operating mode or fault state for classification; generating the probability distribution of each category through the Softmax activation function and determining the operating state or fault type of the current micro-motor by selecting the maximum probability value to achieve abnormal sound classification of the micro-motor means that the fully connected layer uses multiple layers of neurons to calculate the global feature vector and finally outputs a five-dimensional vector, where each dimension represents the predicted value for the operating mode or fault category; the last layer of the fully connected layer uses the Softmax activation function to convert the output five-dimensional vector into the probability distribution of each category, and the sum is 1; the probability value of each category represents the possibility of that operating mode or fault category; by selecting the maximum probability value in the Softmax output, the operating state or fault type of the current micro-motor is determined to achieve abnormal sound classification of the micro-motor.

[0048] In the fifth step, after determining the operating state or fault type of the current micro-motor, a confusion matrix and a t-SNE graph are output, and the probability values of the operating state or fault type are displayed to visualize the classification results.

[0049] A micro-motor abnormal sound classification device based on multi-scale feature fusion and attention mechanism, characterized by comprising:

[0050] An acoustic sensor acquisition module for acquiring the sound signals generated by the micro-motor under different working states; the acquired sound signals are: normal operation sound signals, rotor imbalance sound signals, bearing wear sound signals, gear friction sound signals, and sound signals mixed with foreign objects;

[0051] Signal preprocessing module: used to perform noise reduction on the collected sound signals to obtain sound data, so as to reduce the influence of environmental noise and interference;

[0052] Feature extraction module: used to extract features from the sound data after preprocessing of the sound signals; the extracted features include spectrogram, Mel Frequency Cepstral Coefficients (MFCC) and Zero Crossing Rate (ZCR), and are fused through feature splicing to form a comprehensive feature vector;

[0053] Multi-scale convolutional neural network module; used to extract multi-level feature information from the comprehensive feature vector by using convolutional neural sub-networks with three different scales of convolutional kernels to obtain feature maps; in each convolutional neural sub-network, a channel attention mechanism and a spatial attention mechanism are integrated, the importance of each channel feature vector is weighted through the channel attention mechanism, and the attention area of the feature vector is dynamically adjusted in the spatial dimension through the spatial attention mechanism;

[0054] Classification and output module: used to obtain a global feature vector by flattening and regularizing the feature maps; extract the global feature vector by using multiple layers of neurons, and map the features to the operating mode or fault state for classification; generate the probability distribution of each category through the Softmax activation function, and determine the operating state or fault type of the current micro-motor by selecting the maximum probability value, so as to realize the abnormal sound classification of the micro-motor; among them, the operating mode is the normal state, and the fault states are rotor imbalance state, bearing wear state, gear friction state and foreign object mixing state.

[0055] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0056] 1. Improve the ability to extract multi-scale features:

[0057] By introducing multiple convolutional neural sub-networks with different sizes of convolutional kernels, it is possible to simultaneously extract low-frequency, medium-frequency and high-frequency features in the sound signals of the micro-motor at multiple scales. This technical effect enables the present invention to comprehensively analyze and understand the different frequency components of complex fault sound signals or operating mode sound signals, and significantly improves the accuracy of micro-motor fault diagnosis. Compared with traditional methods, a single-scale convolutional kernel cannot effectively capture the features of each frequency band of the signal, resulting in limitations in feature extraction.

[0058] 2. Enhance noise robustness:

[0059] The present invention integrates a channel attention mechanism and a spatial attention mechanism, which can dynamically adjust the feature weights of each channel and spatial region. This mechanism can focus on the most discriminative feature regions, thereby effectively suppressing the interference of irrelevant noise and reducing the impact of background noise on fault diagnosis. Especially in the case of more noise components in the signal, the attention mechanism can help the model enhance the influence of important features and improve the sensitivity and accuracy of micro-motor fault diagnosis.

[0060] 3. Improved the classification accuracy and robustness of the model:

[0061] By combining multi-scale feature extraction with the attention mechanism, the present invention effectively improves the classification accuracy of the classification model when facing various different fault modes. Different fault types such as normal operation, rotor imbalance, bearing wear, gear friction, and the presence of foreign objects are more accurately distinguished by the model. In addition, the attention mechanism also enhances the model's focusing ability on the fault features of micro-motors, making the classification results more robust under various environmental conditions.

[0062] 4. Improved the generalization ability of the model and suppression of overfitting:

[0063] After the output of the multi-scale convolutional neural network, the Dropout technique is added for regularization, effectively reducing the risk of overfitting during the training process of the model and further improving the generalization ability of the model. This technical means enables the model to better adapt to various different working environments of micro-motors and improves the application performance on different data sets. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 is a flowchart of the method for classifying abnormal sounds of micro-motors based on multi-scale feature fusion and attention mechanism of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0065] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments.

[0066] Embodiment 1

[0067] As Figure 1 shown, the method for classifying abnormal sounds of micro-motors based on multi-scale feature fusion and attention mechanism of the present invention includes the following steps:

[0068] The first step, sound signal data acquisition: Using a combination of a high-precision microphone, an audio acquisition card and a computer, the sound signals generated by the micro-motor under different working states are collected in real time in a soundproof room; the collected sound signals are: normal operation sound signals, rotor imbalance sound signals, bearing wear sound signals, gear friction sound signals, and sound signals with foreign objects mixed in.

[0069] Step 2, preprocessing of the sound signal: Denoise the collected sound signal to obtain sound data, so as to reduce the influence of environmental noise and interference;

[0070] Step 3, feature extraction of the sound data: Extract features from the sound data after preprocessing the sound signal; The extracted features include spectrogram, Mel-frequency cepstral coefficients MFCC and zero-crossing rate ZCR, and are fused through feature splicing to form a comprehensive feature vector;

[0071] Step 4, multi-scale feature fusion: Use convolutional neural subnetworks with three different scales of convolutional kernels to extract multi-level feature information from the comprehensive feature vector to obtain feature maps; In each convolutional neural subnetwork, a channel attention mechanism and a spatial attention mechanism are integrated. The importance of each channel feature vector is weighted through the channel attention mechanism, and the attention area of the feature vector is dynamically adjusted in the spatial dimension through the spatial attention mechanism;

[0072] Step 5, feature classification and judgment of the current operating state or fault type of the micro-motor: Flatten and regularize the feature map to obtain a global feature vector; Use multiple layers of neurons to extract the global feature vector, and map the features to the operating mode or fault state for classification; Generate the probability distribution of each category through the Softmax activation function, and determine the current operating state or fault type of the micro-motor by selecting the maximum probability value, realizing the classification of abnormal sounds of the micro-motor; Among them, the operating mode is the normal state, and the fault states are rotor imbalance state, bearing wear state, gear friction state, and state with foreign objects mixed in. Finally, output the confusion matrix and t-SNE graph, and display the probability values of the operating state or fault type to realize the visualization of the classification results.

[0073] Specifically, in Step 2, the VSNLMS algorithm is used to denoise the collected sound signal:

[0074] (1) Gradient term calculation: At each sampling point k, calculate the gradient term h i (k), which represents the product of the current error signal and the input signal, and is used for subsequent step size adjustment;

[0075] h i (k) = d(k)·y(k - i)

[0076] where h i (k) is the gradient term at the current sampling point, representing the product of the error and the input signal, and is used for the dynamic update of the step size; d(k) represents the current error signal, and y(k - i) is the delayed sample of the input signal at the kth moment;

[0077] (2) Update of the step size factor: The step size factor u(k) is dynamically adjusted at each sampling point to adapt to the current error change, and the formula is as follows:

[0078] u(k + 1)=u(k)+α·h i (k)·h i (k - 1)

[0079] Wherein, u(k) is the step size factor, which controls the amplitude of weight update and is dynamically adjusted within each sampling period; α is the adjustment rate factor, which is used to control the rate of step size update;

[0080] (3) Calculation of the upper limit of the step size factor: To prevent the filter from becoming unstable due to an overly large step size factor, VSNLMS sets an upper limit q for the step size max , and the calculation formula is:

[0081]

[0082] Wherein, y T (k)·y(k) represents the power of the input signal at the k-th moment;

[0083] (4) Range limitation of the step size factor: The value of the step size factor u(k) is restricted within the range of [q min , q max to ensure stability; the specific constraint conditions are as follows:

[0084]

[0085] q max and q min are the upper and lower limits of the step size factor, ensuring the stability of the step size under different states;

[0086] (5) Weight update: Finally, the adjusted step size factor u(k) is used to update the filter weights to reduce the error and make the filter adapt to the changes in the input signal:

[0087] z i (k + 1)=z i (k)+2·u(k)·h i (k)

[0088] z i (k) is the weight coefficient of the filter, which is updated according to the step size in each sampling period to minimize the error; z i (k + 1) represents the updated value of the i-th weight coefficient at the k + 1 moment.

[0089] In the third step, spectrogram features are generated through the short-time Fourier transform STFT to capture the frequency distribution of the sound data for detecting low-frequency and high-frequency components:

[0090]

[0091] Among them, S(t, f) represents the spectral amplitude of the signal at time t and frequency f, representing the value of the spectrogram; x[n] is the input discrete-time signal; w[n - t] is the window function, whose center is at time t. The window function is used to intercept the local segment of the signal, defining the analysis range at the current moment; e -j2πfn is the kernel function of the Fourier transform, which converts the time-domain signal into the frequency-domain signal, where j is the imaginary unit, 2πf is the angular frequency.

[0092] In the third step, the Mel Frequency Cepstral Coefficients MFCC are extracted to capture the timbre information of the sound data, which helps to distinguish different fault types:

[0093] First, the Fourier transform is performed on the sound data to obtain the spectrum; secondly, the power spectrum is converted to the Mel scale through the Mel filter bank to be more in line with the auditory characteristics of the human ear. Among them, the conversion formula of the Mel scale is:

[0094]

[0095] The Mel filter bank is composed of a group of triangular filters. Each Mel filter extracts the energy in a different frequency range, and then takes the logarithm of the output energy E of each filter m and performs the Discrete Cosine Transform DCT on the result to obtain the Mel cepstral coefficients MFCC:

[0096]

[0097] where MFCC c is the c-th MFCC coefficient; M is the number of Mel filters; E m is the energy of the m-th Mel filter; C represents the number of MFCC coefficients extracted.

[0098] In the third step, the Zero Crossing Rate ZCR is extracted to capture the roughness or frequency characteristics of the sound data. The calculation formula of the Zero Crossing Rate ZCR for each frame is as follows:

[0099]

[0100] Among them, ZCR is the zero crossing rate of this frame, representing the number of times the signal crosses the zero point within this frame; N is the number of sampling points within this frame, that is, the frame length; x[n] is the value of the discrete-time signal at the n-th sampling point; sgn(x[n - 1]) is the sign function, representing the positive and negative of the signal.

[0101] In the fourth step, the convolutional neural subnetworks with three different-scale convolutional kernels are respectively the convolutional neural subnetwork one, the convolutional neural subnetwork two, and the convolutional neural subnetwork three;

[0102] Among them, the first convolutional neural sub-network uses larger convolutional kernels to capture low-frequency components and includes three convolutional layers. The comprehensive feature vector first passes through 24 5×5 convolutional kernels of the first convolutional layer, then performs 4×2 max pooling, and uses the ReLU activation function. Then it passes through 48 5×5 convolutional kernels of the second convolutional layer, undergoes max pooling and ReLU activation. Finally, it passes through 48 5×5 convolutional kernels of the third convolutional layer, and after pooling and ReLU activation, it outputs a feature map.

[0103] The second convolutional neural sub-network uses medium-sized convolutional kernels to focus on mid-frequency features and consists of two convolutional layers. The comprehensive feature vector first passes through 32 3×3 convolutional kernels of the first convolutional layer, uses L2 regularization to reduce the risk of overfitting, and undergoes 4×2 max pooling and ReLU activation. Then it passes through 48 3×3 convolutional kernels of the second convolutional layer, and after ReLU activation and max pooling operations, it outputs a feature map.

[0104] The third convolutional neural sub-network uses smaller convolutional kernels to capture high-frequency components and includes eight convolutional layers. The comprehensive feature vector passes through 32 3×3 convolutional kernels of the first two convolutional layers, uses ReLU activation and 2×2 pooling. Then it passes through 64 3×3 convolutional kernels of the next four convolutional layers and continues to use ReLU activation. Finally, it passes through 128 3×3 convolutional kernels of the last two convolutional layers, and after ReLU activation, it outputs a feature map.

[0105] Multi-level feature information extraction of the comprehensive feature vector is achieved through convolutional neural sub-networks with three different scales of convolutional kernels.

[0106] In the fifth step, obtaining the global feature vector from the feature map through flattening and regularization means that the feature maps respectively output by the first convolutional neural sub-network, the second convolutional neural sub-network, and the third convolutional neural sub-network are flattened, converted into feature vectors and input into the fully connected layer. The fully connected layer uses the Dropout technique for regularization to reduce the risk of overfitting and improve the generalization ability of the model. After being aggregated by the fully connected layer, a global feature vector is obtained.

[0107] In the fifth step, multi-layer neurons are used to extract the global feature vector, and the features are mapped to the operating mode or fault status for classification; the probability distribution of each category is generated by the Softmax activation function, and the current operating state or fault type of the micromotor is determined by selecting the maximum probability value. The classification of micromotor abnormal sound means: the fully connected layer uses multi-layer neurons to calculate the global feature vector, and finally outputs a five-dimensional vector, in which each dimension represents the predicted value of the operating mode or the fault category; the last layer of the fully connected layer uses the Softmax activation function to convert the output five-dimensional vector into the probability distribution of each category, and the sum is 1; the probability value of each category represents the possibility of the operating mode or the fault category; by selecting the maximum probability value in the Softmax output, the operating state or fault type of the current micromotor is determined to achieve the classification of micromotor abnormal sound.

[0108] The advantages of the micromotor abnormal sound classification method based on multi-scale feature fusion and attention mechanism of the present invention are:

[0109] 1. Improved the ability to extract multi-scale features:

[0110] By introducing multiple convolutional neural subnetworks with convolution kernels of different sizes, the low-frequency, medium-frequency and high-frequency features in the micromotor sound signal can be extracted simultaneously at multiple scales. This technical effect enables the present invention to comprehensively analyze and understand the different frequency components of complex fault sound signals or operating mode sound signals, significantly improving the accuracy of micromotor fault diagnosis. Compared with traditional methods, a single-scale convolution kernel cannot effectively capture the characteristics of each frequency band of the signal, resulting in limitations in feature extraction.

[0111] 2. Enhanced noise robustness:

[0112] The present invention integrates the channel attention mechanism and the spatial attention mechanism, and can dynamically adjust the feature weights of each channel and spatial region. This mechanism can focus on the most discriminative feature area, thereby effectively suppressing the interference of irrelevant noise and reducing the impact of background noise on fault diagnosis. Especially when there are many noise components in the signal, the attention mechanism can help the model enhance the impact of important features and improve the sensitivity and accuracy of micromotor faults.

[0113] 3. Improved the classification accuracy and robustness of the model:

[0114] By combining multi-scale feature extraction with an attention mechanism, the present invention effectively improves the classification accuracy of the classification model when facing various different fault modes. Different fault types such as normal operation, rotor imbalance, bearing wear, gear friction, and the presence of foreign objects are more accurately distinguished by the model. In addition, the attention mechanism enhances the model's focusing ability on the fault features of the micro-motor, thus making the classification results more robust under various environmental conditions.

[0115] 4. Improved generalization ability and overfitting suppression of the model:

[0116] After the output of the multi-scale convolutional neural network, the Dropout technique is added for regularization, effectively reducing the risk of overfitting during the training process of the model and further improving the generalization ability of the model. This technical means enables the model to better adapt to various different working environments of micro-motors and improves the application performance on different data sets.

[0117] Embodiment 2

[0118] The micro-motor abnormal sound classification device based on multi-scale feature fusion and attention mechanism of the present invention includes:

[0119] An acoustic sensor acquisition module for acquiring the sound signals generated by the micro-motor under different working states; the acquired sound signals are: normal operation sound signals, rotor imbalance sound signals, bearing wear sound signals, gear friction sound signals, and sound signals with foreign objects mixed in.

[0120] A signal preprocessing module: for performing noise reduction processing on the acquired sound signals to obtain sound data, so as to reduce the influence of environmental noise and interference.

[0121] A feature extraction module: for extracting features from the sound data after preprocessing of the sound signals; the extracted features include spectrograms, Mel-frequency cepstral coefficients MFCC, and zero-crossing rates ZCR, and are fused through feature splicing to form a comprehensive feature vector.

[0122] A multi-scale convolutional neural network module; for using convolutional neural subnetworks with three different scales of convolutional kernels to perform multi-level feature information extraction on the comprehensive feature vector to obtain feature maps; in each convolutional neural subnetwork, a channel attention mechanism and a spatial attention mechanism are integrated, the importance of each channel feature vector is weighted through the channel attention mechanism, and the attention area of the feature vector is dynamically adjusted in the spatial dimension through the spatial attention mechanism.

[0123] Classification and output module: used to obtain the global feature vector by flattening and regularizing the feature map; extract the global feature vector using multiple layers of neurons, and map the features to the operating mode or fault state for classification; generate the probability distribution of each category through the Softmax activation function, and determine the operating state or fault type of the current micro-motor by selecting the maximum probability value, realizing the abnormal sound classification of the micro-motor; among them, the operating mode is the normal state, and the fault states are the rotor imbalance state, the bearing wear state, the gear friction state, and the state of being mixed with foreign objects.

[0124] The above embodiments are the preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A micromotor noise classification method based on multi-scale feature fusion and attention mechanism, characterized by: The following steps are involved: The first step is to collect sound signal data: collect the sound signals generated by the micromotor under different working conditions; the collected sound signals are: normal operation sound signals, rotor imbalance sound signals, bearing wear sound signals, gear friction sound signals and sound signals mixed with foreign matter; The second step is sound signal preprocessing: noise reduction processing is performed on the collected sound signals to obtain sound data to reduce the impact of environmental noise and interference; The third step is to extract the features of the sound data: extract the features of the sound data after the sound signal preprocessing; the extracted features include the spectrum, Mel frequency cepstral coefficient MFCC and zero crossing rate ZCR, and merge them through feature splicing to form a comprehensive feature vector; The fourth step is multi-scale feature fusion: three convolutional neural sub-networks with convolutional kernels of different scales are used to extract multi-level feature information from the comprehensive feature vector to obtain a feature map; channel attention and spatial attention mechanisms are integrated in each convolutional neural sub-network, the importance of each channel feature vector is weighted through the channel attention mechanism, and the focus area of ​​the feature vector is dynamically adjusted in the spatial dimension through the spatial attention mechanism; The fifth step is to classify the features and determine the current operating state or fault type of the micromotor: obtain the global feature vector by flattening and regularizing the feature map; use multi-layer neurons to extract the global feature vector, and map the features to the operating mode or fault state for classification; generate the probability distribution of each category through the Softmax activation function, and determine the current operating state or fault type of the micromotor by selecting the maximum probability value to achieve the classification of micromotor abnormal noise; among them, the operating mode is the normal state, and the fault state is the rotor imbalance state, bearing wear state, gear friction state and foreign matter mixed state.

2. The micromotor abnormal sound classification method based on multi-scale feature fusion and attention mechanism according to claim 1 is characterized in that: In the second step, the VSNLMS algorithm is used to perform noise reduction on the collected sound signal: (1) Gradient term calculation: At each sampling point k, calculate the gradient term h i (k), represents the product of the current error signal and the input signal, which is used for subsequent step size adjustment; h i (k)=d(k)·y(k-i) Among them, h i (k) is the gradient term of the current sampling point, which represents the product of the error and the input signal and is used for the dynamic update of the step size; d(k) represents the current error signal, and y(ki) is the delayed sample of the input signal at the kth moment; (2) Update of step factor: The step factor u(k) is dynamically adjusted at each sampling point to adapt to the current error change. The formula is as follows: u(k+1)=u(k)+α·h i (k)·h i (k-1) Among them, u(k) is the step size factor, which controls the amplitude of weight update and is dynamically adjusted in each sampling period; α is the adjustment rate factor, which is used to control the rate of step size update; (3) Calculation of the upper limit of the step size factor: In order to prevent the filter from being unstable due to the step size factor being too large, VSNLMS sets an upper limit q for the step size. max , the calculation formula is: Among them, y T (k)·y(k) represents the power of the input signal at the kth moment; (4) Range restriction of step factor: The value of step factor u(k) is limited to [q min ,q max ] to ensure stability; the specific constraints are as follows: q max and q min It is the upper and lower limits of the step size factor, ensuring the stability of the step size under different conditions; (5) Weight update: Finally, the filter weights are updated using the adjusted step size factor u(k) to reduce the error and make the filter adaptive to changes in the input signal: z i (k+1)=z i (k)+2·u(k)·h i (k) z i (k) is the weight coefficient of the filter, which is updated according to the step size in each sampling period to minimize the error; z i (k+1) represents the updated value of the i-th weight coefficient at time k+1.

3. The micromotor abnormal sound classification method based on multi-scale feature fusion and attention mechanism according to claim 1 is characterized in that: In the third step, the spectrogram feature is generated by short-time Fourier transform (STFT) to capture the frequency distribution of the sound data to detect low-frequency and high-frequency components: Where S(t,f) represents the spectrum amplitude of the signal at time t and frequency f, representing the value of the spectrum graph; x[n] is the input discrete time signal; w[nt] is the window function, whose center is at time t. The window function is used to intercept a local segment of the signal and limit the analysis range at the current moment; e -j2πfn is the kernel function of Fourier transform, which converts the time domain signal into the frequency domain signal, where j is the imaginary unit, 2πf is the angular frequency.

4. The micromotor abnormal sound classification method based on multi-scale feature fusion and attention mechanism according to claim 1 is characterized in that: In the third step, Mel-frequency cepstral coefficients (MFCCs) are extracted to capture the timbre information of the sound data, which helps to distinguish different fault types: First, the sound data is Fourier transformed to obtain the spectrum; secondly, the power spectrum is converted to the Mel scale through the Mel filter bank to better conform to the auditory characteristics of the human ear; the conversion formula of the Mel scale is: The Mel filter bank is composed of a set of triangular filters. Each Mel filter extracts energy in a different frequency range, and then the output energy E of each filter is m Take the logarithm and perform discrete cosine transform DCT on the result to get the Mel cepstral coefficient MFCC: Among them, MFCC c is the cth MFCC coefficient; M is the number of Mel filters; E m is the energy of the mth Mel filter; C represents the number of extracted MFCC coefficients.

5. The micromotor abnormal sound classification method based on multi-scale feature fusion and attention mechanism according to claim 1 is characterized in that: In the third step, the zero crossing rate ZCR is extracted to capture the roughness or frequency characteristics of the sound data. The calculation formula of the zero crossing rate ZCR in each frame is as follows: Where ZCR is the zero crossing rate of the frame, which indicates the number of times the signal crosses the zero point in the frame; N is the number of sampling points in the frame, that is, the frame length; x[n] is the value of the discrete-time signal at the nth sampling point; sgn(x[n-1] is the sign function, which indicates the positive or negative nature of the signal.

6. The micromotor abnormal sound classification method based on multi-scale feature fusion and attention mechanism according to claim 1 is characterized in that: In the fourth step, the three convolutional neural sub-networks with convolutional kernels of different scales are convolutional neural sub-network 1, convolutional neural sub-network 2 and convolutional neural sub-network 3 respectively; Among them, the convolutional neural network uses larger convolution kernels to capture low-frequency components, including three convolution layers; the comprehensive feature vector first passes through the 24 5×5 convolution kernels of the first convolution layer, followed by 4×2 maximum pooling and ReLU activation function; then passes through the 48 5×5 convolution kernels of the second convolution layer, after maximum pooling and ReLU activation; finally, passes through the 48 5×5 convolution kernels of the third convolution layer, after pooling and ReLU activation, the feature map is output; The convolutional neural network 2 uses medium-sized convolution kernels to focus on medium-frequency features and contains two convolution layers. The comprehensive feature vector first passes through 32 3×3 convolution kernels in the first convolution layer, uses L2 regularization to reduce the risk of overfitting, and undergoes 4×2 maximum pooling and ReLU activation. Then it passes through 48 3×3 convolution kernels in the second convolution layer, and outputs the feature map after ReLU activation and maximum pooling operations. Convolutional neural network 3 uses smaller convolution kernels to capture high-frequency components and contains eight convolution layers. The comprehensive feature vector passes through 32 3×3 convolution kernels in the first two convolution layers, using ReLU activation and 2×2 pooling. Then it passes through 64 3×3 convolution kernels in the next four convolution layers, continuing to use ReLU activation. Finally, it passes through 128 3×3 convolution kernels in the last two convolution layers, and outputs the feature map after ReLU activation. Multi-level feature information extraction of the comprehensive feature vector is achieved through a convolutional neural sub-network with three convolution kernels of different scales.

7. The micromotor abnormal sound classification method based on multi-scale feature fusion and attention mechanism according to claim 6 is characterized in that: In the fifth step, the feature map is flattened and regularized to obtain a global feature vector, which means that the feature maps output by convolutional neural sub-network 1, convolutional neural sub-network 2, and convolutional neural sub-network 3 are flattened, converted into feature vectors and input into the fully connected layer. The fully connected layer uses the Dropout technology for regularization to reduce the risk of overfitting and improve the generalization ability of the model. A global feature vector is obtained by summarizing the fully connected layer.

8. The micromotor abnormal sound classification method based on multi-scale feature fusion and attention mechanism according to claim 7 is characterized in that: In the fifth step, multi-layer neurons are used to extract the global feature vector, and the features are mapped to the operating mode or fault status for classification; the probability distribution of each category is generated by the Softmax activation function, and the current operating state or fault type of the micromotor is determined by selecting the maximum probability value. The classification of micromotor abnormal sound means: the fully connected layer uses multi-layer neurons to calculate the global feature vector, and finally outputs a five-dimensional vector, in which each dimension represents the predicted value of the operating mode or the fault category; the last layer of the fully connected layer uses the Softmax activation function to convert the output five-dimensional vector into the probability distribution of each category, and the sum is 1; the probability value of each category represents the possibility of the operating mode or the fault category; by selecting the maximum probability value in the Softmax output, the operating state or fault type of the current micromotor is determined to achieve the classification of micromotor abnormal sound.

9. The micromotor abnormal sound classification method based on multi-scale feature fusion and attention mechanism according to claim 1 is characterized in that: In the fifth step, after determining the current operating state or fault type of the micromotor, the confusion matrix and t-SNE graph are output, and the probability value of the operating state or fault type is displayed to achieve visual classification results.

10. A micromotor abnormal sound classification device based on multi-scale feature fusion and attention mechanism, characterized in that: include: An acoustic sensor acquisition module is used to collect sound signals generated by the micromotor under different working conditions; The sound signals collected are: normal operation sound signal, rotor imbalance sound signal, bearing wear sound signal, gear friction sound signal and sound signal mixed with foreign matter; Signal preprocessing module: used to perform noise reduction processing on the collected sound signals to obtain sound data, so as to reduce the influence of environmental noise and interference; Feature extraction module: used to extract features from the sound data after sound signal preprocessing; the extracted features include spectrum, Mel frequency cepstral coefficient MFCC and zero crossing rate ZCR, and are fused through feature splicing to form a comprehensive feature vector; Multi-scale convolutional neural network module; used to extract multi-level feature information from the comprehensive feature vector using convolutional neural sub-networks with three different-scale convolution kernels to obtain feature maps; channel attention and spatial attention mechanisms are integrated in each convolutional neural sub-network, the importance of each channel feature vector is weighted through the channel attention mechanism, and the focus area of ​​the feature vector is dynamically adjusted in the spatial dimension through the spatial attention mechanism; Classification and output module: used to obtain the global feature vector by flattening and regularizing the feature map; use multi-layer neurons to extract the global feature vector, and map the features to the operating mode or fault state for classification; generate the probability distribution of each category through the Softmax activation function, and determine the current micromotor operating state or fault type by selecting the maximum probability value to achieve micromotor abnormal sound classification; among which, the operating mode is the normal state, and the fault state is the rotor imbalance state, bearing wear state, gear friction state and foreign matter mixed state.

Citation Information

Patent Citations

  • Sound event positioning and detecting method based on attention mechanism

    CN116543754A

  • Transformer fault diagnosis method and system based on voiceprint and infrared feature fusion

    CN117292716A

  • Belt conveyor carrier roller fault diagnosis method and system based on multi-scale feature fusion and residual mask convolution attention algorithm

    CN117421581A

  • Gearbox fault diagnosis method based on multi-scale dynamic convolutional neural network

    CN118706436A

  • Apparatus of applying medication

    KR1020230023570A

Cited By

  • Fan abnormal sound detection method and device based on artificial intelligence neural network

    CN120853618A

  • Vehicle abnormal sound detection method and device based on convolutional neural network, equipment and medium

    CN121148424A