Bearing voiceprint fault detection and diagnosis method for multi-motor operation site

Through the improved Conv-tasnet and CAM++ networks, combined with variational modal decomposition and feature extraction technology, the separation and diagnosis of bearing soundprint faults in multi-motor operation sites have been solved, and efficient and accurate fault monitoring and diagnosis have been achieved, reducing costs.

CN120336825APending Publication Date: 2025-07-18YANSHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510507069.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In multi-motor operation sites, it is difficult for the existing technology to effectively separate and identify bearing soundprint faults in complex environments, especially in the face of environmental noise, voice and equipment interference, resulting in low fault detection and diagnosis efficiency and high cost.

Method used

The improved Conv-tasnet network and the improved CAM++ network are used, combining variational modal decomposition, energy entropy screening, Fbank filter group, EMAtrs network and SplitCAM block to realize the separation and feature extraction of multi-motor soundprint signals, adapting to fault diagnosis at different distances.

Benefits of technology

It improves the accuracy of bearing fault detection, reduces monitoring costs, and can monitor the working conditions of multiple sets of motor bearings at the same time, enhancing the fault diagnosis capability in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336825A_ABST
    Figure CN120336825A_ABST
Patent Text Reader

Abstract

The invention discloses a bearing voiceprint fault detection and diagnosis method for a multi-motor operation site, and belongs to the field of motor bearing voiceprint fault detection and diagnose.The method comprises the steps that firstly, recorded voiceprint signals are decomposed through variational mode decomposition, and components with different frequencies as the center are constructed; secondly, screening the components by using an energy entropy, removing an environmental sound component and a load abnormal sound component, and reconstructing a voiceprint signal; then performing source number decomposition on the reconstructed signal by using an improved Conv-tasnet network, and outputting a decomposed signal; performing feature extraction on each decomposed signal by using an Fbank filter bank; an EMAtrs network is used to carry out adaptive soft thresholding weighting processing on the features; and finally, performing grouping weight convolution on different weight features by using SplitCAM block to obtain features capable of being used for classification, and determining equipment working conditions. According to the method, the multi-motor sound aliasing features can be better separated and extracted, so that the accuracy of fault monitoring and diagnosis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of motor bearing acoustic signature fault detection and diagnosis, especially for acoustic signature fault detection and diagnosis at the multi-motor operation site. Background Art

[0002] There is a wide variety of industrial equipment with complex operating conditions, and the stable operation of the equipment directly affects work efficiency and safety. As a widely used component of industrial equipment, once a motor fails, it may lead to equipment shutdown, production delay, or even safety accidents. Therefore, motor fault detection and prediction are extremely important. Motor faults are mainly divided into electrical faults and mechanical faults. Electrical faults are easily detected by circuit fault detection, while mechanical faults often occur inside the motor protection shell and usually require specific methods for detection. As the area of the motor most prone to mechanical faults, the mechanical fault incidence rate of bearings is as high as 40%.

[0003] For motor bearing fault diagnosis, the most common method at present is to detect based on vibration signals. Compared with vibration signal detection, acoustic signature feature detection has the advantage of not requiring contact, and in the case of a production site with monitoring equipment, it has the convenience of low-cost online monitoring. Shan S et al. proposed the Mel-CNN (Mel Convolutional Neural Network) model for fault diagnosis, using variational mode decomposition technology to process motor noise signals and combining Mel spectrum feature extraction with the CNN deep model. Luo Y et al. proposed an enhanced feature extraction network by combining convolutional filters of different sizes with the Context-aware masking (CAM) attention mechanism. Wang H et al. proposed a fast and efficient network for speaker verification using context-aware masking (CAM++) network, taking the densely connected time delay neural network (D-TDNN) network as its core architecture, enhancing the feature extraction ability of the model through dense connection, and at the same time adopting strategies such as context-aware masking technology and multi-granularity pooling, greatly improving the efficiency of fault diagnosis.

[0004] However, the above technical solutions are still very far from being deployed in the industrial field. For industrial acoustic signature recognition, the difficulty lies in the more complex on-site production situation, including environmental noise, human voices, and mutual interference of multiple on-site devices. There are also problems such as the difficulty in unifying the acoustic signature signals collected by the acquisition equipment due to differences in distance. Summary of the Invention

[0005] The object of the present invention is to overcome the deficiencies of the above-mentioned prior art and provide a bearing acoustic fingerprint fault detection and diagnosis method for multi-motor operation sites. This method provides a sound source separation and full-distance fault diagnosis algorithm for multi-device production sites, which can separate the mixed acoustic fingerprints of multiple production devices and adapt to different recording distances, has strong adaptability, reduces the monitoring cost, improves the fault detection and diagnosis efficiency, and meets the needs of the production site.

[0006] The technical solution adopted by the present invention is as follows: a bearing acoustic fingerprint fault detection and diagnosis method for multi-motor operation sites, which uses an improved Conv-tasnet (Convolutional Time-domain Audio Separation Network) network and an improved CAM++ network, and includes the following steps:

[0007] S1. Input the acoustic fingerprint signal collected by the acquisition device, and use variational mode decomposition to decompose the collected acoustic fingerprint signal to construct components centered on different frequencies;

[0008] S2. Use energy entropy to screen the components, remove the environmental sound components and load abnormal sound components, and reconstruct the acoustic fingerprint signal;

[0009] S3. Use the improved Conv-tasnet network to decompose the source number of the reconstructed signal and output the decomposed signal;

[0010] S4. Use the Fbank filter bank to extract features from each decomposed signal;

[0011] S5. Use the EMAtrs network to perform adaptive soft-thresholding weighting processing on the features;

[0012] S6. Use the SplitCAM block to perform grouped weight convolution on different weighted features to obtain features available for classification;

[0013] S7. Determine the working conditions of the motor bearings according to the comparison between different fault features and each sound source feature.

[0014] A further improvement of the technical solution of the present invention lies in: in step S3, the connection mode of the time-domain convolutional layer in the Conv-tasnet network is improved. The main structure of the Conv-tasnet network includes an encoding layer, a time-domain convolutional network layer with an improved convolutional order, and a decoding layer; the specific improvement method in the time-domain convolutional network layer is: improving a serial convolution in which a group of dilation rates of dilated convolutions are connected in ascending order to a multi-scale serial convolution with multiple groups of different maximum dilation rates of dilated convolutions, and the dilation rates of dilated convolutions within each group are connected in ascending order, then descending order.

[0015] A further improvement of the technical solution of the present invention lies in that: in the step S4, the Fbank filter bank structure is successively: pre-emphasis, Hamming window function, Fourier transform, power spectrum calculation, Mel filter bank. Feature extraction of the decomposed signal is performed through the above process, and feature extraction of the decomposed signal is performed through the above process.

[0016] The low-frequency components in the signal are reduced through pre-emphasis, the high-frequency resolution is improved, and the signal-to-noise ratio is improved. The pre-emphasis calculation formula is:

[0017] y(t) = x(t) - αx(t - 1)

[0018] where α is the pre-emphasis coefficient, usually 0.95 or 0.97, and x(t) and x(t - 1) are the speech signals of the current and the previous sampling points respectively, and it is essentially a high-pass filter;

[0019] The calculation formula of the Hamming window is:

[0020] x w (n) = y(n)·w(n)

[0021]

[0022] where x w (n) is the windowed signal, y(n) is the original signal, w(n) is the value of the Hamming window function, and N is the window length, Take 0.46;

[0023] The signal is transformed from the time domain to the frequency domain, and the audio signal is further observed through the energy distribution on the spectrum. The calculation formula of the FFT transformation is:

[0024]

[0025] where S(k) is the frequency-domain signal, and k and n are the frequency-domain and time indices respectively;

[0026] The calculation formula of the power spectrum is:

[0027]

[0028] Finally, Mel filtering calculation is performed. The conversion formula between Mel frequency and linear frequency is:

[0029]

[0030] The Mel filter bank is a series of triangular filter banks. According to its center frequency and bandwidth, its frequency response can be determined. The center frequency of the Mel triangular filter is 1 and decreases on both sides. Its specific formula is:

[0031]

[0032] Multiply and accumulate the energy spectrum with the frequency of the filter and take the logarithm to obtain E m , which is the Fbank feature of the audio data:

[0033]

[0034] A further improvement of the technical solution of the present invention lies in that: in step S5, the EMAtrs network performs adaptive soft thresholding weighted processing on the features based on the ordinary ResNet network. Its specific structure includes a backbone structure, a ChannelGate module, and an EMA module. The backbone structure is two groups of convolutional structures connected in series. Each group includes a 3×3 convolutional layer, a batch normalization layer, and a ReLU activation function layer, and skip connections are used to connect the backbone input and the backbone output; the Channel Gate module is located between the second group of convolutional structures and the backbone output. The main structure is successively a global average pooling layer, a fully connected layer, a Sigmoid activation function, and a weighting layer; the EMA module is similar in position to the Channel Gate module, and the two are parallel. The main structure is an absolute value processing layer (ABS), a horizontal and vertical pooling layer, a feature connection and convolutional layer, and a cross-space learning layer.

[0035] In the Channel Gate module, the fully connected layer maps the global average value to a new space to generate the channel weight W. The weighting processing layer multiplies the channel weight with the original feature map to achieve adaptive weighting. Its calculation formula is:

[0036]

[0037] where F is the network feature map, represents element-wise multiplication, W is the channel weight, and σ(W) is the Sigmoid activation function;

[0038] The EMA module enhances the feature expression ability through multi-scale feature recombination. The calculation formula of the absolute value processing layer is:

[0039] X abs = |X|

[0040] where X is the feature vector.

[0041] The horizontal and vertical pooling layers perform pooling on the output of the absolute value processing layer in the horizontal and vertical directions respectively. Its calculation formula is:

[0042] X h-pool = Pool(X abs , direcyion = horizontal)

[0043] X v-pool = Pool(Xabs , direcyion = vertical)

[0044] Among them, Pool() is the pooling function;

[0045] The feature connection and convolutional layer connect the features of horizontal and vertical pooling, and perform feature recombination through 1x1 convolution. Its calculation formula is:

[0046] X conv = Conv(Concat(X h-pool , X v-pool ))

[0047] Among them, Concat() is the feature concatenation function, and Conv() is the 1×1 convolution;

[0048] The cross-space learning layer combines X conv with the backbone 3×3 convolutional feature map to achieve cross-space feature interaction and obtain the recombined feature map. Its calculation formula is:

[0049] F final = F conv3×3 + X conv

[0050] Among them, F conv3×3 is the backbone output;

[0051] Finally, the skip connection of the backbone network adds the input feature map, weighted feature map, and recombined feature map to obtain the complete adaptive soft thresholding. Its calculation formula is:

[0052] F output = F input + F wighted + F final

[0053] A further improvement of the technical solution of the present invention is that in the step S6, the SplitCAM block splits the original input features into groups, each group applies a separate convolution kernel for feature extraction, and uses the CAM attention mechanism to correct the separated weights through dense connection operations and softmax operations respectively, and performs dispersed weighted feature combination, which can be expressed as:

[0054]

[0055]

[0056] Among them, S j,i represents the attention weight, σ is the activation function, W z is the learnable weight, b z is the bias, y iis the output feature;

[0057] Use a supervised learning strategy to train the network so that it can learn the characteristics of seven working conditions of the motor bearing, namely normal, outer ring fracture, inner ring fracture, rolling element fracture, less oil, cinder intrusion, and assembly misalignment, through the above network structure; during production and use, extract the unknown working conditions on site, separate the sound source features of each through steps S1 to S6, and compare them with the learned features, and output the working condition label with the highest similarity of each sound source to complete the monitoring of the bearing working condition.

[0058] Due to the adoption of the above technical solutions, the technical effects achieved by the present invention are as follows:

[0059] The method of the present invention improves the connection method of the time-domain convolutional layer of the Conv-tasnet network and improves the ResNet network, and can better separate and extract the multi-motor sound aliasing features, thereby improving the accuracy of fault monitoring and diagnosis.

[0060] The bearing acoustic fingerprint fault detection and diagnosis method for the multi-motor operation site in this application weakens the interference of complex conditions such as production site environmental noise and human voices on the bearing working condition monitoring, and improves the accuracy of the model in identifying the bearing working condition; this application can eliminate the mutual interference between devices, enabling a monitoring device to simultaneously monitor the working conditions of multiple groups of motor bearings, greatly improving the detection efficiency and reducing the deployment cost. Brief Description of the Drawings

[0061] Figure 1 is a flowchart of a bearing acoustic fingerprint fault detection and diagnosis method for a multi-motor operation site according to the present invention;

[0062] Figure 2 is a structural diagram of the improved Conv-tasnet network;

[0063] Figure 3 is a structural diagram of the EMAtrs residual network;

[0064] Figure 4 is a structural diagram of the SplitCAM network.

[0065] Figure 5 is a mixed frequency spectrum diagram of the normal bearing working condition and rolling element fault;

[0066] Figure 6 is the separated frequency spectrum diagram of the normal bearing working condition;

[0067] Figure 7 is the separated frequency spectrum diagram of the bearing rolling element fault;

[0068] Figure 8 is a mixed frequency spectrum diagram of the normal bearing working condition and inner ring fault;

[0069] Figure 9The spectrogram of the bearing under normal operating conditions obtained by separation;

[0070] Figure 10 The spectrogram of the inner ring fault of the bearing obtained by separation. Specific implementation manner

[0071] The present invention will be further described in detail below with reference to examples:

[0072] As Figure 1 shown, for the bearing acoustic fingerprint fault detection and diagnosis method in the multi-motor operation site, the following steps are included:

[0073] S1. Input the acoustic fingerprint signal collected by the acquisition device, including environmental noise, motor aliasing sound, and load abnormal sound. Use variational mode decomposition to decompose the input, and construct components centered on the different characteristic frequencies of environmental noise, motor aliasing sound, and load abnormal sound;

[0074] S2. Use energy entropy to screen the components, and its formula is:

[0075]

[0076] where n is the time index, N is the number of sampling points, p i is the normalized energy distribution, i is the sampling point index, and H is the energy entropy. The characteristics of environmental noise are high frequency and low energy entropy, the characteristics of load abnormal sound are low frequency and low energy entropy, and the characteristics of motor aliasing sound are high frequency and high energy entropy. Use the components that conform to the motor aliasing characteristics to reconstruct the input signal;

[0077] S3. Use the improved Conv-tasnet network to decompose the reconstructed signal by the number of sources, and output the decomposed signal of the number of sources. The number of sources is prior information. The main structure of the Conv-tasnet network includes an encoding layer, a time-domain convolutional network layer with an improved convolutional order, and a decoding layer. The specific improvement method in the time-domain convolutional network layer is: improving a serial convolution in which a group of dilation rates of dilated convolutions are connected in ascending order to a multi-scale serial convolution with multiple groups of different maximum dilation rates of dilated convolutions, and the dilation rates of dilated convolutions within each group are connected in ascending order, then descending order.

[0078] As Figure 2 shown, taking the depth of the dilated convolution k = 2 as an example, the input enters the separator after passing through the encoder to perform 3-scale serial dilated convolutions: the dilation rate of the scale 1 dilated convolution: 1 (d = 1 is the yellow Conv 1×D convolution block); the dilation rate of the scale 2 dilated convolution: 1-2-1 (d = 2 is the pink Conv 1×D convolution block); the dilation rate of the scale 3 dilated convolution: 1-2-4-2-1 (d = 4 is the blue Conv 1×D convolution block). The convolution results are fused and enter the decoder to complete the sound source separation;

[0079] S4. Use the Fbank filter bank to extract features from each decomposed signal. The processing of the input signal by the filter bank is as follows: pre-emphasis, Hamming window function, Fourier transform, power spectrum calculation, and Mel filter bank. Feature extraction of the decomposed signal is performed through the above process.

[0080] The role of pre-emphasis is to reduce the low-frequency components in the signal, improve the high-frequency resolution, and thus improve the signal-to-noise ratio (SNR). Its calculation formula is:

[0081] y(t) = x(t) - αx(t - 1)

[0082] where α is the pre-emphasis coefficient, usually 0.95 or 0.97, and x(t) and x(t - 1) are the speech signals at the current and previous sampling points respectively, which is essentially a high-pass filter.

[0083] The Hamming window avoids the problem of spectral leakage after signal framing. Its calculation formula is:

[0084] x w (n) = y(n)·w(n)

[0085]

[0086] where x w (n) is the windowed signal, y(n) is the original signal, w(n) is the value of the Hamming window function, and N is the window length, taking 0.46. Convert the signal from the time domain to the frequency domain, and further observe the audio signal through the energy distribution on the spectrum. The calculation formula for FFT conversion is:

[0087]

[0088] where S(k) is the frequency-domain signal, and k and n are the frequency-domain and time indices respectively. The calculation formula for the power spectrum is:

[0089]

[0090] Finally, perform Mel filter calculation. The conversion formula between Mel frequency and linear frequency is:

[0091]

[0092] The Mel filter bank is a series of triangular filter banks. According to its center frequency and bandwidth, its frequency response can be determined. The center frequency of the Mel triangular filter is 1 and decreases on both sides. Its specific formula is:

[0093]

[0094] Multiply and accumulate the energy spectrum with the frequency of the filter and take the logarithm to obtain E m , which is the Fbank feature of the audio data:

[0095]

[0096] S5. The EMAtrs network is used to perform adaptive soft-thresholding weighting on the features. The specific structure of the EMAtrs network includes a backbone structure, a Channel Gate module, and an EMA module. The backbone structure is two sets of convolutional structures connected in series. Each set includes a 3×3 convolutional layer, a batch normalization (BN) layer, a ReLU activation function layer, and skip connections are used to connect the backbone input and the backbone output; the Channel Gate module is located between the second set of convolutional structures and the backbone output. The main structures are, in sequence, a global average pooling layer (GAP), a fully connected layer (FC), a Sigmoid activation function, and a weighting layer; the EMA module is similar in position to the Channel Gate module and they are parallel. The main structures are, in sequence, an absolute value processing layer (ABS), horizontal and vertical pooling layers, a feature connection and convolutional layer, and a cross-space learning layer. Among them, the fully connected layer in the Channel Gate module maps the global average value to a new space to generate the channel weight W. The weighting processing layer multiplies the channel weight with the original feature map to achieve adaptive weighting. Its calculation formula is:

[0097]

[0098] where F is the network feature map, denotes element-wise multiplication, W is the channel weight, and σ(W) is the Sigmoid activation function.

[0099] The EMA module enhances the feature expression ability through multi-scale feature recombination. The calculation formula of the absolute value processing layer is:

[0100] X abs = |X|

[0101] where X is the feature vector.

[0102] The horizontal and vertical pooling layers perform pooling on the output of the absolute value processing layer in the horizontal and vertical directions respectively. Their calculation formulas are:

[0103] X h-pool = Pool(X abs , direcyion = horizontal)

[0104] X v-pool = Pool(X abs , direcyion = vertical)

[0105] Among them, Pool() is the pooling function.

[0106] The feature connection and convolutional layer connect the features of horizontal and vertical pooling and perform feature recombination through 1x1 convolution. Its calculation formula is:

[0107] X conv = Conv(Concat(X h-pool , X v-pool ))

[0108] Among them, Concat() is the feature concatenation function, and Conv() is the 1×1 convolution.

[0109] The cross-space learning layer combines X conv with the backbone 3×3 convolutional feature map to achieve cross-space feature interaction and obtain the recombined feature map. Its calculation formula is:

[0110] F final = F conv3×3 + X conv

[0111] Among them, F conv3×3 is the backbone output.

[0112] Finally, the skip connection of the backbone network adds the input feature map, the weighted feature map, and the recombined feature map to obtain the complete adaptive soft thresholding. Its calculation formula is:

[0113] F output = F input + F wighted + F final

[0114] The process of adaptive soft thresholding is as Figure 3 shown. The improved network structure introduces the effect of regularization, making it not overly dependent on the features of certain specific channels, and alleviating the problem of model overfitting to a certain extent. At the same time, empirical mode decomposition (EMA) is introduced. After the input vector is processed by the absolute value, it passes through the EMA module. By performing horizontal and vertical pooling connection and convolution operations on the input feature vector respectively, and performing feature recombination with the feature vector after 3*3 convolution through cross-space learning, the performance of the model can be effectively improved;

[0115] S6. Use the SplitCAM block to perform grouped weight convolution on different weight features to obtain features available for classification. The structure of the SplitCAM block is as Figure 4As shown in the figure, the original input features are grouped and split, and each group applies a separate convolutional kernel for feature extraction. The CAM attention mechanism is used to correct the weight vectors of the separated weights through dense connection operations and softmax operations respectively, and perform decentralized weighted feature combination, which can be expressed as:

[0116]

[0117] where S j,i represents the attention weight, σ is the activation function, W z is the learnable weight, b z is the bias, and y i is the output feature.

[0118] S7. Use a supervised learning strategy to train the network so that it learns the characteristics of seven working conditions of the motor bearing, namely normal, outer ring fracture, inner ring fracture, rolling element fracture, less oil, cinder intrusion, and assembly misalignment, through the above network structure.

[0119] During the production and use process, extract the unknown working conditions on site, separate the characteristics of each sound source through steps S1 to S6 and compare them with the learned characteristics, and output the working condition label with the highest similarity of each sound source to complete the monitoring of the bearing working conditions.

[0120] Taking the rolling element fault of the bearing as an example, there are two motor bearings working simultaneously at the production site. One bearing is working normally and the other bearing has a rolling element fault. 4 seconds of mixed audio is collected. As Figure 5 shown, the characteristics of the mixed audio collected by the two through the audio acquisition device are mixed, and the bearing working conditions cannot be directly monitored. After being processed by this technical solution, the sound sources of the two bearings are separated and identified respectively. As Figure 6 shown, the energy distribution of the separated sound source A is uniform and the frequency is stable, and the network identifies it as a normally working bearing; as Figure 7 shown, the energy of the separated sound source B is concentrated in the middle and low frequencies, and the network identifies it as a rolling element fault.

[0121] Taking the inner ring fault of the bearing as an example, there are two motor bearings working simultaneously at the production site. One bearing is working normally and the other bearing has an inner ring fault. 4 seconds of mixed audio is collected. As Figure 8 shown, the characteristics of the mixed audio collected by the two through the audio acquisition device are mixed, and the bearing working conditions cannot be directly monitored. After being processed by this technical solution, the sound sources of the two bearings are separated and identified respectively. As Figure 9 shown, the energy distribution of the separated sound source A is uniform and the frequency is stable, and the network identifies it as a normally working bearing; as Figure 10 shown, the energy of the separated sound source B is concentrated in the middle frequency, and the network identifies it as an inner ring fault.

Claims

1. A bearing acoustic fingerprint fault detection and diagnosis method for a multi-motor operation site, characterized in that: It includes the following steps: S1. Input the voiceprint signal collected by the acquisition device, decompose the collected voiceprint signal using variational mode decomposition, and construct components centered on different frequencies; S2. Use energy entropy to screen the components, remove the environmental sound components and load abnormal sound components, and reconstruct the voiceprint signal; S3. Use the improved Conv-tasnet network to decompose the source number of the reconstructed signal and output the decomposed signal; S4. Use the Fbank filter bank to extract features from each decomposed signal; S5. Use the EMAtrs network to perform adaptive soft thresholding weighted processing on the features; S6. Use the SplitCAM block to perform grouped weight convolution on different weighted features to obtain features available for classification; S7. Determine the motor bearing condition by comparing different fault features with each sound source feature.

2. According to claim 1, a bearing soundprint fault detection and diagnosis method for a multi-motor operation site is characterized in that: The main structure of the Conv-tasnet network in step S3 includes an encoding layer, a time-domain convolutional network layer with an improved convolutional order, and a decoding layer; the specific improvement method in the time-domain convolutional network layer is: improving a serial convolution in which a set of dilation rates of dilated convolutions are connected in ascending order to a multi-scale serial convolution with multiple groups of different maximum dilation rates of dilated convolutions, and the dilation rates of dilated convolutions within each group are connected in ascending order and then descending order.

3. A bearing acoustic fingerprint fault detection and diagnosis method for a multi-motor operation site according to claim 1, characterized in that: The processing process of the Fbank filter bank for the input signal in step S4 is as follows: pre-emphasis, Hamming window function, Fourier transform, power spectrum calculation, Mel filter bank, and the features of the decomposed signal are extracted through the above process.

4. A bearing acoustic fingerprint fault detection and diagnosis method for a multi-motor operation site according to claim 3, characterized in that: Reduce the low-frequency components in the signal through pre-emphasis, improve the high-frequency resolution, and improve the signal-to-noise ratio. The pre-emphasis calculation formula is: y(t) = x(t) - αx(t - 1) where α is the pre-emphasis coefficient, usually 0.95 or 0.97, and x(t) and x(t - 1) are the voice signals of the current and the previous sampling points respectively, and it is essentially a high-pass filter; The calculation formula of the Hamming window is: x w (n) = y(n) · w(n) where x w (n) is the windowed signal, y(n) is the original signal, w(n) is the value of the Hamming window function, and N is the window length, take 0.46; Convert the signal from the time domain to the frequency domain, and further observe the audio signal through the energy distribution on the spectrum. The calculation formula of the FFT conversion is: where S(k) is the frequency-domain signal, and k and n are the frequency-domain and time indices respectively; The calculation formula of the power spectrum is: Finally, perform Mel filter calculation. The conversion formula between Mel frequency and linear frequency is: The Mel filter bank is a series of triangular filter banks. According to its center frequency and bandwidth, its frequency response can be determined. The center frequency of the Mel triangular filter is 1 and decreases on both sides. Its specific formula is: Multiply and accumulate the energy spectrum with the frequency of the filter and take the logarithm to obtain E m , which is the Fbank feature of the audio data:

5. A bearing acoustic fingerprint fault detection and diagnosis method for a multi-motor operation site according to claim 1, characterized in that: In step S5, the EMAtrs network performs adaptive soft-thresholding weighting processing on features based on the ordinary ResNet network. Its specific structure includes a backbone structure, a Channel Gate module, and an EMA module. The backbone structure is composed of two groups of convolutional structures connected in series. Each group includes a 3×3 convolutional layer, a batch normalization layer, and a ReLU activation function layer, and skip connections are used to connect the backbone input and the backbone output. The Channel Gate module is located between the second group of convolutional structures and the backbone output. Its main structure includes a global average pooling layer, a fully connected layer, a Sigmoid activation function, and a weighting layer in sequence. The EMA module is similar in position to the Channel Gate module and they are parallel. Its main structure includes an absolute value processing layer, horizontal and vertical pooling layers, a feature connection and convolutional layer, and a cross-space learning layer in sequence.

6. A bearing acoustic fingerprint fault detection and diagnosis method for a multi-motor operation site according to claim 5, characterized in that: In the Channel Gate module, the fully connected layer maps the global average value to a new space to generate the channel weight W. The weighting processing layer multiplies the channel weight by the original feature map to achieve adaptive weighting. Its calculation formula is: Among them, F is the network feature map, denotes element-wise multiplication, W is the channel weight, and σ(W) is the Sigmoid activation function; The EMA module enhances the feature expression ability through multi-scale feature recombination. The calculation formula of the absolute value processing layer is: X abs = |X| where X is the feature vector; The horizontal and vertical pooling layers perform pooling on the output of the absolute value processing layer in the horizontal and vertical directions respectively. Its calculation formula is: X h-pool = Pool(X abs , direction = horizontal) X v-pool = Pool(X abs , direction = vertical) where Pool() is the pooling function; The feature connection and convolutional layer connects the features of the horizontal and vertical pooling and performs feature recombination through a 1x1 convolution. Its calculation formula is: X conv = Conv(Concat(X h-pool , X v-pool )) where Concat() is the feature concatenation function and Conv() is the 1×1 convolution; The cross-space learning layer combines X conv with the backbone 3×3 convolutional feature map to achieve cross-space feature interaction and obtain a recombined feature map. The calculation formula is as follows: F final = F conv3×3 + X conv Among them, F conv3×3 is the main output; Finally, the skip connection of the backbone network adds the input feature map, the weighted feature map, and the recombined feature map to obtain a complete adaptive soft-thresholding. Its calculation formula is: F output = F input + F wighted + F final 。 7. A bearing acoustic fingerprint fault detection and diagnosis method for a multi-motor operation site according to claim 1, characterized in that: In step S6, the SplitCAM block splits the original input features into groups, and each group applies a separate convolutional kernel for feature extraction. The CAM attention mechanism is used to correct the weight vectors of the separated weights through dense connection operations and softmax operations respectively to perform decentralized weighted feature combination, which can be expressed as: Among which S j,i represents the attention weight, σ is the activation function, W z is the learnable weight, b z is the bias, and y i is the output feature.

8. A bearing acoustic fingerprint fault detection and diagnosis method for a multi-motor operation site according to claim 1, characterized in that: Use a supervised learning strategy to train the network so that it learns the characteristics of seven working conditions of the motor bearing, namely normal, outer ring fracture, inner ring fracture, rolling element fracture, less oil, coal slag intrusion, and assembly misalignment, through the above network structure. During production and use, extract the unknown working conditions on site, separate the features of each sound source through steps S1 to S6 and compare them with the learned features, and output the working condition label with the highest similarity of each sound source to complete the monitoring of the bearing working conditions.