Machine equipment abnormal sound detection method, medium, equipment and product
Through convolutional neural network technology with data augmentation and multi-feature fusion, the generalization problem of abnormal sound detection of machine equipment in a diverse industrial environment is solved, and efficient abnormal sound detection is achieved, adapting to complex working conditions and reducing the need for labeling data.
Patent Information
- Application Number
- CN202510599683.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-22
AI Technical Summary
Existing abnormal sound detection technology for machinery and equipment has the problem of insufficient generalization capabilities in complex and diverse industrial environments, especially when data distributions vary greatly in different production scenarios, traditional models are difficult to adapt, and obtaining high-quality labeled data is expensive.
Data augmentation technology is used to process raw audio, and multiple convolutional neural networks are constructed to extract FFT spectrum maps, amplitude spectrum maps and logarithmic Meer spectrum features, and the similarity score is calculated through K-Means clustering, combining CBAM attention mechanism and Mixup data augmentation to achieve multi-feature fusion and domain generalization.
It improves the detection performance of the model in complex industrial environments, can effectively deal with diversified working conditions, reduces dependence on labeled data, and improves the accuracy and adaptability of detection.
Smart Images

Figure CN120526801A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of abnormal sound detection, and in particular to a method, medium, equipment and product for detecting abnormal sound of machine equipment. Background Art
[0002] Abnormal sound detection technology mainly includes steps such as acoustic signal acquisition, signal preprocessing, feature extraction, and fault diagnosis. Abnormal sounds from machinery and equipment are usually caused by various reasons such as blade damage, bearing wear, and poor gear meshing. The sound emitted by the equipment at this time is usually quite different from the sound characteristics of the machine when it is working normally. The acoustic signal detection system can be used to identify and monitor faults. At the same time, because acoustic signals contain rich fault information, the equipment can be collected non-contactly through a microphone, which is highly safe and low-cost. Therefore, using the acoustic signals of the equipment to detect abnormal sounds has extremely high application prospects. Diagnostic methods based on acoustic signals are mainly divided into traditional signal analysis methods and deep learning methods based on neural networks.
[0003] Fault diagnosis technology based on traditional signal analysis methods usually relies on multi-dimensional analysis and processing of the collected signals to extract feature information closely related to the equipment operation status, including time domain analysis, frequency domain analysis, and time-frequency domain analysis. By extracting relevant features such as amplitude, frequency, and energy spectrum from complex signals, the equipment fault status can be evaluated and identified. Traditional signal analysis methods have the following problems: (1) They rely on manual feature design and require domain expert experience to extract features (such as time-frequency domain features, MFCC, etc.). This is time-consuming and highly subjective, and may ignore important information or introduce redundant features; (2) They have limited ability to process complex patterns and poor adaptability to nonlinear and non-stationary signals (such as transient shocks in mechanical failures). It is difficult to capture high-order features or long-term dependencies; (3) They are not sensitive to noise and have insufficient robustness. Feature extraction is easily affected by background noise. If preprocessing is insufficient (such as imperfect filtering), it may lead to misjudgment; (4) They are difficult to process high-dimensional data. Traditional models have limited processing capabilities for high-dimensional features (such as multi-channel acoustic signals) and are prone to dimensionality disasters.
[0004] Deep learning technology can automatically learn and extract features from a large number of raw signals, thereby improving the accuracy and real-time performance of fault diagnosis. However, existing abnormal sound detection technologies based on deep learning have the following problems: (1) large data requirements, relying on a large amount of labeled data for training. In the industrial field, obtaining high-quality labeled data (such as rare fault samples) is costly and difficult; (2) overfitting and generalization problems. When the sample size is small or the data distribution is uneven, it is easy to overfit (such as learning noise instead of real features). Cross-scenario migration requires additional domain adaptation technology; (3) it is sensitive to input quality. Changes in signal sampling rate, length or noise may significantly affect performance, requiring strict data alignment and enhancement strategies.
[0005] With the advancement of industrial automation and intelligence, the operating environment of equipment is becoming increasingly complex. Equipment may face a variety of operating conditions, different noise backgrounds, and complex equipment states. This makes it difficult for traditional abnormal sound diagnosis models trained on fixed datasets to cope with diverse operating conditions in real-world applications. Traditional models often assume that training data and test data come from the same distribution. In reality, data distributions in different production scenarios can vary significantly, resulting in poor generalization performance of the model in new environments. In many industrial scenarios, obtaining high-quality labeled data is time-consuming and expensive, especially in diverse production environments, where the cost of labeling data for different scenarios is even higher. Summary of the Invention
[0006] The purpose of the present invention is to solve the problem of generalization of abnormal sound detection models with less labeled data in complex and diverse industrial environments. A method for detecting abnormal sounds in machine equipment is proposed, comprising the following steps:
[0007] S1. Obtain the original audio and labels of the machine device during operation, use data enhancement technology to process the original audio to obtain enhanced audio, and divide it into source domain samples and target domain samples;
[0008] S2, converting the waveform features of the enhanced audio into FFT spectrogram features, amplitude spectrogram features and logarithmic Mel spectrogram features respectively;
[0009] S3, constructing the first convolutional neural network, the second convolutional neural network, and the third convolutional neural network to respectively extract the embedded features of the FFT spectrum feature, the amplitude spectrum feature, and the logarithmic Mel spectrum feature, and connecting the three embedded features to obtain the splicing feature;
[0010] S4. Calculate the samples of each category in the source domain through K-Means clustering to obtain multiple cluster centers. Based on the splicing features of the source domain samples, the minimum value of the cosine distance between the source domain samples and all cluster centers of each category is used as the similarity score of the source domain; based on the splicing features of the target domain samples, the minimum value of the cosine distance between the target domain samples and the target domain sample mean is used as the similarity score of the target domain, and the smaller value of the similarity score of the source domain and the similarity score of the target domain is used as the final anomaly score; the anomaly score range is [0,2]. According to the set threshold a, when the anomaly score is in [0,a], the sample is normal, and when the anomaly score is in (a,2], the sound sample is abnormal.
[0011] Furthermore, the data enhancement technology is used to process the original audio. The random mixed samples are obtained by randomly mixing the samples for data enhancement, which is expressed as:
[0012]
[0013] Among them, y p Represents the mixed label, p represents the smoothing value applied to the original label, N represents the number of sample categories, y(n) represents the category label corresponding to the sample category, and n represents the index of the category.
[0014] Furthermore, the waveform features are converted from the time domain to the frequency domain through fast Fourier transform, the amplitude information of the spectrum is obtained by taking the modulus in the frequency domain, and the positive frequency part of the amplitude information of the spectrum is taken to obtain the FFT spectrum graph features.
[0015] Furthermore, the waveform features are divided into multiple time frames, and short-time Fourier transform processing is performed, and the positive frequency part is extracted to obtain the amplitude spectrum features.
[0016] Furthermore, the waveform features are subjected to short-time Fourier transform, converted to the Mel frequency scale and logarithmized to obtain the logarithmic Mel spectrogram features.
[0017] Furthermore,
[0018] The first convolutional neural network consists of three consecutive one-dimensional convolutional layers and five consecutive dense layers, with a flattening layer between the consecutive one-dimensional convolutional layers and the dense layers;
[0019] The second convolutional neural network consists of a two-dimensional convolutional layer and five consecutive residual blocks. The CBAM attention mechanism is added after the first and second residual blocks. The five consecutive residual blocks are followed by a maximum pooling layer and a flattening layer.
[0020] The third convolutional neural network consists of a two-dimensional convolutional layer and five consecutive residual blocks. The first and second residual blocks are embedded with the CBAM attention mechanism. The five consecutive residual blocks are followed by a maximum pooling layer and a flattening layer.
[0021] The present invention also provides a system for detecting abnormal sound of machine equipment, comprising:
[0022] The data acquisition module is used to obtain the original audio and labels of the machine device during operation, process the original audio using data enhancement technology to obtain enhanced audio, and divide it into source domain samples and target domain samples;
[0023] The feature conversion module is used to convert the waveform features of the enhanced audio into FFT spectrogram features, amplitude spectrogram features and logarithmic Mel spectrogram features respectively;
[0024] A feature splicing module is used to construct the first convolutional neural network, the second convolutional neural network, and the third convolutional neural network to extract the embedded features of the FFT spectrum feature, the amplitude spectrum feature, and the logarithmic Mel spectrum feature, respectively, and connect the three embedded features to obtain the splicing feature;
[0025] The anomaly score calculation module is used to calculate the samples of each category in the source domain through K-Means clustering to obtain multiple cluster centers. Based on the splicing features of the source domain samples, the minimum cosine distance between the source domain samples and all cluster centers of each category is used as the similarity score of the source domain; based on the splicing features of the target domain samples, the minimum cosine distance between the target domain samples and the target domain sample mean is used as the similarity score of the target domain, and the smaller value of the similarity score of the source domain and the similarity score of the target domain is used as the final anomaly score; the anomaly score range is [0, 2]. According to the set threshold a, when the anomaly score is in [0, a], the sample is normal, and when the anomaly score is in (a, 2], the sound sample is abnormal.
[0026] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned abnormal sound detection method for machine equipment is implemented.
[0027] The present invention also proposes an electronic device, comprising a processor and a memory, wherein the processor and the memory are interconnected, wherein the memory is used to store a computer program, the computer program includes computer-readable instructions, and the processor is configured to call the computer-readable instructions to execute the above-mentioned abnormal sound detection method for machine equipment.
[0028] The present invention also provides a computer program product, including a computer program / instruction, characterized in that when the computer program / instruction is executed by a processor, the steps of the above-mentioned abnormal sound detection method for machine equipment are implemented.
[0029] The beneficial effects brought about by the technical solution provided by the present invention are:
[0030] The present invention constructs independent neural networks to extract the embedded features of the audio's FFT spectrogram features, amplitude spectrogram features, and logarithmic Mel-spectrogram features, and performs multi-feature fusion. The independent neural networks effectively integrate the advantages of various features, which helps to capture multi-level information in the data, improves the overall detection performance of the model, can effectively cope with domain generalization situations, and can cope with complex and diverse industrial environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 This is a flow chart of a method for detecting abnormal sound of machine equipment according to an embodiment of the present invention;
[0032] Figure 2 is a structural diagram of a first convolutional neural network according to an embodiment of the present invention;
[0033] Figure 3 is a structural diagram of a second convolutional neural network according to an embodiment of the present invention;
[0034] Figure 4 is a structural diagram of a third convolutional neural network according to an embodiment of the present invention;
[0035] Figure 5 4 is a residual block structure diagram of the CBAM attention mechanism embedded in the third convolutional neural network of an embodiment of the present invention;
[0036] Figure 6 It is a block diagram of an electronic device in an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0037] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0038] The flowchart of the abnormal sound detection method of machine equipment according to the embodiment of the present invention is as follows: Figure 1 , specifically including the following steps:
[0039] S1. Obtain the original audio and labels of the machine device during operation, use data enhancement technology to process the original audio, obtain enhanced audio, and divide it into source domain samples and target domain samples.
[0040] Raw audio typically requires preprocessing before analysis. This typically involves converting the original audio format (such as MP3, WAV, or FLAC) into a suitable format. For example, WAV is a lossless audio format often used as an intermediate format for audio processing because it contains complete audio information, facilitating various operations.
[0041] Sampling rate adjustment: The sampling rate refers to the number of times the audio signal is sampled per second. If the sampling rate of the original audio is not suitable for subsequent processing or playback devices, it can be adjusted to a suitable sampling rate.
[0042] Quantization bit conversion: The number of quantization bits indicates the precision with which the audio signal amplitude at each sampling point is quantized. Common quantization bit counts include 8, 16, and 24 bits. A higher number of quantization bits provides more accurate audio representation, but also consumes more storage space. During preprocessing, the number of quantization bits can be converted based on actual needs.
[0043] Noise reduction: Raw audio may contain various noises, such as background noise and electrical noise. Noise reduction can be achieved by using noise reduction algorithms or filters to reduce the effects of these noises. Common noise reduction methods include spectral analysis-based noise reduction algorithms and wavelet denoising.
[0044] De-reverberation: Excessive reverberation in audio can affect the clarity and intelligibility of the sound. De-reverberation can reduce the effects of reverberation and make the sound clearer by using methods such as deconvolution.
[0045] Normalization: Normalization adjusts the amplitude of an audio signal to a specific range, usually adjusting its maximum amplitude to 1 or -1. This ensures that the audio signal will not be overloaded or distorted during subsequent processing, and also facilitates comparison and processing of different audio signals.
[0046] Framing: For some audio processing tasks, such as speech recognition and audio feature extraction, it is usually necessary to divide the audio signal into several frames for processing. The purpose of framing is to convert the continuous audio signal into a discrete frame sequence so that each frame can be analyzed and processed independently.
[0047] Windowing: Based on framing, a window function is typically applied to each frame to reduce discontinuities at frame boundaries. The window function gradually reduces the amplitude at both ends of the frame, making the transition between frames smoother and reducing problems such as spectral leakage.
[0048] The pre-processed original audio first goes through the improved Mixup data enhancement technology. By randomly mixing the samples, random mixed samples are obtained as simulated abnormal samples. It is necessary to identify the confusion category and mixing coefficient so that the model must correctly identify the original samples and mixed samples, and must be able to distinguish between the original samples and mixed samples. The formula is as follows:
[0049]
[0050] Among them, y p Represents the mixed label, p represents the smoothing value applied to the original label, N represents the number of sample categories, y(n) represents the category label corresponding to the sample category, and n represents the index of the category.
[0051] This data augmentation method effectively avoids overfitting of the model to certain categories while relaxing the strict distinction between non-mixed and mixed data samples. During training, its advantage is that the input feature representation uses the same mixing coefficient regardless of whether the mixing coefficient is used, thus ensuring input consistency.
[0052] S2. Convert the waveform features of the enhanced audio into FFT spectrum features, amplitude spectrum features and logarithmic Mel spectrum features respectively.
[0053] FFT spectrogram features are often used to display the energy distribution of audio signals at different frequencies. The original signal is converted from the time domain to the frequency domain through the Fast Fourier Transform (FFT), and then modulo it to obtain the amplitude information of the spectrum. Finally, the positive frequency part of the result is taken to obtain the FFT spectrogram feature. This feature can provide the global spectrum of the audio signal and is often used in tasks such as audio quality analysis and spectrum feature extraction.
[0054] The amplitude spectrogram feature is a common representation of audio signals in the frequency domain and is widely used in fields such as speech recognition and sound event detection. The extraction process first divides the original signal into multiple time frames and performs a short-time Fourier transform (STFT) on it, transforming it into a signal with time-varying characteristics. Finally, only the positive frequency component is retained to obtain the amplitude spectrogram. Unlike the FFT spectrogram feature, the amplitude spectrogram uses the STFT transform to divide the signal into multiple time windows, obtaining the frequency components at each moment. Therefore, it can simultaneously provide characteristic information of the audio in both the frequency and time domains.
[0055] The log-mel spectrogram feature is a time-frequency representation of an audio signal. It is created by performing a short-time Fourier transform on the audio signal, converting it to the mel-frequency scale, and then taking its logarithm. The log-mel spectrogram effectively captures the human ear's sensitivity to sounds of different frequencies. It not only preserves the audio signal's time and frequency domain characteristics, but also effectively reduces its dimensionality and enhances features relevant to human hearing.
[0056] S3: Construct a first convolutional neural network based on the FFT spectrum features to extract embedded features of the FFT spectrum features. Construct a second convolutional neural network based on the amplitude spectrum features to extract embedded features of the amplitude spectrum features. Construct a third convolutional neural network based on the log-Mel spectrum features to extract embedded features of the log-Mel spectrum features. Concatenate the three embedded features to obtain the concatenated features.
[0057] The structural diagram of the first convolutional neural network of the embodiment of the present invention is referenced Figure 2 , including three consecutive one-dimensional convolutional layers and four consecutive dense layers (Dense), with a flatten layer between the consecutive one-dimensional convolutional layers and dense layers. Each one-dimensional convolutional layer has a LeakyReLU activation function, and the dense layers are followed by a batch normalization layer and a LeakyReLU activation function.
[0058] The structural diagram of the second convolutional neural network of the embodiment of the present invention is referenced Figure 3The network consists of a 2D convolutional layer and five consecutive residual blocks. The 2D convolutional layer is followed by a maximum pooling operation (MaxPooling2D). The first and second residual blocks are followed by the CBAM attention mechanism. The five consecutive residual blocks are followed by a maximum pooling layer and a flattening layer. The 2D convolutional layer first captures local features in the input feature map. Then, by adding consecutive residual blocks with the CBAM attention mechanism and connecting CBAMs across blocks, the network globally weights features across blocks, improving the overall feature modeling capabilities of the network and enabling more refined feature capture during feature extraction. Each residual block contains a two-dimensional convolution layer, a BN (Batch Normalization) layer, and a LeakyReLU activation function. By using different convolution kernels and step sizes, the feature processing process gradually transforms from low-level features to high-level features. Each residual block has a residual connection to enhance the learning ability of the convolution layer, promote the effective transmission of information, and avoid the problem of gradient disappearance or gradient explosion. After processing by multiple residual blocks, all spatial features are flattened (Flattened) after maximum pooling (MaxPooling 2D) and mapped through a fully connected layer to generate embedded features containing the key information of the audio signal. Subsequently, the spatial dimension of the feature is reduced by using a pooling operation and the key information can be extracted.
[0059] The structural diagram of the third convolutional neural network of the embodiment of the present invention is shown in FIG. Figure 4 , including a two-dimensional convolution layer and five consecutive residual blocks, wherein there is a maximum pooling operation (MaxPooling 2D) after the two-dimensional convolution layer, and the CBAM attention mechanism is embedded in the first residual block and the second residual block. There is also a maximum pooling layer (MaxPooling 2D) and a flattening layer (Flatten) after the five consecutive residual blocks. The residual blocks are composed of multiple convolution layers, BN layers, and LeakyReLU activation functions. CBAM is embedded in the first and second residual blocks of the network instead of the cross-block connection part. This method enables the logarithmic Mel spectrum features to play a greater advantage in the processing process. Reference is made to the residual block structure diagram with the CBAM attention mechanism embedded in the third convolutional neural network of the embodiment of the present invention. Figure 5 .
[0060] The residual block consists of two residual connection parts, and the CBAM attention mechanism is located between the two residual connection parts of the residual block.
[0061] The first residual connection part includes a BN layer and two 2D convolutional layers connected in sequence. The first 2D convolutional layer is connected to the LeakyReLU activation function, and the second 2D convolutional layer is connected to the BN layer and the LeakyReLU activation function. The output of the BN layer is jump-connected to the output of the second 2D convolutional layer through a 2D convolutional layer connected to the maximum pooling (MaxPooling 2D). The second residual connection part includes a BN layer and two 2D convolutional layers connected in sequence. The first 2D convolutional layer is connected to the LeakyReLU activation function, and the second 2D convolutional layer is connected to the BN layer and the LeakyReLU activation function. The output of the BN layer is jump-connected to the output of the second 2D convolutional layer of the second residual connection part.
[0062] After initial convolution and pooling, features are processed by residual blocks containing CBAM attention, which improves feature selectivity and avoids potential feature conflicts caused by cross-block connections in this subnetwork. In this subnetwork, placing CBAM within the block better cooperates with residual connections, promoting interaction between input features and convolved features, optimizing information flow within the residual block, and improving feature expression capabilities.
[0063] The sub-clustering AdaCos loss function is used to train the model. By calculating the median angles within and between categories during training, the feature embeddings of each category are reasonably distributed with other categories in the angular space, thereby optimizing the discriminability of the features.
[0064] S4. Calculate the samples of each category in the source domain through K-Means clustering to obtain multiple cluster centers. These multiple cluster centers serve as the multimodal distribution features of the categories. Based on the splicing features of the source domain samples, the minimum cosine distance between the source domain samples and all cluster centers of each category is used as the source domain similarity score; based on the splicing features of the target domain samples, the minimum cosine distance between the target domain samples and the target domain sample mean is used as the target domain similarity score. The smaller value of the source domain similarity score and the target domain similarity score is used as the final anomaly score. Among them, the source domain is a field with a large amount of labeled data and the task has been fully solved, while the target domain is a new field that uses the source domain knowledge to improve its performance. The source domain samples are labeled data, and the target domain samples are unlabeled new scene data.
[0065] The smaller value between the similarity score of the source domain and the similarity score of the target domain is used as the final anomaly score. This can simultaneously consider the similarity information of the source domain and the target domain, avoiding the bias caused by relying on only one domain.
[0066] Cosine distance is used to evaluate the similarity between samples. For the abnormal sound detection task, the cosine distance between the input sound sample and the training set sample is calculated to determine whether the sound is abnormal. The calculation formula is as follows:
[0067]
[0068] CD=1-CS
[0069] Where CS represents cosine similarity, A·B represents the dot product of two feature vectors A and B, and ||A|| and ||B|| represent the moduli of vectors A and B. When calculating the source domain similarity score, A and B clearly represent the feature vectors of the concatenated features of the source domain samples and the cluster center, respectively. When calculating the target domain similarity score, A and B clearly represent the feature vectors of the target domain samples and their mean. CD represents the cosine distance, which has a value range of [0, 2]. The smaller of the source domain similarity score and the target domain similarity score is used as the final anomaly score. The anomaly score range is also [0, 2]. An anomaly score of 0 indicates that the collected sound sample is completely normal. A larger anomaly score indicates a higher degree of anomaly. In practice, the anomaly score threshold a is set to determine whether a sound sample is anomaly. When the anomaly score is in the range of [0, a], the sound sample is normal. When the anomaly score is in the range of (a, 2], the sound sample differs significantly from normal samples and is abnormal.
[0070] The present invention also provides a system for detecting abnormal sound of machine equipment, comprising:
[0071] The data acquisition module is used to obtain the original audio and labels of the machine device during operation, process the original audio using data enhancement technology to obtain enhanced audio, and divide it into source domain samples and target domain samples;
[0072] The feature conversion module is used to convert the waveform features of the enhanced audio into FFT spectrogram features, amplitude spectrogram features and logarithmic Mel spectrogram features respectively;
[0073] A feature splicing module is used to construct the first convolutional neural network, the second convolutional neural network, and the third convolutional neural network to extract the embedded features of the FFT spectrum feature, the amplitude spectrum feature, and the logarithmic Mel spectrum feature, respectively, and connect the three embedded features to obtain the splicing feature;
[0074] The anomaly score calculation module is used to calculate the samples of each category in the source domain through K-Means clustering to obtain multiple cluster centers. Based on the splicing features of the source domain samples, the minimum cosine distance between the source domain samples and all cluster centers of each category is used as the similarity score of the source domain; based on the splicing features of the target domain samples, the minimum cosine distance between the target domain samples and the target domain sample mean is used as the similarity score of the target domain, and the smaller value of the similarity score of the source domain and the similarity score of the target domain is used as the final anomaly score; the anomaly score range is [0, 2]. According to the set threshold a, when the anomaly score is in [0, a], the sample is normal, and when the anomaly score is in (a, 2], the sound sample is abnormal.
[0075] In an exemplary embodiment, a computer-readable storage medium is included. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the above-mentioned abnormal sound detection method for machine equipment is implemented.
[0076] See also Figure 6 In an exemplary embodiment, an electronic device is also included, including at least one processor, at least one memory, and at least one communication bus.
[0077] The memory stores a computer program including computer-readable instructions. The processor calls the computer-readable instructions stored in the memory through a communication bus to execute the above-mentioned abnormal sound detection method for machine equipment.
[0078] In an exemplary embodiment, a computer program product is further included, including a computer program / instruction, characterized in that when the computer program / instruction is executed by a processor, the steps of the above-mentioned abnormal sound detection method for machine equipment are implemented.
[0079] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for detecting abnormal sound of machine equipment, characterized in that: The following steps are involved: S1. Obtain the original audio and labels of the machine device during operation, use data enhancement technology to process the original audio to obtain enhanced audio, and divide it into source domain samples and target domain samples; S2, converting the waveform features of the enhanced audio into FFT spectrogram features, amplitude spectrogram features and logarithmic Mel spectrogram features respectively; S3, constructing the first convolutional neural network, the second convolutional neural network, and the third convolutional neural network to respectively extract the embedded features of the FFT spectrum feature, the amplitude spectrum feature, and the logarithmic Mel spectrum feature, and connecting the three embedded features to obtain the splicing feature; S4. Calculate the samples of each category in the source domain through K-Means clustering to obtain multiple cluster centers. Based on the splicing features of the source domain samples, the minimum value of the cosine distance between the source domain samples and all cluster centers of each category is used as the similarity score of the source domain; Based on the splicing features of the target domain samples, the minimum value of the cosine distance between the target domain samples and the target domain sample mean is used as the target domain similarity score, and the smaller value of the source domain similarity score and the target domain similarity score is used as the final anomaly score; the anomaly score range is [0, 2]. According to the set threshold a, when the anomaly score is in [0, a], the sample is normal, and when the anomaly score is in (a, 2], the sample is abnormal.
2. A method for detecting abnormal sound of machine equipment according to claim 1, characterized in that: Data enhancement is performed by randomly mixing samples to obtain random mixed samples, which can be expressed as: Among them, y p Represents the mixed label, p represents the smoothing value applied to the original label, N represents the number of sample categories, y(n) represents the category label corresponding to the sample category, and n represents the index of the category.
3. The method for detecting abnormal sound of a machine device according to claim 1, characterized in that: The waveform features are converted from the time domain to the frequency domain through fast Fourier transform, and the amplitude information of the spectrum is obtained by taking the modulus in the frequency domain. The positive frequency part of the amplitude information of the spectrum is taken to obtain the FFT spectrum graph features.
4. The method for detecting abnormal sound of machine equipment according to claim 1, characterized in that: The waveform features are divided into multiple time frames, and short-time Fourier transform processing is performed, and the positive frequency part is extracted to obtain the amplitude spectrum features.
5. The method for detecting abnormal sound of machine equipment according to claim 1, characterized in that: The waveform features are short-time Fourier transformed, converted to the Mel frequency scale and logarithmized to obtain the logarithmic Mel spectrogram features.
6. The method for detecting abnormal sound of machine equipment according to claim 1, characterized in that: The first convolutional neural network consists of three consecutive one-dimensional convolutional layers and five consecutive dense layers, with a flattening layer between the consecutive one-dimensional convolutional layers and the dense layers; The second convolutional neural network consists of a two-dimensional convolutional layer and five consecutive residual blocks. The CBAM attention mechanism is added after the first and second residual blocks. The five consecutive residual blocks are followed by a maximum pooling layer and a flattening layer. The third convolutional neural network consists of a two-dimensional convolutional layer and five consecutive residual blocks, where the CBAM attention mechanism is embedded in the first residual block and the second residual block, and the five consecutive residual blocks are followed by a maximum pooling layer and a flattening layer.
7. A machine equipment abnormal sound detection system, characterized in that: include: The data acquisition module is used to obtain the original audio and labels of the machine device during operation, process the original audio using data enhancement technology to obtain enhanced audio, and divide it into source domain samples and target domain samples; The feature conversion module is used to convert the waveform features of the enhanced audio into FFT spectrogram features, amplitude spectrogram features and logarithmic Mel spectrogram features respectively; A feature splicing module is used to construct the first convolutional neural network, the second convolutional neural network, and the third convolutional neural network to extract the embedded features of the FFT spectrum feature, the amplitude spectrum feature, and the logarithmic Mel spectrum feature, respectively, and connect the three embedded features to obtain the splicing feature; The anomaly score calculation module is used to calculate the samples of each category in the source domain through K-Means clustering to obtain multiple cluster centers. Based on the splicing features of the source domain samples, the minimum value of the cosine distance between the source domain samples and all cluster centers of each category is used as the similarity score of the source domain; Based on the splicing features of the target domain samples, the minimum value of the cosine distance between the target domain samples and the target domain sample mean is used as the target domain similarity score, and the smaller value of the source domain similarity score and the target domain similarity score is used as the final anomaly score; the anomaly score range is [0, 2]. According to the set threshold a, when the anomaly score is in [0, a], the sample is normal, and when the anomaly score is in (a, 2], the sound sample is abnormal.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the processor and the memory are interconnected, wherein the memory is used to store a computer program, the computer program includes computer-readable instructions, and the processor is configured to call the computer-readable instructions to execute the method according to any one of claims 1 to 6.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Escalator abnormal sound detection method and system based on domain invariant feature transfer and clustering
CN122511299A
Escalator abnormal sound detection method and system based on domain invariant feature transfer and clustering
CN122511299B