STA-TFT network-based unsupervised industrial abnormal sound detection method

By building a time-frequency change data set and a time-frequency Transformer network, the problem that the existing technology is difficult to take into account local details and global information is solved, and more efficient and robust industrial anomaly sound detection is achieved.

CN120199276APending Publication Date: 2025-06-24GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510271582.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-09
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing industrial anomaly sound detection methods are difficult to take into account local details and global information, resulting in limited detection accuracy and relying on a large number of abnormal samples for training, resulting in uneven training data.

Method used

An unsupervised anomaly sound detection method based on time-frequency change data set (STA) and time-frequency Transformer (TFT) networks is proposed. By constructing a pseudo-exception data set and time-frequency Transformer network for feature extraction and classification, the accuracy and robustness of anomaly detection are improved.

Benefits of technology

Through joint modeling of simulated exception data sets and time-frequency Transformer networks, the global and local time-frequency characteristics of sound signals can be better captured, and the accuracy and robustness of abnormal detection are improved, and suitable for different types of industrial equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199276A_ABST
    Figure CN120199276A_ABST
Patent Text Reader

Abstract

The invention provides an unsupervised industrial abnormal sound detection method based on an STA-TFT network, and meets the monitoring requirements of intelligent manufacturing equipment. According to the method, a time-frequency change data set (STA) is constructed to simulate anomalies, features are extracted and classified in combination with a time-frequency Transform (TFT) network, and the detection efficiency and stability are improved. The method comprises the following operation processes: firstly, preprocessing collected equipment sound into a logarithmic Mel spectrogram, then generating a pseudo-abnormal data set (STA) by using normal audio time-frequency characteristics, and strengthening model learning; then, the TFT network carries out comprehensive modeling on the sound time-frequency features, and feature weight distribution is optimized by means of extended space attention (EPSA), so that the precision is improved. And finally, a multi-head self-attention mechanism is used for measuring feature relevance, and an anomaly judgment method is used for evaluating the audio anomaly degree. The method gets rid of the dependence of traditional governor learning on abnormal data, enhances the model adaptability, shows efficient and reliable detection capability in various industrial environments, and has wide application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine condition monitoring, and particularly to an unsupervised industrial abnormal sound detection method based on a time-frequency change data set and an attention mechanism, which is applicable to equipment abnormal detection in an intelligent industrial environment. Background Art

[0002] In the era of Industry 4.0, intelligent manufacturing and intelligent operation and maintenance have become important directions for industrial development. Machine Condition Monitoring (MCM) is a key technology to ensure the normal operation of equipment and reduce maintenance costs. Among them, the abnormal detection method based on sound signals has been widely used in the operation state analysis of various industrial equipment. However, traditional abnormal sound detection methods mainly rely on supervised learning and require a large number of abnormal samples for training. In an industrial environment, abnormal sound data is usually difficult to obtain, resulting in serious imbalance in training data.

[0003] Currently, deep learning methods have become the mainstream direction for industrial abnormal sound detection. For example, an Autoencoder (AE) evaluates abnormality by reconstructing the log-Mel spectrogram of normal sound and calculating the reconstruction error. However, AE has certain limitations in capturing complex features of sound signals. In addition, a Convolutional Neural Network (CNN) can model time-frequency information, but it performs poorly in processing long-term dependent signals.

[0004] In recent years, due to its excellent performance in natural language processing and image processing tasks, the Transformer model has been introduced into the abnormal sound detection task. For example, a Transformer-based model can effectively capture the dependencies between time features. However, existing Transformer models are difficult to balance the capture of local details and global information, resulting in limited accuracy of abnormal detection. Therefore, there is an urgent need for an abnormal sound detection method that can combine local and global information simultaneously and take into account computational efficiency. Summary of the Invention

[0005] The present invention mainly aims at the abnormal sound detection requirements in the industrial intelligent manufacturing environment, and proposes an unsupervised industrial abnormal sound detection method based on the STA-TFT network. This method constructs a time-frequency change data set (STA) as pseudo-abnormal data, and combines a time-frequency Transformer (TFT) network for feature extraction and classification, improving the accuracy and robustness of abnormal detection.

[0006] Based on deep learning methods, the method of the present invention is designed as the following steps:

[0007] Step 1: Data preprocessing, extract the log-Mel spectrogram, and prepare for subsequent feature modeling;

[0008] Step 2: Based on Step 1, design and implement a pseudo-anomaly data generation method (STA dataset), and generate training anomaly data by modifying the time-frequency information of normal audio data;

[0009] Step 3: Based on Step 2, design and implement a TFT network, model the time dimension and frequency dimension respectively, and fuse the feature information;

[0010] Step 4: Based on Step 3, construct a classifier, dynamically allocate time-frequency feature weights through an extended spatial attention (EPSA) module, and finally determine whether the input sound is abnormal by predicting the probability.

[0011] This design method simulates abnormal sounds by constructing an STA dataset, extracts time-frequency features with a TFT network, and combines an attention mechanism for classification, improving the accuracy and generalization ability of anomaly detection. Specifically, the above steps include the following sub-contents.

[0012] Specifically, the above steps include the following sub-contents.

[0013] Step 1 specifically includes the following contents:

[0014] First, convert the original audio signal into waveform data, calculate its spectral information using the short-time Fourier transform (STFT), then perform Mel filtering on the spectral data and convert it into a log-Mel spectrogram to enhance the representation ability of human auditory features. Then, set appropriate numbers of FFT components, hop lengths, and Mel filters to optimize the extraction of time-frequency features, making the data more suitable for the training and analysis of subsequent deep learning models.

[0015] Step 2 specifically includes the following contents:

[0016] First, adjust the signal amplitude according to the set enhancement factor (range 0.2 to 2.0) to vary the audio data within different amplitude ranges. Then, randomly select the starting position, frequency band, and time bandwidth in the audio data to ensure that the generated pseudo-anomaly data can cover different time-frequency feature variations. Next, enhance or weaken the audio signal in the selected area to simulate possible abnormal patterns and increase the data diversity. Finally, generate multiple varying versions of the audio data so that the model can better learn the key features of normal and abnormal data, improving the model's adaptability and generalization performance. Step 3: Time-frequency Transformer (TFT) network

[0017] Step 3 specifically includes the following content:

[0018] First, in terms of time-domain modeling, a Transformer encoder is used to model the time dimension, capture long-term dependency information, and enhance the time-series structure through positional encoding to ensure the integrity of feature representation. Second, in terms of frequency-domain modeling, a Transformer is used to encode frequency-domain features, thereby extracting local and global features, and combined with segment embedding to make the model pay more attention to key frequency regions and improve the accuracy of anomaly detection. Finally, in the feature fusion stage, the relationship between time and frequency features is integrated to ensure that the model can adapt to different types of industrial equipment and improve the stability and robustness of detection. Step 4: Classification decision

[0019] Step 4 specifically includes the following content:

[0020] First, the EPSA module is used to enhance the time-frequency feature expression ability. Through horizontal and vertical convolution operations, the model can dynamically adjust the weights of different features and extract key features more accurately. Then, a multi-layer Transformer encoder is adopted to calculate the correlation between time-frequency features through the multi-head self-attention mechanism to improve the accuracy of anomaly detection. Finally, based on the anomaly determination method of prediction probability, the classification probability is calculated, and the anomaly detection decision is made by setting a threshold. The sigmoid activation function is combined to optimize the classification result to ensure the accuracy and robustness of anomaly detection.

[0021] Advantages of the present invention

[0022] (1) The present invention proposes and constructs an unsupervised abnormal sound detection network (STA-TFT) based on a time-frequency change dataset (STA) and a time-frequency Transformer (TFT). By changing random time periods or frequency bands of normal sound data, a pseudo-abnormal dataset is generated, thereby enhancing the model's detection ability for abnormal sounds and being able to adapt to the sound characteristics in different fields, improving the robustness and accuracy of the model;

[0023] (2) The present invention designs and implements a pseudo-abnormal data generation method. By randomly changing the time-frequency domain characteristics of normal sound data, a time-frequency change dataset (STA) is generated, solving the problem of lack of abnormal data. And through data augmentation and domain generalization strategies, the detection ability of the model in different environments is improved, making this method more versatile;

[0024] (3) The present invention designs and constructs a time-frequency Transformer network (TFT). By sequentially modeling time-frequency features, the time-domain and frequency-domain features of sound signals are effectively fused, thereby enhancing the model's ability to capture global and local time-frequency features of sound signals and further improving the performance of abnormal sound detection;

[0025] (4)The present invention proposes a data augmentation method based on Mixup. By performing linear interpolation between the source domain and the target domain, mixed samples are generated, optimizing the generalization ability of the model, making the model perform more stably under different data distributions, reducing the influence of domain shift, and improving the performance of the model.

[0026] (5)The present invention designs and implements a feature classifier. By combining the EPSA module and a multi-layer Transformer encoder, the weights of time-frequency features are dynamically allocated, thereby enhancing the adaptability of the model to different machine types and further improving the classification accuracy and robustness. Description of the Drawings

[0027] Figure 1 Overall framework of the time-frequency variation dataset (STA) and the time-frequency Transformer (TFT) network.

[0028] Figure 2 Schematic diagram of the pseudo-anomaly data generation process.

[0029] Figure 3 Structure diagram of the time-frequency Transformer network (TFT).

[0030] Figure 4 Structure diagram of the feature classifier. Detailed Embodiment

[0031] S1: Data preprocessing. Preprocess the collected operating sounds of the device to generate a log-Melspectrogram.

[0032] S2: Pseudo-anomaly data generation. Generate a time-frequency variation dataset (STA) by changing random time periods or frequency bands of normal sound data to simulate anomaly data.

[0033] S3: Feature extraction. Use the time-frequency Transformer (TFT) network to sequentially model the time and frequency dimensions of the acoustic features and fuse the obtained features.

[0034] S4: Anomaly detection. Dynamically allocate weights through the attention mechanism to distinguish different types of machines, and use the prediction probability to determine whether the input audio is abnormal.

[0035] The specific steps of S1 are as follows:

[0036] S1.1 Read the audio data and convert it into a numerical array to represent the waveform signal.

[0037] S1.2 Use the short-time Fourier transform (STFT) to obtain spectral information and convert it into a logarithmic Mel spectrogram;

[0038] S1.3 Set the number of FFT components, hop length, and number of Mel filters to optimize the extraction of time-frequency features.

[0039] According to the method described in claim 1, wherein step S2 specifically includes the following steps:

[0040] S2.1 Randomly select the starting position, frequency, and time-bandwidth range in the normal sound samples;

[0041] S2.2 Adjust the signal amplitude according to the set enhancement factor to perform enhancement or weakening operations on the randomly selected time-frequency region;

[0042] S2.3 Generate a pseudo-abnormal data set through data augmentation methods to ensure that the model can learn the key differences between normal and abnormal data.

[0043] The said S3 specifically includes the following steps:

[0044] S3.1 Use segment embedding and positional encoding to input the time-frequency feature data into the Transformer network;

[0045] S3.2 Through multiple layers of Transformer encoders, use the self-attention mechanism to capture long-range dependencies;

[0046] S3.3 Combine frequency-domain and time-domain information to extract local and global features and improve the feature expression ability.

[0047] The said S4 specifically includes the following steps:

[0048] S4.1 Use the EPSA module to combine horizontal and vertical convolution operations to enhance time-frequency features;

[0049] S4.2 Calculate the dependence relationship between features through the multi-head self-attention mechanism to improve the accuracy of anomaly detection;

[0050] S4.3 Calculate the classification probability through the sigmoid activation function and use the anomaly score to determine the sound state.

[0051] Wherein, the method is applicable to the DCASE 2022 Task 2 data set and supports the detection of the sounds of various industrial devices, including the sounds of different types of machines such as toy cars, fans, bearings, gearboxes, sliders, valves, etc.

[0052] The present invention proposes an unsupervised industrial abnormal sound detection method based on the STA-TFT network. This method constructs a time-frequency change dataset (STA), and uses the time-frequency Transformer (TFT) network for feature extraction and classification to improve the accuracy and robustness of abnormal detection. First, the running sound of the device is preprocessed to generate a log-Mel spectrogram, and then a pseudo-abnormal dataset (STA) is constructed based on this. By randomly modifying the time-frequency features of normal audio data, abnormal situations are simulated. Subsequently, the TFT network is used to jointly model the time-frequency information of the sound signal, and an extended spatial attention (EPSA) module is adopted to dynamically adjust the weights of time-frequency features to improve the classification accuracy. Finally, combined with the multi-head self-attention mechanism and the abnormal determination method, the sound state of the device is intelligently analyzed to achieve efficient and accurate abnormal detection.

[0053] This method can effectively solve the problem of strong dependence of traditional supervised methods on abnormal data, and at the same time improve the adaptability of the model to the sounds of different industrial devices. Experimental results show that the method of the present invention can achieve efficient and reliable abnormal sound detection in multiple industrial scenarios, and has high practical application value. For the specific implementation manners of the present invention, it should be noted that for those of ordinary skill in the art, without departing from the inventive concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. An unsupervised industrial abnormal sound detection method based on STA-TFT network, characterized in that: This method uses the attention mechanism and time-frequency variation dataset to detect abnormal sounds, including the following steps: S1: Data preprocessing: preprocess the collected device operation sound to generate a log-Melspectrogram; S2: Pseudo abnormal data generation, by changing the random time period or frequency band of normal sound data, a time-frequency variation data set (STA) is generated to simulate abnormal data; S3: Feature extraction, using the time-frequency Transformer (TFT) network to sequentially model the time and frequency dimensions of acoustic features and fuse the obtained features; S4: Anomaly detection, dynamically assigns weights through the attention mechanism, distinguishes different types of machines, and uses the predicted probability to determine whether the input audio is abnormal.

2. According to claim 1, an unsupervised industrial abnormal sound detection method based on STA-TFT network is characterized in that: The S1 specifically includes the following steps: S1.1 reads the audio data and converts it into a numerical array to represent the waveform signal; S1.2 uses short-time Fourier transform (STFT) to obtain spectrum information and convert it into a logarithmic Mel spectrum graph; S1.3 Set the number of FFT components, hop length, and number of Mel filters to optimize the extraction of time-frequency features.

3. The unsupervised industrial abnormal sound detection method based on STA-TFT network according to claim 1 is characterized in that: Step S2 specifically includes the following steps: S2.1 Randomly select the starting position and frequency and time bandwidth range in the normal sound sample; S2.2 adjusts the signal amplitude according to the set enhancement factor, and enhances or weakens the randomly selected time-frequency region; S2.3 Generate pseudo-abnormal datasets through data augmentation methods to ensure that the model can learn the key differences between normal and abnormal data.

4. The unsupervised industrial abnormal sound detection method based on STA-TFT network according to claim 1 is characterized in that: The S3 specifically includes the following steps: S3.1 uses segment embedding and position encoding to input time-frequency feature data into the Transformer network; S3.2 uses a multi-layer Transformer encoder to capture long-range dependencies using a self-attention mechanism; S3.3 Combine frequency domain and time domain information to extract local and global features and improve feature expression capabilities.

5. The unsupervised industrial abnormal sound detection method based on STA-TFT network according to claim 1 is characterized in that: The S4 specifically comprises the following steps: S4.1 uses the EPSA module, combining horizontal and vertical convolution operations to enhance time-frequency features; S4.2 Calculate the dependencies between features through the multi-head self-attention mechanism to improve the accuracy of anomaly detection; S4.3 Calculate the classification probability through the sigmoid activation function, and use the anomaly score to determine the sound status.

6. The unsupervised industrial abnormal sound detection method based on STA-TFT network according to claim 1 is characterized in that: The method is applicable to the DCASE 2022 Task 2 dataset and supports sound detection of various industrial equipment, including different types of machine sounds such as toy cars, fans, bearings, gearboxes, sliders, valves, etc.

7. The unsupervised industrial abnormal sound detection method based on STA-TFT network according to claim 1 is characterized in that: The STA-TFT network can achieve adaptive learning in changing environments of different domains and improve domain generalization capabilities.

Citation Information

Cited By

  • Audio data management system and method based on sound console

    CN120544607A